How do I configure pyspark to write to HDFS by default?

Viewed 2996

I am trying to make spark write to HDFS by default. Currently, when I call saveAsTextFile on an RDD, it writes to my local filesystem. Specifically, if I do this:

rdd = sc.parallelize( [1,2,3,4,5] )
rdd.saveAsTextFile("/tmp/sample")

it will write to a file on my local file system called /tmp/sample. But, if I do

rdd = sc.parallelize( [1,2,3,4,5] )
rdd.saveAsTextFile("hdfs://localhost:9000/tmp/sample")

then it saves to the appropriate spot on my local hdfs instance.

Is there a way to configure or initialize spark such that

rdd.saveAsTextFile("/tmp/sample")

will save to HDFS by default?

To answer a commenter below, when I run

hdfs getconf -confKey fs.defaultFS

I see

17/11/28 09:47:18 WARN util.NativeCodeLoader: Unable to load native-hadoop   library for your platform... using builtin-java classes where applicable
hdfs://localhost:9000
3 Answers
Related