Why am I able to map a RDD with the SparkContext

Viewed 297

The SparkContext is not serializable. It is meant to be used on the driver, thus can someone explain the following ?

Using the spark-shell, on yarn, and Spark version 1.6.0

val rdd = sc.parallelize(Seq(1))
rdd.foreach(x => print(sc))

Nothing happens on the client (prints executors-side)

Using the spark-shell, local master, and Spark version 1.6.0

val rdd = sc.parallelize(Seq(1))
rdd.foreach(x => print(sc))

Prints "null" on the client

Using pyspark, local master, and Spark version 1.6.0

rdd = sc.parallelize([1])
def _print(x):
    print(x)
rdd.foreach(lambda x: _print(sc))

Throws an Exception

I also tried the following :

Using the spark-shell, and Spark version 1.6.0

class Test(val sc:org.apache.spark.SparkContext) extends Serializable{}
val test = new Test(sc)
rdd.foreach(x => print(test))

Now it finally throws a java.io.NotSerializableException: org.apache.spark.SparkContext


Why does it works in Scala when I only print sc ? Why do I have a null reference when it should have thrown a NotSerializableException (or so I thought ...)

0 Answers
Related