Why Spark generate many MapPartitionsRDDs within a simple JDBC transformation?

Viewed 100

I'm learning Spark's data linegae, I wrote a simple JDBC read transformation as bellow, and use RDD.toDebugString to get the data lineage.

val paramDf = spark.read.jdbc(url, "(select * from tb limit 500) t", connectionProperties)
System.out.println(paramDf.rdd.toDebugString)

I found that, the result DataFrame's dependencies have 5 RDDs, as bellow

(1) MapPartitionsRDD[4] at rdd at SparkApp.scala:25 []
 |  SQLExecutionRDD[3] at rdd at SparkApp.scala:25 []
 |  MapPartitionsRDD[2] at rdd at SparkApp.scala:25 []
 |  MapPartitionsRDD[1] at rdd at SparkApp.scala:25 []
 |  JDBCRDD[0] at rdd at SparkApp.scala:25 []

Why Spark generate so many MapPartitionsRDDs within a simple JDBC transformation?

0 Answers
Related