I'm learning Spark's data linegae, I wrote a simple JDBC read transformation as bellow, and use RDD.toDebugString to get the data lineage.
val paramDf = spark.read.jdbc(url, "(select * from tb limit 500) t", connectionProperties)
System.out.println(paramDf.rdd.toDebugString)
I found that, the result DataFrame's dependencies have 5 RDDs, as bellow
(1) MapPartitionsRDD[4] at rdd at SparkApp.scala:25 []
| SQLExecutionRDD[3] at rdd at SparkApp.scala:25 []
| MapPartitionsRDD[2] at rdd at SparkApp.scala:25 []
| MapPartitionsRDD[1] at rdd at SparkApp.scala:25 []
| JDBCRDD[0] at rdd at SparkApp.scala:25 []
Why Spark generate so many MapPartitionsRDDs within a simple JDBC transformation?