in pyspark, key value pairs are used to define a RDD. But are they conceptually same as indexes in dataframes?
in pyspark, key value pairs are used to define a RDD. But are they conceptually same as indexes in dataframes?
RDDs in a more general way are a distributed collection of elements of data distributed across the cluster. Key/value pairs define a pair RDD, a type of RDD.
Spark Dataframes are inherently unordered so there is no concept of index like in pandas. Each row is treated as an independent collection of structured data.