SpiceQA
Questions
Tags
Users
Badges
rdd
162 Questions
Newest
Active
Unanswered
Frequent
More
Score
View
Card
Compact
Best approach to transform Dataset[Row] to RDD[Array[String]] in Spark-Scala?
user_13149806
0
•
asked Jan 8, 2021
3
3
518
apache-spark-dataset
rdd
apache-spark-sql
apache-spark
scala
Why "collect" action in spark triggers data collection to driver?
user_5256983
0
•
asked Dec 26, 2020
1
1
350
rdd
pyspark
apache-spark-sql
apache-spark
hadoop
AWS Glue RDD.saveAsTextFile() raises Class org.apache.hadoop.mapred.DirectOutputCommitter not found
user_913624
0
•
asked Dec 22, 2020
3
3
798
rdd
aws-glue
apache-spark
scala
getPersistentRDDs returns Map of cached RDDs and DataFrames in Spark 2.2.0, but in Spark 2.4.7 - it returns Map of cached RDDs only
user_7618482
0
•
asked Dec 19, 2020
2
1
189
rdd
apache-spark
scala
Differences between persist(DISK_ONLY) vs manually saving to HDFS and reading back
user_8788071
0
•
asked Oct 20, 2020
3
1
530
rdd
apache-spark
How to convert text log which contains partially json string to the structured in pyspark?
user_14280476
0
•
asked Oct 16, 2020
2
2
179
rdd
pyspark
apache-spark-sql
apache-spark
python-3.x
Difference in caching RDD vs. caching DataFrame in Spark
user_12208910
0
•
asked Oct 8, 2020
4
0
93
rdd
apache-spark-sql
apache-spark
How to use forEachPartition on pyspark dataframe?
user_8229534
0
•
asked Sep 9, 2020
2
1
2601
rdd
pyspark
how spark handles out of memory error when cached( MEMORY_ONLY persistence) data does not fit in memory?
user_14158600
0
•
asked Aug 25, 2020
3
1
2073
rdd
apache-spark
out-of-memory
partitioning
caching
Spark: Replicate each row but with change in one column value
user_14122030
0
•
asked Aug 17, 2020
3
3
952
rdd
apache-spark-sql
apache-spark
scala
Prev
Prev
6
7
8
(current)
9
10
Next
Next
Hot Questions