I'm running a simple groupby on 350GB of data. Since I'm running this on a single node (I'm on an HPC cluster), I requested computing resource of 400GB and then running the spark job by setting spark.driver.memory to 350 GB.
Since it's running on a single node, the Driver node acts as both master and slave. The job is currently taking more than 6 hours to complete. All it does is a simple groupby operation followed by merging it into a single parquet:
val data = spark.read.parquet("path_to_folder/*")
val grouped = data.groupBy("i","j").agg(sum("count").alias("count"))
grouped.write.parquet("output_folder_path")
Is there a way to make this process more optimal. Specifically, is there a way to force the driver node to make multiple slaves even as the driver node is acting as both master and slave (0 workers) so that the grouping is more efficient?