I'm fairly new to Spark in cluster-mode and want to know what is wrong with my current configuration since resource utilization seems extremely low.
I am using Dataproc in GCP. I am comparing two cases in reading 10+gb data that is split into 100K+ files in JSON format into a Spark dataframe.
4 vCPU, 16gb memory for both master and 10 worker nodes. 10 executors with 3 vCPU and 11gb memory each are spawned (assuming 1 vCPU reserved for each worker node and 30% memory overhead)
4 vCPU, 16gb memory for master, 16 vCPU, 64gb memory for 4 worker nodes. 12 executors with 5 vCPU and 15gb memory each are spawned.
In both cases, I tested a few different parallelism constant C, where parallelism = C * <num_executors> * <num_cores_per_executor>. Driver node is set 11gb memory and 3 vCPU.
Case 1 result:
- 2/10 containers used
- <memory:7168, vCores:2> out of <memory:110000, vCores:27> available executor resource used
Case 2 result
- 2/4 containers used
- <memory:26624, vCores:2> out of <memory:180000, vCores:60> available executor resource used.
In either case, resource utilization (confirmed through YARN console) is extremely low.
I thought I am missing some major configurations, and tried the following but none of them showed any identifiable difference.
{
"spark.dynamicAllocation.enabled": "false",
"yarn.scheduler.maximum-allocation-mb": "11 or 15gb",
"yarn.nodemanager.resource.memory-mb": "11g or 15gb"
"spark.submit.deployMode": "cluster"
}
What am I missing/doing wrong?