How to increase resource utilization in Spark cluster

Viewed 215

I'm fairly new to Spark in cluster-mode and want to know what is wrong with my current configuration since resource utilization seems extremely low.

I am using Dataproc in GCP. I am comparing two cases in reading 10+gb data that is split into 100K+ files in JSON format into a Spark dataframe.

  1. 4 vCPU, 16gb memory for both master and 10 worker nodes. 10 executors with 3 vCPU and 11gb memory each are spawned (assuming 1 vCPU reserved for each worker node and 30% memory overhead)

  2. 4 vCPU, 16gb memory for master, 16 vCPU, 64gb memory for 4 worker nodes. 12 executors with 5 vCPU and 15gb memory each are spawned.

In both cases, I tested a few different parallelism constant C, where parallelism = C * <num_executors> * <num_cores_per_executor>. Driver node is set 11gb memory and 3 vCPU.

Case 1 result:

  • 2/10 containers used
  • <memory:7168, vCores:2> out of <memory:110000, vCores:27> available executor resource used

Case 2 result

  • 2/4 containers used
  • <memory:26624, vCores:2> out of <memory:180000, vCores:60> available executor resource used.

In either case, resource utilization (confirmed through YARN console) is extremely low.

I thought I am missing some major configurations, and tried the following but none of them showed any identifiable difference.

{
  "spark.dynamicAllocation.enabled": "false",
  "yarn.scheduler.maximum-allocation-mb": "11 or 15gb",
  "yarn.nodemanager.resource.memory-mb": "11g or 15gb"
  "spark.submit.deployMode": "cluster"
}

What am I missing/doing wrong?

0 Answers
Related