How to properly set spark cluster properties in Databricks

Viewed 59

I have a cluster in Databricks for my spark workflow and I wanted some help in setting up right for optimal use. Following are the details of my cluster.

RUNTIME: 10.4 LTS (includes Apache Spark 3.2.1, Scala 2.12)
DRIVER TYPE: c5a.8xlarge (64GB Memory, 32 Cores)
WORKER TYPE: c5a.4xlarge (32GB Memory, 16 Cores)
(Min worker 1, Max Workers 5) 

This is a going to process large amount of data (not sure about the exact numbers). Here are the current properties that i am using. I think these are not optimal.

spark.driver.extraJavaOptions -Xss64M
spark.executor.cores 7
spark.executor.memory 20G
spark.driver.maxResultSize 20G
spark.sql.shuffle.partitions 35
spark.driver.memory 48G
spark.sql.execution.arrow.pyspark.enabled true
spark.sql.execution.arrow.pyspark.fallback.enabled true
spark.executor.memoryOverhead 1G

Is there a rule or a guide as how to set proper values to get the maximum performance?

0 Answers
Related