I'm trying to execute a query on multiple data frames with each frame consists of around 4 parquets except one of them consists of around 1800 parquet files.
The EMR instances are configured to autoscale. When I try to run a query with more than 3 joins, the execution gets stuck.
I tried everything it could, increasing timeouts, enabling shuffling, dynamic allocation and broadcasting . Below is the spark config:
spark.network.timeout=4800
spark.executor.heartbeatInterval=4200
spark.sql.broadcastTimeout=3600
spark.sql.autoBroadcastJoinThreshold=209715200
spark.shuffle.service.enabled=true
spark.dynamicAllocation.enabled=true
This is the output log I keep getting at the end with no further error/exception.
.Logging$class.logInfo(Logging.scala:54)ogger{39} : Shuffle files lost for host: ip-152-20-116-20.eu-central-2.compute.internal (epoch 7)
.Logging$class.logInfo(Logging.scala:54)ogger{39} : Shuffle files lost for host: ip-172-23-216-128.eu-central-1.compute.internal (epoch 8)
.Logging$class.logInfo(Logging.scala:54)ogger{39} : Shuffle files lost for host: ip-172-23-219-85.eu-central-1.compute.internal (epoch 9)
.Logging$class.logInfo(Logging.scala:54)ogger{39} : Shuffle files lost for host: ip-172-23-218-123.eu-central-1.compute.internal (epoch 4)
.Logging$class.logInfo(Logging.scala:54)ogger{39} : Shuffle files lost for host: ip-172-23-216-84.eu-central-1.compute.internal (epoch 5)
.Logging$class.logInfo(Logging.scala:54)ogger{39} : Shuffle files lost for host: ip-172-23-218-159.eu-central-1.compute.internal (epoch 6)
.Logging$class.logInfo(Logging.scala:54)ogger{39} : Shuffle files lost for host: ip-172-23-219-86.eu-central-1.compute.internal (epoch 7)