spark thrift server sql not scheduled for running for long time

Viewed 29

I am running spark thrift on EMR (6.6), with managed scaling enabled. from time to time we have SQL that stack for a long time (45m) until a new request comes to the server and releases it.

when that happens we see that there is one executor on a task node that EMR ask to kill.

What could be the reason for such behavior? How could it be avoided?

1 Answers

it turned out that AWS has a feature that prevents Spark from sending tasks to executors that run on DECOMMISSIONING node.

so in our case, we have min-executor = 1 and the last one was on DECOMMISSIONING node. so spark does not send to it any tasks but it does not ask for new resources because it has that executor.

so that seems to be like an EMR bug.

Related