Spark driver failed to start within 900 seconds

Viewed 301

We are trying to run a Spark job in Azure Databricks but getting this error:

Failure type: User configuration issue Details: Databricks execution failed with error state: InternalError, error message: INTERNAL_ERROR: The Spark driver failed to start within 900 seconds.

How do we resolve it?

1 Answers

Databricks job has an error message: "The Spark driver failed to start within 900 seconds"

Steps to Debug

  1. Copy the url of the job execution from the Incident and check the job logs.
  2. Sometimes the job page does not contain an error message, for that an entire spark cluster log to be be checked.
  3. Now check both: Standard error and Log4j output

Outcome of the Analysis Such an issue is a sign that Databricks cannot allocate an Azure VM (for Azure) or EC2 instance (for AWS) to start a driver. Often the root cause is lack of free resources (memory) on cloud side in the region where the Databricks workspace is running. This incident is to be resolved by Cloud Engineers.

To check if it solved, please open the job and check that the latest execution is successful or not.

Related