Google data proc logs error about insufficient resources but not failing

Viewed 174

I run apache spark java job on google dataproc. The job creates spark context, analyses logs and finally closes the spark context. Then creates another spark context for another set of analysis. This continues for 50-60 times. Sometimes I get the error Initial job has not accepted any resources; check your cluster UI to ensure that workers are registered and have sufficient resources repeatedly.

Based on answers on SO, this occurs when there isn't enough resource available while starting the job. But this usually happens mid job.

I want the dataproc job to error out and exit. But instead the job just logs this error. How can I make the job to fail. Also how can I prevent this error.

1 Answers

This could happen during job execution, because Spark on Dataproc runs driver in client mode (outside of Yarn), only when it needs to launch executors, Spark requests containers for its AppMaster and executors from YARN.

The error simply indicate there is not enough resources, usually you can find in the cluster's monitoring tab YARN Pending Memory > 0. You can manually scale the cluster up 1, or enable autoscaling 2.

Related