Cloud Run finishes but Cloud Scheduler thinks that job has failed

Viewed 1825

I have a Cloud Run service setup and I have a Cloud Scheduler task that calls an endpoint on that service. When the task completes (http handler returns), I'm seeing the following error:

The request failed because the HTTP connection to the instance had an error.

However, the actual handler returns HTTP 200 and successfully exists. Does anyone know what this error means and under what circumstances it shows up?

I'm also attaching a screenshot of the logs.

Imgur

3 Answers

Does your job take longer than 120 seconds? I was having the same issue and figured out node versions prior to 13 has 120 seconds server.timeout limit. I installed node 13 on docker and problem is gone.

  1. Error 503 is returned by the Google Frontend (GFE). The Cloud Run service either has a transient issue, or the GFE has determined that your service is not ready or not working correctly.
  2. In your log entries, I see a POST request. 7 ms later is the error 503. This tells me your Cloud Run application is not yet ready (in a ready state determined by Cloud Run).
  3. One minute, 8 seconds before, I see ReplaceService. This tells me that your service is not yet in a running state and that if you retry later, you will see success.

I've run an incremental sleep test on my FLASK endpoint which returns 200 within 1 min, 2 min and 10 min of waiting time. Having triggered the endpoint via the Cloud Scheduler, the job failed only in the 10 min test. I've found that it was one of the properties of my Cloud Scheduler job causing the failure. The following solved my issue.

gcloud scheduler jobs describe <my_test_scheduler>

There, you'll see a property called 'attemptDeadline' which was set to 180 seconds by default.

You can update that property using:

gcloud scheduler jobs update http <my_test_scheduler> --attempt-deadline 1000s

Ref: scheduler update

Related