Google CloudSQL for PostgreSQL HA cluster downtime due to maintenance without failover

Viewed 1738

This morning we experienced a downtime of a little over 5 minutes on our Google CloudSQL for PostgreSQL High Available (HA) cluster. This was during the maintenance period that Google requires you to provide.

Google is clear on why they need the maintenance window (see here). What struck us was the duration of the downtime and that no failover was performed.

The documentation is clear on that the maintenance is performed on an instance (and not on the cluster as a whole). So why was the fallback not performed like is documented here? It could take up to 60 seconds, they say. But it took a little over 5 minutes.

And then again; it is a scheduled maintenance. Automated failover should not have to take place if you anticipate.

Did we misinterpret the documentation, do we have unrealistic expectations or did we misconfigure our application?

1 Answers

As described in the document that you are referring, it's intended only for the event of an instance or zone failure. In other words, only if the instance fails (become unresponsive) or if there is an issue in the zone where the MySQL/PostgreSQL instance is located that causes that the instance cannot be reached, then Cloud SQL will automatically switch to serving data from the standby instance.

Also, in the same document is indicated that the primary instance must be in a normal operating state, this is mentioned in the requirements section.

Related