GCP - Rolling update max unavailable

Viewed 63

I am trying to understand the reasoning behind the GCP error message. To give you the context,

  • I have 3 instances running 1 instance per zone using managed instance group.
  • I want to do an update. I would like to do the update one by one. So max unavailable should be 1. However GCP does not seem to like it.

How to achieve high availability here if I give max unavailable 3?

enter image description here

2 Answers

The reasoning behind the error is because when you initiate an update to a regional MIG, the Updater always updates instances proportionally and evenly across each zone, as described in the official documentation. If you set the number of instances lower than the number of zones, then the update could not be proportionally and evenly across zones.

Now, as you said, it does not make much sense from the high availability stand point; but this is because you are keeping the instance names when replacing them, and this forces the Replacement method to be RECREATE instead of SUBSTITUTE. The Maximum Surge for the RECREATE method should be 0 and that is because the original VM should be terminated before the new one is created in order to use the same name.

On the other hand, using the SUBSTITUTE method allows configuring a maximum surge that will be enforced during the update process, creating new VMs with a different name before terminating the old ones, and thus always having VMs available.

The recommendation then is to use the SUBSTITUTE method instead to achieve high availability during your Rolling Updates; if for some reason you need to preserve the instance names, then you can achieve high availability by instantiating more than 1 VM per zone.

I don't think that's really achievable, in your context since there is only 1 instance per zone.. in a managed instance group, it would not be highly available if 33% of your instances would be unavailable, so rather it will be 99% and after the update the high availability is on again.

I would suggest giving a good good read to [1] in order to properly understand how MIGs availability is defined on GCP, essentially you could of have had 2 2 2 and then have 2 2 2 update and again 2 2 2.

Also please check [2] As it's a proven example of my 33% statement above.

[1] https://cloud.google.com/compute/docs/instance-groups/regional-migs#provisioning_a_regional_managed_instance_group_in_three_or_more_zones [2]https://cloud.google.com/compute/docs/instance-groups/regional-migs#:~:text=Use%20the%20following%20table%20to%20determine%20the%20minimum%20recommended%20size%20for%20your%20group%3A

Related