GKE Metadata server errors

Viewed 792

I have a GKE with Workload identity enabled. Most of our workloads use Cloud Storage or Cloud logging GCP packages which means actually using the Workload identity for GCP access.

Recently we’ve started adding Secret Manager to the stack and started encountering random errors for the Metadata Server on workload startup. It happens on different frameworks.

Python:

File "/venv/lib/python3.8/site-packages/google/auth/compute_engine/credentials.py", line 117, in refresh six.raise_from(new_exc, caught_exc) File "<string>", line 3, in raise_from google.auth.exceptions.RefreshError: ("Failed to retrieve http://metadata.google.internal/computeMetadata/v1/instance/service-accounts/default/?recursive=true from the Google Compute Enginemetadata service. Status: 404 Response:\nb'Not Found\\n'", <google.auth.transport.requests._Response object at 0x7f3a3084dd60>)

NodeJS:

failed to initialize. exiting. Error: 16 UNAUTHENTICATED: Failed to retrieve auth metadata with error: Could not refresh access token: network timeout at: http://169.254.169.254/computeMetadata/v1/instance/service-accounts/default/token?scopes=https%3A%2F%2Fwww.googleapis.com%2Fauth%2Fcloud-platform at Object

I’m trying to understand why it's happening.

First, 404 Not Found means we are trying to get metadata which does not exist/deleted. The thing is it recovers a few seconds later so I'm not sure how exactly.

Based on documentation, sometimes it takes some time for the metadata server to be available, and hence the error which ‘recover’ afterwards. So recommendation is to add delays on the app code or using init Containers until the Metadata server is operated.

I wonder if that's really the best approach, to add an init container to all of our workloads, and if it's really our use case as the error code is a bit misleading. Also, not quite sure why its only started when adding the secret manager.

1 Answers

This sometimes happens due to OOM issues on Metadata server. you can check status of the pod running metadata server using:

kubectl -n kube-system describe pods <pod_name>

you can get the pod_name using:

kubectl get pods --namespace kube-system . the pod name will start with a prefix gke-metadata-server-

if you see something like following in output when you describe the pod:

Last State: Terminated

Reason: OOMKilled

then that would indicate OOM issue.

Some mitigations that you can try:

  1. check if you have un-used ServiceAccounts in your cluster and if you can remove em.
  2. check if you are creating too many clients (new one for every API request). sharing clients if possible will reduce token refresh calls to Metadata server thus, saving memory.
  3. check if you can find metadata server's definition under /etc/kubernetes/addons/. if you can, update the memory to increase it and apply the updated config.
Related