gitlab job pod exits unexpectedly

Viewed 24

I currently have gitlab runner deployed in my kubernetes cluster with 2 replicas.

When I run a job in gitlab, the runners are successful in spawning pods that run the pipeline. But in some cases, after the pipeline runs the job, I suddenly get the error

Running after_script
00:00
Uploading artifacts for failed job
00:00
Cleaning up project directory and file based variables
00:00
ERROR: Job failed (system failure): pods "runner-hzfiusrx-project-37057717-concurrent-21gs8vm" not found

When I have a look at the runner logs, all I see is

 gitlab-runners-exchange-587cdbf898-pkgt2 | grep "runner-hzfiusrx-project-37057717-concurrent-21gs8vm"
WARNING: Error streaming logs exchange/runner-hzfiusrx-project-37057717-concurrent-21gs8vm/helper:/logs-37057717-2986450184/output.log: command terminated with exit code 137. Retrying...  job=2986450184 project=37057717 runner=hzFiusRx
WARNING: Error streaming logs exchange/runner-hzfiusrx-project-37057717-concurrent-21gs8vm/helper:/logs-37057717-2986450184/output.log: pods "runner-hzfiusrx-project-37057717-concurrent-21gs8vm" not found. Retrying...  job=2986450184 project=37057717 runner=hzFiusRx
WARNING: Error while executing file based variables removal script  error=couldn't get pod details: pods "runner-hzfiusrx-project-37057717-concurrent-21gs8vm" not found job=2986450184 project=37057717 runner=hzFiusRx
ERROR: Job failed (system failure): pods "runner-hzfiusrx-project-37057717-concurrent-21gs8vm" not found  duration_s=2067.525269137 job=2986450184 project=37057717 runner=hzFiusRx
WARNING: Failed to process runner                   builds=32 error=pods "runner-hzfiusrx-project-37057717-concurrent-21gs8vm" not found executor=kubernetes runner=hzFiusRx

Im trying to understand the issue here.

My kubernetes runner config is

  [runners.kubernetes]
    host = ""
    bearer_token_overwrite_allowed = true
    image = "ubuntu:20.04"
    namespace = "exchange"
    namespace_overwrite_allowed = ""
    privileged = true
    cpu_request = "5"
    memory_request = "25Gi"

The nodes on which the job pods get scheduled have the following capacity

Capacity:
  attachable-volumes-aws-ebs:  25
  cpu:                         8
  ephemeral-storage:           20959212Ki
  hugepages-1Gi:               0
  hugepages-2Mi:               0
  memory:                      32523380Ki
  pods:                        58

So what exactly might be the issue here ? The cpu and memory dimensioning for the nodes seem correct.

looking at the utilization, everything seems good too

enter image description here

So what might be the issue here ? Is it that kubernetes/gitlab is not gracefully killing the job pod ? Or does it need more memory ?

0 Answers
Related