What is the difference between"Pod The node had condition: [DiskPressure]" and "The node was low on resource: ephemeral-storage"

Viewed 2309

When the pod is Evicted by disk issue, I found there are two reasons:

  1. The node had condition: [DiskPressure]
  2. The node was low on resource: ephemeral-storage. Container NAME was using 16658224Ki, which exceeds its request of 0.

I found Node conditions for DiskPressure.

What is the difference?

3 Answers

Both reasons are caused by the same error - that the worker node has ran out of disk space. However, the difference is when they are exactly happening.

To answer this question I decided to dig inside Kubernetes source code.

Starting with The node had condition: [DiskPressure] error.

We can find that it is used in pkg/kubelet/eviction/helpers.go file, line 44:

nodeConditionMessageFmt = "The node had condition: %v. "

This variable is used by Admit function in pkg/kubelet/eviction/eviction_manager.go file, line 137.

Admit function is used by canAdmitPod function in pkg/kubelet/kubelet.go file, line 1932:

canAdmitPod function is used by HandlePodAdditions function, also in the kubelet.go file, line 2195.

Comments in the code in Admit and canAdmitPod functions:

canAdmitPod determines if a pod can be admitted, and gives a reason if it cannot. "pod" is new pod, while "pods" are all admitted pods. The function returns a boolean value indicating whether the pod can be admitted, a brief single-word reason and a message explaining why the pod cannot be admitted.

and

Check if we can admit the pod; if not, reject it.

So based on this analysis we can conclude that The node had condition: [DiskPressure] error message happens when kubelet agent won't admit new pods on the node, that means they won't start.


Now moving on to the second error - The node was low on resource: ephemeral-storage. Container NAME was using 16658224Ki, which exceeds its request of 0

Similar as before, we can find it in pkg/kubelet/eviction/helpers.go file, line 42:

nodeLowMessageFmt  =  "The node was low on resource: %v. "

This variable is used in the same file by evictionMessage function, line 1003.

evictionMessage function is used by synchronize function in pkg/kubelet/eviction/eviction_manager.go file, line 231

synchronize function is used by start function, in the same file, line 177.

Comments in the code in the synchronize and start functions:

synchronize is the main control loop that enforces eviction thresholds. Returns the pod that was killed, or nil if no pod was killed.

Start starts the control loop to observe and response to low compute resources.

So we can conduct that the error The node was low on resource: error message happens when kubelet agent decides to kill currently running pods on the node.


It is worth emphasising that both error messages comes from node conditions (which are set in function synchronize, line 308, to the values detected by eviction manager). Then kubelet's agent makes the decisions that results in these two error messages.


To sum up:

Both errors are due to insufficient disk space, but:

  • The node had condition: error is related to the pods that are about to start on the node, but they can't
  • The node was low on resource: error is related to the currently running pods that must be terminated

Both reasons refer to the same cause where the underlying worker node has ran out of disk space.

For the second error you mentioned, it's caused by too much stuff in emptyDir or too much logs of the pods

But I have no idea if the second error could also trigger the first one...

Related