skip to content

During a Kubernetes rollout of a 7-replica Deployment on a 210-node cluster, four new pods run but three sit in ImagePullBackOff. How do you find why only some nodes fail?

level: seniorimportance: should knowfreq 44%

answer

  1. status word is not the cause
  2. Failed event carries runtime error
  3. map pods to nodes first
  4. warm cache hides cold-node failures
  5. backoff keyed on pod UID

basics

~20 s

Map the failing pods to their nodes, then read each pod's Failed event for the actual pull error. The error class (not found, 401/403, 429, network or TLS) combined with what differs about those nodes (cache, egress IP, pool, credentials) explains the split.

solid answer

~50 s

`ImagePullBackOff` only means the kubelet is waiting before its next attempt. The pod-and-image backoff starts at 10s and doubles up to 5 minutes. The cause is in the earlier `Failed` event, `Failed to pull image "...": <runtime error>`, and the container's `state.waiting.message` repeats it. First run `kubectl get pods -o wide` to see which nodes the three pods landed on. Then match the error class to a per-node difference. A not-found error is usually the same everywhere, so a split means the working nodes had the image cached under `IfNotPresent`. A 401/403 points at credentials, often a node pool with a different kubelet credential provider. A 429 comes from a registry that limits per source IP, and only nodes that actually pull, behind a busy NAT IP, hit it. A timeout, name-resolution failure or x509 error points at that node pool's network, DNS or CA trust. To confirm, pull from the node itself with `kubectl debug node/...`. After fixing, delete the stuck pods to reset the backoff.

code

bash · 8 lines
bash
kubectl get pods -n payments-authz -l app=payments-authz -o wide
kubectl get events -n payments-authz \
  --field-selector involvedObject.kind=Pod,reason=Failed \
  --sort-by=.lastTimestamp
kubectl debug node/worker-pool-b-0173 -it --image=busybox:1.36
# inside the debug pod:
#   chroot /host
#   crictl pull registry.example.com/payments/authz:2.14.3

go deeper

for a junior

Recall that ImagePullBackOff means a pull failed and the kubelet is waiting to retry, and that the real error is in kubectl describe pod under Events.

for a middle

Explain the ErrImagePull to ImagePullBackOff cycle, the 10s-to-300s backoff, and where the runtime's error text appears in Events and in the container's waiting message.

for a senior

Demonstrate turning 'only some nodes' into a hypothesis about node pool, egress IP, cache warmth, CA trust or architecture, confirm it from the node, and recover without making the incident worse.

for a principal

Discuss which node-level differences a platform should remove, such as uniform credential plugins, egress and CA trust, so that pull failures stop depending on where a pod lands.

## What the two statuses actually tell you When a pull fails, the kubelet sets the container's waiting reason to **`ErrImagePull`**. That means an attempt just failed. It then records a backoff keyed on the **pod UID plus the image**. Until that backoff expires, each sync reports **`ImagePullBackOff`** instead of trying again. The delay starts at **10 seconds** and doubles up to a **300-second** cap. Neither status names the cause. The cause is in the pod's Events: - `Pulling`: an attempt started. - `Failed` with message `Failed to pull image "<ref>": <error from the runtime>`: **this is the line you need**. - `Failed` with message `Error: ErrImagePull`, and later `Error: ImagePullBackOff`. - `BackOff` with message `Back-off pulling image "<ref>"`. This one is a Normal-type event, so a filter on Warning events alone hides it. During backoff, `status.containerStatuses[].state.waiting.message` also carries the previous pull error. That helps once the Events have expired. ## Step 1: map pods to nodes In the payments-authorization rollout, four of the seven new replicas run and three do not. Before reading any error, get the placement: `kubectl get pods -l app=payments-authz -o wide`. If the three failures sit on nodes that share a node pool, zone, NAT gateway or node image, the cause is almost certainly a property of those nodes, not of the Deployment. ## Step 2: classify the pull error | Error class in the `Failed` event | Why it can hit only some nodes | Confirm by | |---|---|---| | Reference **not found** (wrong tag or digest, typo, or no manifest for the node's OS/architecture) | Normally hits every node. A split means the working nodes had the image cached under `IfNotPresent`, or the failing pool has a different CPU architecture | Compare `kubernetes.io/arch` on the nodes; check whether the tag exists in the registry | | **401 / 403** or an access-denied message | That node pool lacks the kubelet credential provider or node identity the others have, or the working nodes reused a cached private image | Compare kubelet `--image-credential-provider-config` across pools; check the pod's `imagePullSecrets` | | **429 Too Many Requests** | The registry counts pulls per source identity. Only nodes that actually pull, often behind one busy NAT IP, run out of budget | Check which egress IP the failing nodes use. The registry's own limits are that registry's topic | | **i/o timeout, name-resolution failure, connection refused** | A firewall rule, proxy setting or DNS resolver differs on that pool | Reproduce from the node itself (below) | | **x509** certificate error | The node image does not trust the registry's CA | Reproduce from the node; compare the node images | | **no space left on device** | The node's image filesystem is full | Check the node's `DiskPressure` condition and disk usage | The node cache is the factor people most often miss. A node that already holds the image under `IfNotPresent` never contacts the registry, so every class above hides on warm nodes and shows up only on cold ones. The kubelet reports up to 50 images per node in `status.images` by default (`nodeStatusMaxImages`), which gives you a quick warm-versus-cold comparison. ## Step 3: reproduce from the failing node If the event is ambiguous, pull from the node itself: 1. Run `kubectl debug node/<node> -it --image=<a debug image>`. The node's root filesystem is mounted at `/host`. 2. Run `chroot /host`, then run the runtime's CLI (for example `crictl pull <ref>`, if the node image includes it) to test reachability and TLS without the pod layer in the way. 3. Compare the result with the same test on one of the four healthy nodes. ## Step 4: recover cleanly - **Fix the cause first.** Retrying against a 429 or a missing tag only adds more backoff. - **Reset the backoff by deleting the stuck pods.** The backoff key includes the pod UID, so the ReplicaSet's replacement pods pull immediately instead of waiting out a delay of up to 5 minutes. - **Keep the rollout safe.** A Deployment's `maxUnavailable` keeps old replicas serving while new ones are stuck. If the fix will take a while, `kubectl rollout undo` returns to the previous image, which the nodes most likely still have cached. - **Watch the pull load.** By default the kubelet pulls one image at a time (`serializeImagePulls: true`). A cold node pulling a large image can look stuck in `Pulling` without failing, which is a different symptom from `ImagePullBackOff`. ## What a strong answer shows - It reads the `Failed` event, not the status word. - It treats "only some nodes" as a clue about **node properties**: cache, pool, egress, CA trust, architecture. - It confirms from the node before changing anything cluster-wide.

  • You fixed the cause, but the three pods still show ImagePullBackOff for a few minutes. Why, and what do you do?
    The kubelet's pull backoff is keyed on pod UID plus image. It started at 10s and doubled toward a 300s cap, so the existing pods wait out their current delay before retrying. Deleting them makes the ReplicaSet create new pods with new UIDs, which pull immediately. Alternatively, wait up to five minutes for the next retry.
  • The failing pods all landed on an arm64 node pool, and the error says the reference was not found. What is going on?
    The tag exists, but its image index has no manifest for linux/arm64, so the runtime on those nodes finds nothing to pull. The amd64 nodes succeed with the same reference. Confirm with the nodes' `kubernetes.io/arch` label. Fix it by publishing a multi-architecture image, or by constraining the Deployment with a `nodeSelector` on that label until one exists.

saying these in an interview costs you the question

  • ImagePullBackOff and ErrImagePull are two different root causes
  • If the tag exists in the registry, the pull error cannot be not found
  • A registry rate limit would fail every node equally
  • Restarting the kubelet is the first step for an ImagePullBackOff
  • The BackOff event is a Warning, so a Warning filter catches everything
  • Waiting longer always resolves ImagePullBackOff on its own