skip to content

CrashLoopBackOff and ImagePullBackOff

A container that keeps dying leaves a trail: the previous container's logs, an exit code (137 for an OOM kill, 143 for a clean SIGTERM), and a backoff that stretches toward five minutes. Interviewers make you walk that ladder without guessing.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

6

A pod is stuck in ImagePullBackOff, and earlier its events showed ErrImagePull. What are the usual causes of a failed container image pull in Kubernetes, and how do you confirm which one you are hitting?

level: juniorimportance: must knowfreq 76%

answer

  1. ErrImagePull = first failure · ImagePullBackOff = retry backoff
  2. Read the event text — it names the cause
  3. unauthorized → imagePullSecrets, same namespace, dockerconfigjson
  4. manifest unknown → tag never pushed / implicit docker.io/library
  5. toomanyrequests → Docker Hub limit, use a mirror

basics

~20 s

The kubelet cannot fetch the image. Usual causes: wrong repository name, tag or digest; the tag does not exist; a private registry with missing or wrong imagePullSecrets; registry unreachable or rate-limiting. Run kubectl describe pod and read the exact pull error in Events.

solid answer

~60 s

`ErrImagePull` is the first failed attempt; `ImagePullBackOff` is the kubelet backing off before retrying. The event message in `kubectl describe pod` names the cause almost verbatim, so read it before theorising. Common causes: - **Typo or missing tag** — `manifest unknown` / `not found`. Someone deployed `:v1.2.4` that was never pushed, or the CI build failed. - **Private registry without credentials** — `unauthorized` or `authentication required`. The pod needs an `imagePullSecrets` entry referencing a `kubernetes.io/dockerconfigjson` Secret **in the same namespace**, or that secret attached to its ServiceAccount. - **Wrong registry host or a missing project path** — the default is Docker Hub, so `myapp:1.0` is resolved as `docker.io/library/myapp:1.0`. - **Rate limiting** — `toomanyrequests`, typical for anonymous Docker Hub pulls on a big cluster; fix with a pull-through mirror or authenticated pulls. - **Network** — no route to the registry from the nodes, proxy or TLS issues, air-gapped cluster. Confirm by copying the exact image reference and pulling it yourself from a node or a debug pod, then by checking the Secret exists, is in the right namespace and is actually referenced.

code

bash · 8 lines
bash
kubectl create secret docker-registry regcred \
  --docker-server=registry.example.com \
  --docker-username=ci-bot \
  --docker-password="$REG_TOKEN" \
  -n team-a

kubectl get pod web-6c7f9-5tzqp -n team-a \
  -o jsonpath='{.spec.imagePullSecrets[*].name}{"\n"}{.spec.containers[*].image}{"\n"}'

go deeper

for a junior

Name the main causes — wrong image/tag, missing credentials, unreachable registry — and show you would read the describe output first.

for a middle

Explain imagePullSecrets mechanics (namespace scope, dockerconfigjson, ServiceAccount attachment) and the tag-dependent imagePullPolicy default.

for a senior

Distinguish credential-provider setups on managed clusters, rate limiting and mirrors, TLS/proxy failures, and whether the failure is fleet-wide or node-local.

for a principal

Talk about the registry as a hard dependency in the pod-start path: digest pinning, mirrors or caches, pre-pulled base layers, and what happens to autoscaling and incident recovery when the registry is unavailable.

## Two statuses, one problem When the kubelet needs an image it is not holding locally, it asks the container runtime to pull it. A failed attempt surfaces as `ErrImagePull`. After a few failures the kubelet applies the same exponential backoff it uses for crash loops and reports `ImagePullBackOff`. Both mean the container has never started, so there are no application logs and no exit code — all evidence lives in the pod's Events. ``` kubectl describe pod web-6c7f9-5tzqp ... Failed to pull image "registry.example.com/team/web:v2.3.1": rpc error: code = NotFound desc = failed to pull and unpack image ...: manifest unknown ``` That message is the diagnosis. The failure categories map onto it directly. ## Cause 1: the reference is wrong Errors like `manifest unknown`, `not found` or `repository does not exist`. Sources: a tag that was never pushed (the CI image build failed but the deploy went ahead), a typo, or an implicit registry. An image string without a host is resolved against Docker Hub, and a single-segment name gets the `library/` namespace — so `internal-api:1.0` becomes `docker.io/library/internal-api:1.0`, which does not exist. Verify with a registry client (`crane manifest`, `docker manifest inspect`, `skopeo inspect`) using the exact string from the pod spec, copied not retyped. ## Cause 2: authentication `unauthorized`, `authentication required`, `denied`. Private registries need credentials the kubelet can use. The Kubernetes mechanism is a Secret of type `kubernetes.io/dockerconfigjson`, referenced either from the pod spec's `imagePullSecrets` or from the ServiceAccount the pod runs under (which the admission controller then copies into the pod). Two rules trip people up constantly: the Secret is **namespace-scoped**, so a copy is needed in every namespace that pulls the image; and the registry host inside the docker config JSON must match the host in the image reference. Cloud clusters often avoid the Secret entirely via node identity — a kubelet credential provider that mints a token for the cloud registry — in which case the failure is an IAM/role problem on the node pool, not a missing Secret. ## Cause 3: rate limiting and registry availability `toomanyrequests: You have reached your pull rate limit`. Anonymous Docker Hub pulls are limited per source IP, and a cluster behind one NAT address exhausts that quickly, especially during a large rollout or a node scale-up. Remedies: authenticate the pulls, or run a pull-through cache/mirror and point the runtime at it. Registry downtime looks similar but with connection errors; so does a blocked egress path or an intercepting proxy with an untrusted certificate (`x509: certificate signed by unknown authority`). ## Cause 4: policy and caching interactions `imagePullPolicy` decides whether a locally present image is reused. `IfNotPresent` skips the pull when the image is already on the node; `Always` pulls (or at least revalidates) every time; `Never` requires it to be preloaded. The default is `Always` when the tag is `:latest` or omitted, otherwise `IfNotPresent`. Consequences: a mutable tag plus `IfNotPresent` gives you nodes running different builds under the same tag; and with `Never` on a node that lacks the image, you get an `ErrImageNeverPull` status. Pinning by digest (`repo@sha256:...`) removes the ambiguity entirely. ## What is *not* an image pull problem If the image pulls fine but the container exits immediately with `exec format error`, the manifest was for a different CPU architecture (an amd64-only image on arm64 nodes). That is a crash, not a pull failure — though on some multi-arch setups the registry will instead report no matching manifest, which does show up as a pull error. Likewise, a pod stuck because a referenced ConfigMap or Secret is missing reports `CreateContainerConfigError`, not `ImagePullBackOff`. ## A practical checklist 1. `kubectl describe pod` and read the Failed/BackOff event text. 2. Copy the image reference out of `kubectl get pod -o jsonpath='{.spec.containers[*].image}'` and inspect it against the registry yourself. 3. If unauthorized: does the Secret exist **in this namespace**, is its type `kubernetes.io/dockerconfigjson`, is it referenced by the pod or its ServiceAccount, and does its registry host match? 4. If network/TLS: test reachability from a node or a debug pod on the same network path. 5. Check whether every node fails or only some — a subset points to node configuration, credential providers, or a partially rolled-out mirror. The fix is usually one line of YAML or one pushed image, but the failure is worth understanding because the same pull path is the one that breaks during a registry outage, when the whole cluster tries to start pods at once.

  • The imagePullSecret exists in the cluster but pods in one namespace still fail with 'unauthorized'. Why?
    Secrets are namespace-scoped, so a Secret created in `default` is invisible to a pod in `team-a`. Each namespace that pulls from the private registry needs its own copy, either created directly, attached to the namespace's ServiceAccount so it is injected automatically, or replicated by an operator. Also check that the registry host recorded in the dockerconfigjson matches the host in the image reference exactly.
  • What is the default imagePullPolicy, and why does that default depend on the tag?
    It is `Always` when the tag is `:latest` or no tag is given, and `IfNotPresent` for any other explicit tag. The reasoning is that `latest` is assumed mutable, so the kubelet must re-check the registry, while a specific tag is assumed immutable and can be served from the node's cache. Because tags are not actually immutable, pinning by digest is the reliable way to guarantee every node runs the same bits.
  • Your cluster starts failing pulls with 'toomanyrequests' during a node scale-up. What do you do?
    That is registry rate limiting — typically anonymous Docker Hub pulls counted per source IP, and a scaling cluster behind one NAT gateway trips it fast. Short term, authenticate the pulls so the higher authenticated limit applies. Longer term, run a pull-through mirror or a registry cache in your own network and configure the container runtime to use it, which also removes the external dependency from the critical path of every node start.

saying these in an interview costs you the question

  • Guessing at causes instead of reading the exact pull error in the pod's Events
  • Believing an imagePullSecret is cluster-wide rather than namespace-scoped
  • Assuming a container that pulls but fails with 'exec format error' is a pull problem rather than an architecture mismatch
  • Thinking imagePullPolicy: Always is the default for every image regardless of tag
  • Deleting and recreating the pod repeatedly, which just restarts the same failing pull

context

open as a page

A Kubernetes pod shows the status CrashLoopBackOff. What does that status actually mean, and what is your step-by-step method for finding out why the container keeps dying?

level: middleimportance: must knowfreq 82%

basics

~20 s

It means the container keeps exiting and the kubelet is now waiting, with a growing delay, before restarting it again. It is a symptom, not a cause. Use kubectl describe pod for the last exit code and reason, then kubectl logs --previous for the dead container's output.

open as a page

kubectl describe pod shows a container with Last State: Terminated, Reason: OOMKilled, Exit Code: 137, and a rising restart count. Explain exactly what happened and how you would fix it.

level: middleimportance: must knowfreq 68%

basics

~20 s

The container's processes exceeded the memory limit in its cgroup, so the kernel's OOM killer sent SIGKILL — exit 137 is 128+9. Fix by measuring real memory use, then either raising the limit or making the runtime respect it (for example JVM heap sizing).

open as a page

A pod sits in the status CreateContainerConfigError and never produces any application logs. What class of problem does that status indicate, and how is it different from a container that starts and then crashes?

level: middleimportance: should knowfreq 44%

basics

~20 s

The kubelet cannot assemble the container's configuration, almost always because a referenced ConfigMap, Secret or a specific key inside one does not exist. The container is never created, so there is no exit code and no logs — the detail is in the pod's Events.

open as a page

A container starts normally, serves traffic for about a minute, is then killed and restarted, and after several cycles the pod reports CrashLoopBackOff — yet the application logs show no error before each death. How do you investigate?

level: seniorimportance: should knowfreq 54%

basics

~20 s

A silent kill after a healthy period usually means the kubelet is killing the container because its liveness probe failed. Check kubectl describe pod for Unhealthy events and the probe's path, port, timeout and thresholds, then call the same endpoint from inside the pod.

open as a page

Your Kubernetes platform has had several incidents where a bad rollout or a registry problem left large numbers of pods failing to start. What guardrails would you put in place so that pod startup failures stop becoming outages?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Make failures non-destructive and detectable: readiness-gated rolling updates with a progress deadline and automatic rollback, digest-pinned images served from a registry mirror, admission rules requiring probes and resource limits, and alerts on restart and pull-failure rates rather than on individual pods.

open as a page