skip to content

A Kubernetes pod shows the status CrashLoopBackOff. What does that status actually mean, and what is your step-by-step method for finding out why the container keeps dying?

level: middleimportance: must knowfreq 82%

answer

  1. CrashLoopBackOff = symptom, exit code = cause
  2. describe → Last State, Reason, Exit Code
  3. logs --previous for the dead instance
  4. 137 SIGKILL/OOM · 143 SIGTERM · 127 not found · 0 clean exit
  5. No logs at all → process never started

basics

~20 s

It means the container keeps exiting and the kubelet is now waiting, with a growing delay, before restarting it again. It is a symptom, not a cause. Use kubectl describe pod for the last exit code and reason, then kubectl logs --previous for the dead container's output.

solid answer

~60 s

CrashLoopBackOff is the kubelet telling you "this container exited, I restarted it, it exited again, so I am backing off before the next try". The backoff starts around 10s and doubles up to 5 minutes. The real information is one level down. My order is: 1. `kubectl describe pod` — read **Last State: Terminated**, its **Reason** (Error, OOMKilled, Completed) and **Exit Code**, plus Restart Count and the Events at the bottom. 2. `kubectl logs <pod> -c <container> --previous` — the crashed instance's stdout/stderr. Without `--previous` you get the new, possibly empty, container. 3. Interpret the exit code: 0 = process finished cleanly (a one-shot command under restartPolicy Always), 1/2 = application error, 137 = SIGKILL (usually OOMKilled), 143 = SIGTERM, 127 = command not found, 126 = not executable. 4. If there are no logs at all, the process never started — suspect entrypoint, permissions, or a failed mount. 5. If needed, override the command with a sleep, or use `kubectl debug`, to get a shell in the same image and environment.

code

bash · 5 lines
bash
kubectl get pods -o wide
kubectl describe pod api-7d9f8-2xk4l
kubectl logs api-7d9f8-2xk4l -c api --previous
kubectl get events --sort-by=.lastTimestamp | tail -20
kubectl get pod api-7d9f8-2xk4l -o jsonpath='{.status.containerStatuses[0].lastState.terminated}'

go deeper

for a junior

Know the definition (container exits, kubelet restarts with growing delay) and the two commands: describe pod and logs --previous.

for a middle

Be able to read the Last State block, interpret common exit codes (0, 1, 137, 143, 127), and separate instant failures from failures after a period of running.

for a senior

Show a discriminating method: node-local versus cluster-wide, probe-driven versus self-inflicted exits, and techniques to hold the environment open (command override, ephemeral containers) when the container dies too fast to inspect.

for a principal

Frame it as observability and workload modelling: last-terminated logs are lost after enough restarts, so log shipping and restart-rate alerts matter; one-shot work belongs in Jobs; slow starts belong in startupProbes rather than in tolerating restarts.

## What the status literally means A pod's `restartPolicy` defaults to `Always`. When a container's main process exits — for any reason, including a clean exit with code 0 — the kubelet on that node starts it again. If the container keeps exiting quickly, the kubelet stops retrying immediately and inserts a delay: roughly 10s, then 20s, 40s, and so on, capped at 5 minutes. While it is waiting, the container state is `Waiting` with reason `CrashLoopBackOff`, and that string is what surfaces in `kubectl get pods`. So CrashLoopBackOff is never the root cause. It is the scheduler of restarts complaining. The cause is whatever made the process exit. The backoff timer resets once the container has run successfully for long enough (10 minutes in current kubelet behaviour), which is why a pod that is slowly degrading can flip between Running and CrashLoopBackOff. ## Reading the evidence `kubectl describe pod <name>` is the primary tool. In the container block you get: - **State**: `Waiting`, Reason `CrashLoopBackOff` — the current situation. - **Last State**: `Terminated`, with `Reason`, `Exit Code`, `Started` and `Finished` timestamps. The gap between Started and Finished tells you whether the app died instantly (config/entrypoint) or after running for a while (probe, memory, dependency loss). - **Restart Count**. - **Events** at the bottom: `BackOff`, `Unhealthy`, `Failed`, `FailedMount`. Exit codes carry a lot of signal: - `0` — the process finished normally. Common when someone containerises a script or a `docker run` one-shot and deploys it as a Deployment. Nothing is broken except the workload type; a Job is the right object. - `1` / `2` — generic application failure. Look at the logs. - `126` — the entrypoint exists but is not executable; `127` — the entrypoint or a shell it needs is missing (typical on distroless/scratch images where `sh` does not exist). - `137` — 128+9, killed by SIGKILL. If `Reason: OOMKilled`, the container hit its memory limit. If `Reason: Error`, the process ignored SIGTERM and was force-killed after the grace period. - `139` — 128+11, segmentation fault (often a native library or an architecture mismatch). - `143` — 128+15, terminated by SIGTERM. ## Getting the logs of the dead container `kubectl logs <pod>` shows the *current* container, which may have just started and printed nothing. `kubectl logs <pod> --previous` (`-p`) shows the previously terminated instance. Add `-c <container>` in multi-container pods, and remember init containers are addressed the same way. Only the last terminated instance is retained on the node, so after many restarts older evidence is gone — which is one practical argument for shipping logs off-node. If `--previous` returns nothing at all, the process almost certainly never reached the point of writing to stdout/stderr: bad entrypoint, missing shared library, permission denied on a mounted file, or `runAsNonRoot` against an image whose files are root-owned. ## Cause families 1. **Configuration and dependencies** — missing environment variable, malformed config file, database or discovery endpoint unreachable at boot. Fails within a second or two, and usually logs a stack trace. 2. **Resources** — `OOMKilled` with exit 137. The limit is too low, or a runtime is not cgroup-aware. 3. **Probe-driven restarts** — a failing liveness probe makes the kubelet kill a container that is otherwise healthy from the app's own point of view. The tell is a Last State of Terminated with no application error, plus `Unhealthy` events. 4. **Image/entrypoint problems** — 126/127, wrong `command`/`args`, read-only root filesystem, wrong CPU architecture. 5. **Not actually a long-running process** — exit 0. ## Narrowing further Ask whether it fails on every node or one node: `kubectl get pods -o wide` across replicas. Every node points to image or config; one node points to something node-local (a mount, disk pressure, a stale image). Compare with a working environment: same image digest? same ConfigMap/Secret? Same resource limits? When the container dies too fast to inspect, keep the environment but replace the process: temporarily set `command: ["sleep", "3600"]` on a copy of the pod, or use `kubectl debug <pod> --copy-to=debug-pod --set-image=...` / an ephemeral container, then exec in and run the real command by hand. You see the actual error on a terminal instead of guessing from a truncated log. ## Fixing versus masking Raising restart tolerance is not a fix. Legitimate remedies include: an init container or retry loop for a dependency that is genuinely slow to appear; a `startupProbe` so a slow-booting app is not killed before it is ready; correct resource limits; and, for one-shot workloads, using a Job instead of a Deployment.

  • kubectl logs returns nothing for a crash-looping container. What are the possible reasons and what do you do next?
    Either you are reading the freshly started container instead of the dead one — fix that with `--previous` — or the process genuinely never produced output. The second case means it died before running: a missing or non-executable entrypoint (exit 126/127), a failed mount, or a permission problem under a non-root securityContext. Fall back to `kubectl describe pod` events and to running the image with an overridden command so you can execute the entrypoint interactively.
  • What is the restart backoff schedule, and when does it reset?
    The kubelet backs off exponentially, starting around 10 seconds and doubling to a 5-minute ceiling between restart attempts. The counter resets after the container has stayed up successfully for about 10 minutes. That reset explains pods that alternate between Running and CrashLoopBackOff: each successful interval clears the penalty and the cycle starts over.
  • A container exits with code 0 but the pod still restarts. Is that a bug?
    No. With the default `restartPolicy: Always`, the kubelet restarts the container regardless of exit status, so even a clean exit produces a restart loop. It usually means a one-shot script or batch command was deployed as a Deployment. The right fix is to model it as a Job or CronJob, or to make the process genuinely long-running.

It is like a car ignition that keeps failing: the dashboard light saying "waiting before retrying the starter" tells you nothing about the engine. You still have to open the bonnet and read what happened on the last attempt.

saying these in an interview costs you the question

  • Treating CrashLoopBackOff as an error in its own right and looking for a Kubernetes fix rather than an application exit reason
  • Running kubectl logs without --previous, seeing nothing, and concluding the app produces no logs
  • Confusing CrashLoopBackOff with ImagePullBackOff — the container never even started in the second case
  • Assuming exit code 0 cannot cause restarts under the default restartPolicy: Always
  • "Just delete the pod" as a diagnosis, which recreates the same failing spec

context