skip to content

A container starts normally, serves traffic for about a minute, is then killed and restarted, and after several cycles the pod reports CrashLoopBackOff — yet the application logs show no error before each death. How do you investigate?

level: seniorimportance: should knowfreq 54%

answer

  1. Killed silently after running = external killer: probe or OOM
  2. Events: 'Liveness probe failed' + 'will be restarted'
  3. timeoutSeconds defaults to 1 — the classic trap
  4. Liveness shallow; dependencies go in readiness
  5. Slow boot → startupProbe, not a huge initialDelaySeconds

basics

~20 s

A silent kill after a healthy period usually means the kubelet is killing the container because its liveness probe failed. Check kubectl describe pod for Unhealthy events and the probe's path, port, timeout and thresholds, then call the same endpoint from inside the pod.

solid answer

~60 s

When the process dies without logging an error, something outside the process killed it. The two candidates are a failing **liveness probe** and an OOM kill; the container status tells them apart (`Reason: OOMKilled` versus events reading `Liveness probe failed: ...` and `Container api failed liveness probe, will be restarted`). For the probe case I check, in order: 1. **Is the probe reachable?** Wrong port or path, or an app bound to `127.0.0.1` instead of `0.0.0.0` — the kubelet dials the pod IP, so a loopback-only listener always fails. 2. **Is it too strict?** `timeoutSeconds` defaults to 1 second and `failureThreshold` to 3. A health endpoint that occasionally takes 1.2s under GC or load will kill a healthy service. 3. **Is it too deep?** A liveness endpoint that checks the database means a database blip restarts every replica at once. Liveness should test only "is this process wedged"; dependency checks belong in readiness. 4. **Is it a slow start?** Use a `startupProbe` rather than stretching `initialDelaySeconds`. Verify by exec'ing into the pod and curling the probe endpoint under load.

code

bash · 4 lines
bash
kubectl describe pod api-7d9f8-2xk4l | grep -A3 -E 'Liveness|Unhealthy|Killing'
kubectl get events --field-selector reason=Unhealthy --sort-by=.lastTimestamp | tail
kubectl exec -it api-7d9f8-2xk4l -- sh -c \
  'for i in 1 2 3 4 5; do time curl -s -o /dev/null -w "%{http_code}\n" -m 1 http://127.0.0.1:8080/healthz; done'

go deeper

for a junior

Know that a liveness probe failure makes the kubelet restart the container, and that describe pod shows Unhealthy events with the probe error.

for a middle

Recite the probe fields and their defaults — especially timeoutSeconds 1 — and explain the difference between liveness, readiness and startup probes.

for a senior

Reason from the timing pattern to the probe arithmetic, test the endpoint under real load, spot synchronised fleet-wide restarts as a shared-dependency signature, and relax the probe as a controlled experiment.

for a principal

Treat probe design as a reliability policy: liveness as last-resort recovery only, dependency checks confined to readiness, defaults enforced by templates or admission, and alerting on restart rate so probes that mask bugs are visible.

## Reading the pattern "Runs fine, then is killed, no error in the logs" is a strong fingerprint. A process that fails on its own terms almost always leaves something behind — a stack trace, a fatal log line, a non-zero exit from its own code path. Silence plus a hard stop means an external signal. The two external killers are the kernel OOM killer (SIGKILL, `Reason: OOMKilled`, exit 137) and the kubelet acting on a failed liveness probe. The kubelet's kill is a SIGTERM followed by SIGKILL after the grace period, so the exit code is 143 or 137 with `Reason: Error`. The decisive evidence is in the events: ``` Warning Unhealthy kubelet Liveness probe failed: Get "http://10.1.4.7:8080/healthz": context deadline exceeded (Client.Timeout exceeded ...) Normal Killing kubelet Container api failed liveness probe, will be restarted ``` The periodicity is a clue too: if the time from start to kill is roughly `initialDelaySeconds + failureThreshold × periodSeconds`, the probe is the cause, and the constant interval across restarts confirms it. ## The probe fields and their defaults - `initialDelaySeconds` (0) — how long before the first check. - `periodSeconds` (10) — how often. - `timeoutSeconds` (**1**) — how long the kubelet waits for a response. This default is the single most common cause of spurious restarts. - `failureThreshold` (3) — consecutive failures before acting. - `successThreshold` (1 for liveness; must be 1). So the default configuration kills a container after three consecutive checks that each take longer than one second, spread over roughly 30 seconds. Any endpoint whose tail latency crosses a second — a JVM in a long GC pause, a thread pool exhausted by real traffic, an event loop blocked by a synchronous call — will trip it exactly when the system is already under stress. That is the worst possible time to remove capacity, and it is how a load spike turns into a rolling restart of every replica. ## Reachability problems The kubelet performs an HTTP probe against the **pod IP** on the container port. Failures follow from: - The app listening on `127.0.0.1` (very common defaults in some frameworks) — connection refused from the kubelet. - A probe `port` naming a container port that does not exist, or a numeric port the app does not serve. - A path that returns 3xx or requires authentication — the kubelet treats 200–399 as success, anything else as failure, and it sends no credentials. - HTTPS with `scheme: HTTP` or vice versa. - A separate admin/management port that the app only opens later in startup. An `exec` probe has its own trap: the command must exist in the image (distroless images have no shell), and each invocation costs a process — a heavy exec probe every few seconds can itself add load. ## Liveness versus readiness versus startup These answer different questions and confusing them causes real outages: - **Liveness** — "is this process unrecoverably stuck?" Failure means *restart me*. It should be shallow, local, and cheap: can the HTTP server accept a connection and return, is the main loop alive. It must **not** check downstream dependencies, because a shared dependency's blip then restarts the entire fleet simultaneously, adding a thundering herd of cold starts to an existing incident. - **Readiness** — "should I receive traffic right now?" Failure removes the pod from Service endpoints without restarting it. Dependency checks belong here, where the consequence is proportionate. - **Startup** — "has boot finished?" While a startup probe is running, liveness and readiness are suspended. This is the correct tool for an application with a long or variable boot (large caches, JIT warm-up, schema migrations) — far better than a large `initialDelaySeconds`, which either under-protects slow boots or delays detection of a genuinely wedged process forever. ## Investigation steps 1. `kubectl describe pod` — Unhealthy/Killing events, and the probe configuration as applied. 2. Compare the restart interval against the probe arithmetic. 3. Exec into a *currently running* replica and call the probe endpoint yourself, ideally repeatedly and while the pod is taking traffic: `kubectl exec -it pod -- curl -sv -m 1 http://127.0.0.1:8080/healthz`. Measure latency, not just status. 4. Reproduce under load — the failure often only appears with concurrency, GC pressure or a cold cache. 5. Check whether all replicas restart at the same moment. Synchronised restarts across the fleet point at a shared dependency inside the liveness check. 6. As an experiment, temporarily remove or relax the liveness probe. If the container then runs indefinitely, the probe was the killer; if it still dies, look elsewhere (memory, a supervisor inside the container, a node-level issue). ## Fixing it well Raise `timeoutSeconds` to something realistic (2–5s), raise `failureThreshold` so transient blips do not count, move dependency checks out of liveness into readiness, add a `startupProbe` for slow boots, and make the health endpoint cheap and independent of business traffic — ideally served from a path that does not queue behind the same thread pool. A liveness probe should fire rarely; if yours is restarting containers regularly, it is being used as a workaround for a bug rather than as a last-resort recovery mechanism.

  • Why is it dangerous for a liveness probe to check a database connection?
    Because the consequence of liveness failure is a restart. If every replica's liveness check depends on the same database, a brief database outage restarts the whole fleet at once — and the restarts add cold starts, reconnection storms and lost in-flight work to an incident that would otherwise have recovered on its own. Dependency health belongs in the readiness probe, whose failure merely removes the pod from Service endpoints and is fully reversible.
  • When would you use a startupProbe instead of simply increasing initialDelaySeconds on the liveness probe?
    Whenever boot time is long or variable. A startupProbe suspends the liveness and readiness probes until it first succeeds, so you can allow generous time to boot (a long failureThreshold) while keeping the liveness probe tight afterwards. A large initialDelaySeconds forces a bad trade instead: it must be as long as the slowest possible boot, which delays detection of a genuinely hung process for that entire window on every restart.

A liveness probe is a defibrillator, not a thermometer. Wiring it to a reading that dips whenever the patient exerts themselves means shocking a perfectly healthy heart every time it beats fast.

saying these in an interview costs you the question

  • Never checking events, and searching the application code for a bug that does not exist
  • Leaving timeoutSeconds at its 1-second default for an endpoint with second-scale tail latency
  • Putting downstream dependency checks in the liveness probe
  • Using a large initialDelaySeconds instead of a startupProbe for slow-starting applications
  • Assuming the kubelet can reach a service bound to 127.0.0.1 inside the container

context