Walk through the Kubernetes Pod phases (Pending, Running, Succeeded, Failed, Unknown) and explain what the Pod-level restartPolicy values Always, OnFailure and Never do when a container exits.
answer
- phase is coarse; conditions carry the truth
- Pending = scheduling/image/init
- Running does not mean healthy
- kubelet restarts in place, same Pod IP
- CrashLoopBackOff = symptom, backoff to 5 min
basics
~20 sPending = accepted but not all containers running (scheduling, image pull, init). Running = bound to a node with at least one container running. Succeeded/Failed = all containers terminated, all zero exit codes or not. Unknown = node unreachable. restartPolicy tells the kubelet whether to restart exited containers in place.
solid answer
~50 s`status.phase` is a coarse, one-word rollup: - **Pending** - the Pod object exists but is not fully up: waiting for scheduling, pulling images, or running init containers. - **Running** - bound to a node, all containers created, at least one running or restarting. - **Succeeded** / **Failed** - every container has terminated; Succeeded means all exited 0 and none will restart. - **Unknown** - the node stopped reporting. The real detail lives in `status.conditions` (PodScheduled, Initialized, ContainersReady, Ready) and in per-container `state`/`lastState`. `restartPolicy` is Pod-wide but applies per container, and it is enforced by the **kubelet on the same node** - it restarts the container inside the existing Pod, it never creates a new Pod or moves it. `Always` (default) restarts on any exit; `OnFailure` only on non-zero exit; `Never` never restarts. Repeated crashes back off exponentially - roughly 10s, 20s, 40s, capped at 5 minutes - and the container shows the waiting reason **CrashLoopBackOff**. Deployments require `Always`; Jobs use `OnFailure` or `Never`.
code
bash · 4 lineskubectl get pod api-7d9 -o jsonpath='{.status.phase}'
kubectl get pod api-7d9 -o jsonpath='{.status.containerStatuses[0].lastState.terminated.exitCode}'
kubectl logs api-7d9 --previous
kubectl describe pod api-7d9go deeper
List the five phases with a one-line meaning each and state that restartPolicy defaults to Always for long-running Pods.
Explain that the kubelet enforces restartPolicy in place with exponential backoff, and that conditions plus container states, not the phase, carry the diagnostic signal.
Turn it into a triage routine: Pending means scheduling or image or storage, CrashLoopBackOff means read the previous logs and exit code, Running-not-Ready means readiness or dependencies.
Discuss which controller semantics each restartPolicy implies, how terminal Pods and their retention interact with cluster-wide object pressure, and what the platform should surface to app teams instead of raw phases.
## Phase versus conditions `status.phase` is deliberately coarse - it answers "has this Pod started, finished, or failed?" and nothing more. - **Pending** - the API server accepted the Pod, but it is not fully running. Real causes: no node satisfies the constraints (insufficient CPU/memory, taints, node selectors, unbound PVC), the image is still being pulled, or init containers are still executing. `kubectl describe pod` events and the `PodScheduled` condition tell you which. - **Running** - the Pod is bound to a node, all containers have been created, and at least one is running, starting or restarting. **A Pod can be Running while your application is completely broken** - one container crash-looping while another runs still yields phase Running. - **Succeeded** - all containers terminated with exit code 0 and none will be restarted. - **Failed** - all containers terminated and at least one failed (non-zero exit, OOMKill, or the Pod was evicted). - **Unknown** - the node hosting the Pod stopped reporting to the control plane. Because phase is blunt, real diagnosis uses **conditions**: `PodScheduled`, `Initialized`, `ContainersReady`, `Ready` - plus per-container `state` (`Running` / `Waiting{reason}` / `Terminated{exitCode,reason}`) and `lastState`, which holds why the *previous* instance died. `restartCount` is the crash counter. ## restartPolicy `spec.restartPolicy` is set once for the whole Pod but acts on each container: - **Always** (default) - restart the container whenever it exits, regardless of exit code. - **OnFailure** - restart only on a non-zero exit code. - **Never** - do not restart; the container stays Terminated and the Pod ends up Succeeded or Failed. Three points candidates commonly miss: 1. **The kubelet does the restarting, locally.** It re-creates the container inside the same Pod sandbox. Pod name, UID, node and IP do not change; `restartCount` increments. "Restart" never means "reschedule elsewhere" - moving to another node requires a *new* Pod created by a controller. 2. **Backoff.** Repeated failures are throttled with exponential backoff - about 10s, 20s, 40s, doubling to a 5-minute cap - and reset once the container has run successfully long enough. While waiting, the container is `Waiting` with reason **CrashLoopBackOff**. That is not a root cause; it is the symptom. The cause is in `kubectl logs --previous` and in the previous termination's exit code and reason (137 = SIGKILL, often OOMKilled; 143 = SIGTERM; 1 = app error). 3. **Controllers constrain it.** Deployments, StatefulSets and DaemonSets require `Always` because they manage long-running services. Jobs and CronJobs must use `OnFailure` or `Never`, which is what lets a completed Pod stay Succeeded instead of restarting forever. ## How phase and restartPolicy interact With `Always`, a Pod essentially never reaches Succeeded - containers keep being restarted, so the Pod stays Running. With `Never`, a container that exits 1 leaves the Pod Failed permanently and nothing recovers it - exactly why a bare Pod is not a workload. With `OnFailure`, a container exiting 0 is left terminated; once all containers are terminated the Pod becomes Succeeded. ## Terminal Pods and cleanup Succeeded/Failed Pods keep existing in etcd, holding their logs, until something deletes them: Job history limits, the Pod garbage collector's terminated-Pod threshold, or you. They consume no node CPU or memory, but they do consume API objects - a cluster full of Failed Pods is an operational smell. ## Reading it in practice A Pod stuck **Pending** is a scheduling, storage or image problem: describe it and read the events. A Pod flapping in **CrashLoopBackOff** is an application problem: read the previous logs and the exit code. A Pod **Running but not Ready** is a readiness or dependency problem. Naming that triage tree is usually what the interviewer is fishing for.
- A container exits with code 137. What does that tell you?137 is 128+9, meaning the process was killed with SIGKILL. In Kubernetes the usual cause is the container exceeding its memory limit, which appears as reason OOMKilled in lastState; the other common cause is a SIGTERM that was ignored until the grace period expired.
- Why can a Pod be in phase Running while the service is down?Phase only reports that containers were created and at least one is running; it says nothing about application health. Serving readiness is expressed through the Ready and ContainersReady conditions, so a Pod can be Running, failing its readiness probe, and receiving no traffic at all.
saying these in an interview costs you the question
- Thinking restartPolicy makes the Pod restart on a different node
- Reading CrashLoopBackOff as a distinct root cause rather than a backoff state
- Believing phase Running implies the app is healthy or serving traffic
- Claiming restartPolicy can be set per container
- Assuming a Failed Pod is automatically retried or cleaned up without a controller