In a Kubernetes Pod spec, what does the `initContainers` list do, and what happens if one of those containers exits with a non-zero status?
answer
- initContainers: sequential, must exit 0
- Init:1/3 → Pod never Ready during init
- Never → Pod Failed; Always/OnFailure → Init:CrashLoopBackOff
- No probes, no lifecycle hooks
- Effective request = max(init) vs sum(app)
basics
~20 sInit containers run to completion one at a time, in list order, before any app container starts. A non-zero exit means the kubelet retries that init container, or fails the Pod if restartPolicy is Never. Later containers never start.
solid answer
~40 s`initContainers` is an ordered list of containers the kubelet runs **sequentially, to completion**, before it starts anything in `containers`. Each must exit 0 before the next begins; the Pod shows `Init:1/3` meanwhile and is never Ready. They share the Pod's network namespace and volumes but have their own image, command and securityContext, so they suit setup that shouldn't live in the app image: waiting for a dependency, running a schema migration, fetching config or certs into an `emptyDir`, `chown`-ing a mounted volume, or setting sysctls with extra privileges. On a non-zero exit: - `restartPolicy: Never` → the whole Pod goes to `Failed`. - `restartPolicy: Always` or `OnFailure` → the kubelet restarts **that** init container with exponential backoff (`Init:CrashLoopBackOff`). Regular init containers accept no `livenessProbe`/`readinessProbe`/`startupProbe` and no `lifecycle` hooks — "done" simply means "exited 0".
code
yaml · 26 linesapiVersion: v1
kind: Pod
metadata:
name: api
spec:
restartPolicy: Always
volumes:
- name: config
emptyDir: {}
initContainers:
- name: wait-for-db
image: busybox:1.36
command: ['sh', '-c', 'until nc -z db 5432; do sleep 2; done']
- name: fetch-config
image: curlimages/curl:8.8.0
command: ['sh', '-c', 'curl -sf http://cfg/app.yaml -o /work/app.yaml']
volumeMounts:
- name: config
mountPath: /work
containers:
- name: api
image: registry.example.com/api:1.4.0
volumeMounts:
- name: config
mountPath: /etc/app
readOnly: truego deeper
Recall the core: ordered, run-to-completion, all must succeed before app containers start; and name one real use like waiting for a dependency or fetching config.
Add the failure matrix by restartPolicy, the Init:N/M status, kubectl logs -c, and the fact that probes and lifecycle hooks are not allowed.
Discuss the cost side: init time is added to every rollout and node-failure recovery, the max-vs-sum resource rule, and why migrations belong in a Job rather than an init container.
Frame it as where platform-enforced preconditions should live — init container vs image build step vs operator vs app-level retry — and the operational tax each choice puts on rollout latency and node churn.
## Two container lists A Pod spec has two lists of containers. `containers` holds the long-running application containers; the kubelet starts them **in parallel** with no ordering guarantee between them. `initContainers` holds containers that run **before** any of those, **one at a time, in written order**, each required to run to completion (exit code 0) before the next one is started. An init container is a normal container in every other respect: it has its own image, command, environment, resource requests, volume mounts and security context. It joins the same Pod network namespace (so `localhost` and the Pod IP are already the ones the app will use) and can mount the same volumes. What it cannot do is stay running — the Pod does not progress until it exits. ## The lifecycle you observe While init containers run, the Pod's phase is `Pending` and `kubectl get pod` shows a status like `Init:0/2`, `Init:1/2`. The Pod is never `Ready` during this window, so a Service will not route traffic to it. Once the last init container exits 0, the kubelet pulls/starts the app containers and the Pod moves to `Running`. Because init containers are a distinct list, you must name them explicitly to read their output: `kubectl logs <pod> -c <init-container-name>`. `kubectl describe pod` shows a separate `Init Containers:` block with each one's state, exit code and restart count — that block is where you diagnose a stuck Pod. ## Failure semantics If an init container exits non-zero, the Pod's `spec.restartPolicy` decides what happens: - **`Never`** — the Pod is marked `Failed` immediately. No retry. For a Job whose template uses `Never`, the Job controller then creates a replacement Pod (subject to `backoffLimit`). - **`Always`** (the default for Deployments and bare Pods) or **`OnFailure`** — the kubelet restarts the *same* init container in place, with exponential backoff, and the Pod status becomes `Init:CrashLoopBackOff`. It keeps retrying forever; subsequent init containers and all app containers stay unstarted. The key point for interviews: failure is *blocking*, not skippable. There is no "continue on error" flag. If you want best-effort setup, the init container itself must swallow the error and exit 0. ## Why the sequencing matters Sequential execution is the whole value proposition. It gives you a place to express "this must be true before the app runs" that is enforced by the platform rather than by application code: - **Dependency gating** — loop until a database or config service answers, then exit; the app never sees a cold dependency on its very first request. - **Schema migration** — run the migration once per Pod start (idempotent migrations only; N replicas means N runs, so real deployments usually put migrations in a Job instead). - **Artifact staging** — clone a repo, decrypt secrets, or download a model into a shared `emptyDir` that the app container mounts read-only. The tooling (git, gpg, curl) stays out of the runtime image, which shrinks it and reduces its attack surface. - **Privilege isolation** — run one short container with `privileged: true` or extra capabilities to set sysctls or fix volume ownership, while the app container keeps a locked-down security context. ## Resources and scheduling Init containers run before app containers, so they do not run *at the same time* as them. Kubernetes therefore computes the Pod's effective request for each resource as `max( largest init container request, sum of app container requests )` — plus, since native sidecars exist, sidecar requests are added into the app-container sum because sidecars do overlap. A single fat init container can therefore raise the whole Pod's scheduling footprint even though it lives for two seconds. ## Limits worth naming - No probes and no `lifecycle` hooks on regular init containers; completion is defined by exit status alone. - They run on **every** start of the Pod, including restarts and reschedules — so they must be idempotent and fast. A 90-second init container is 90 seconds added to every rollout and every node failure recovery. - They are not a scheduling ordering primitive *between* Pods. Waiting for another Pod's Service inside an init container works, but it burns a scheduling slot while it spins; app-level retry is often the better answer. - Setting `restartPolicy: Always` on an entry in `initContainers` changes its meaning entirely — that entry becomes a **native sidecar** that keeps running for the Pod's lifetime rather than a run-to-completion step.
- Why can't a regular init container declare a readinessProbe?Readiness answers "is this long-running process able to serve?", but an init container is defined by running to completion — its success signal is the exit code, so a readiness result would have nothing to gate. The API server rejects `readinessProbe`, `livenessProbe`, `startupProbe` and `lifecycle` on regular init containers. Native sidecars (an init container with `restartPolicy: Always`) do accept them, because those really are long-running.
- How do init containers affect the Pod's resource requests and scheduling?Init containers do not overlap with app containers, so the scheduler uses the maximum of any single init container's request rather than their sum, and compares it with the sum of the app (plus sidecar) container requests — the Pod's effective request per resource is the larger of the two. A one-off init container asking for 4 GiB therefore forces every replica onto a node with 4 GiB free, even though it exits in seconds.
- Your init container runs a database migration. What is wrong with that at three replicas?The init container runs once per Pod start, so three replicas mean three concurrent migration attempts, plus another one on every restart or reschedule. Unless the migration tool takes an advisory lock and is fully idempotent, that races. The conventional fix is a separate Job (or a Helm pre-upgrade hook) that runs the migration once, with the app Pods merely waiting for the schema version.
Think of a pre-flight checklist: each item must be signed off in order before the aircraft is allowed to taxi. Fail an item and you do not skip to the next one — you stop and redo it.
saying these in an interview costs you the question
- Saying init containers run in parallel with each other or with the app containers.
- Believing a failed init container is skipped and the Pod carries on.
- Adding a readinessProbe to an init container and expecting the API server to accept it.
- Assuming init containers run only once for the whole Deployment rather than on every Pod start.
- Thinking the Pod is Ready — and receiving Service traffic — while init containers are still running.