skip to content

Kubernetes allows `restartPolicy: Always` on an entry inside a Pod's `initContainers` list. What does that setting change about how the container behaves, and which problems of the older "just add a second container to the Pod" sidecar pattern does it solve?

level: middleimportance: must knowfreq 55%

answer

  1. initContainers entry + restartPolicy: Always = sidecar
  2. Init ordering + app-container lifetime
  3. Waits for started / startupProbe, not just created
  4. Shutdown: app first, sidecars reverse order
  5. Sidecar requests SUM with app; Jobs still complete

basics

~20 s

It makes that entry a native sidecar: it starts in init order but keeps running for the Pod's whole life instead of running to completion. Later init containers and the app containers wait for it, it does not block Job completion, and it shuts down after the app containers.

solid answer

~50 s

An `initContainers` entry with `restartPolicy: Always` is a **native sidecar**. It gets the ordering of an init container and the lifetime of an app container. Concretely: - **Startup ordering** — the kubelet starts it in list order and waits until it has started (and its `startupProbe`, if present, succeeds) before moving to the next init container or the app containers. Classic sidecars in `containers` start in parallel with the app, so the app can beat its proxy. - **Lifetime** — it keeps running instead of having to exit; it restarts independently if it crashes. - **Termination** — on Pod shutdown the app containers are terminated first, then sidecars in reverse order, so log shippers and proxies outlive the workload they serve. - **Jobs** — the Pod completes when the regular containers finish; sidecars are then killed. Previously a log shipper kept a Job Pod `Running` forever. - They also accept probes and `lifecycle` hooks, which regular init containers cannot have.

code

yaml · 29 lines
yaml
apiVersion: batch/v1
kind: Job
metadata:
  name: nightly-export
spec:
  backoffLimit: 2
  template:
    spec:
      restartPolicy: Never
      volumes:
        - name: logs
          emptyDir: {}
      initContainers:
        - name: log-shipper
          image: fluent/fluent-bit:3.0
          restartPolicy: Always          # <- makes this a sidecar
          startupProbe:
            httpGet: { path: /api/v1/health, port: 2020 }
            periodSeconds: 2
            failureThreshold: 30
          volumeMounts:
            - name: logs
              mountPath: /var/log/app
      containers:
        - name: export
          image: registry.example.com/export:2.1.0
          volumeMounts:
            - name: logs
              mountPath: /var/log/app

go deeper

for a junior

Know that it is declared under initContainers with restartPolicy: Always, starts before the app, and keeps running for the Pod's life.

for a middle

Explain all four wins — ordered start, ordered shutdown, independent restart, Job completion — and that sidecars can carry probes and lifecycle hooks.

for a senior

Add the operational detail: startupProbe is what makes ordering real, sidecar readiness gates Pod readiness, a hung sidecar stalls init or shutdown, and requests sum rather than max.

for a principal

Weigh sidecar-per-Pod against node-level agents (DaemonSet) and out-of-Pod services on capacity, blast radius and upgrade cadence; note the 1.29/1.33 availability floor when standardizing a platform on it.

## The problem the feature fixes Before native sidecars, a "sidecar" was just convention: you added a second entry to `containers` — a service-mesh proxy, a log shipper, a secret-refresher, a metrics adapter — and hoped it behaved. Kubernetes had no notion that it was subordinate to the main container, which produced four recurring bugs: 1. **Startup race.** All entries in `containers` are started concurrently, with no ordering guarantee. An app that calls out through a mesh proxy would frequently fail its first requests because the proxy wasn't listening yet. 2. **Shutdown race.** On termination all containers get SIGTERM at once. The proxy could die while the app was still draining in-flight requests, and the last log lines were lost when the shipper exited first. 3. **Jobs never finish.** A Job Pod is complete only when *all* its containers have exited. A log shipper that runs forever means the Pod stays `Running` and the Job never reports success. Teams worked around this with shared `emptyDir` flag files, `preStop` hooks, or calling a proxy's shutdown endpoint from the main container. 4. **Restart coupling.** With `restartPolicy: Always`, a crashing helper and a crashing app looked identical to the Pod; with `OnFailure`/`Never` templates you could not have a helper restart independently at all. ## What the feature actually is The `restartPolicy` field was made settable on individual entries in `initContainers`, and the only meaningful value there is `Always`. An entry that sets it becomes a **sidecar container**: still declared among the init containers (so it inherits their ordering), but no longer required to exit. The kubelet's algorithm becomes: walk `initContainers` in order; for a normal entry, wait until it exits 0; for an entry with `restartPolicy: Always`, wait until the container has **started** — meaning its `postStart` hook returned and, if a `startupProbe` is declared, that probe has succeeded — then move on. Once the list is exhausted, start all of `containers` in parallel. That single change gives: - **Deterministic startup order.** A proxy declared before an app container is guaranteed to be up first. You can also order sidecars relative to each other (secret fetcher, then proxy, then app). - **Full container features.** Sidecars accept `livenessProbe`, `readinessProbe`, `startupProbe` and `lifecycle` hooks — regular init containers accept none of these. A sidecar's readiness feeds the Pod's readiness, so a broken proxy takes the Pod out of Service endpoints. - **Independent restart.** A sidecar restarts on its own without the Pod restarting, even inside a Job whose Pod template says `restartPolicy: Never`. That is otherwise impossible. - **Reverse-order shutdown.** During termination the kubelet sends SIGTERM to the regular containers first, waits for them (bounded by `terminationGracePeriodSeconds`), then terminates sidecars in reverse declaration order. The proxy and log shipper outlive the app they serve. - **Jobs complete.** Job/Pod completion is evaluated on the regular containers only. When they exit, the kubelet kills the sidecars and the Pod reaches `Succeeded`. All the `emptyDir` flag-file and `/quitquitquit` hacks disappear. ## Resources and scheduling Because sidecars run alongside the app rather than before it, their requests are **added to** the app containers' sum when computing the Pod's effective request — unlike regular init container requests, which are folded in as a max. Adding a 200m/256Mi proxy to every Pod is a real, cluster-wide capacity change; it also participates in the QoS class calculation. ## Practical caveats - **"Started" is not "ready".** Without a `startupProbe`, the kubelet only waits for the container process to be up, not for it to be serving. If ordering actually matters, give the sidecar a `startupProbe` that hits its health endpoint. - **A hanging sidecar blocks the Pod.** During init, a sidecar that never becomes started stalls everything behind it. During termination, a sidecar that ignores SIGTERM is killed when the grace period expires. - **Sidecar readiness affects Pod readiness**, which is usually desirable but can be surprising if the sidecar's health check is flakier than the app's. - **It is still one Pod.** Sidecars share the network namespace, node and lifecycle. Anything that should scale or fail independently belongs in its own Deployment. ## Version awareness The capability arrived as alpha in Kubernetes 1.28, went beta and on-by-default in 1.29, and reached stable in 1.33. On clusters older than 1.29 you cannot rely on it and must keep the classic workarounds; mentioning that boundary is what separates a rehearsed answer from an operational one.

  • Why does the kubelet wait for a sidecar's startupProbe rather than just for the process to launch?
    Because "process launched" says nothing about "able to serve". A mesh proxy takes time to fetch its configuration and bind its listener; if the kubelet moved on the moment the process existed, the app container could still beat it. Declaring a `startupProbe` on the sidecar turns the gate into a real readiness gate for ordering purposes — that is the only way to make ordering meaningful.
  • How do a sidecar's resource requests differ from a regular init container's in scheduling?
    Regular init containers do not overlap with app containers, so the scheduler takes the maximum across them and compares it against the sum of the app containers. A sidecar does overlap, so its requests are added into that app-container sum. Injecting a sidecar into every Pod in the cluster therefore raises total reserved capacity linearly with Pod count.
  • What did teams do about the never-completing Job problem before native sidecars existed?
    They faked the missing lifecycle. Common hacks were a shared `emptyDir` flag file that the main container touched on exit and the sidecar polled before exiting, a `preStop` hook or a wrapper script in the main container calling the sidecar's shutdown endpoint (for example Istio's `/quitquitquit`), or having the main container run the sidecar as a child process. All of them coupled the application image to its sidecar, which is exactly what native sidecars remove.

saying these in an interview costs you the question

  • Claiming a plain second entry in `containers` gives ordered startup — it does not; those start in parallel.
  • Setting `restartPolicy: OnFailure` or `Never` on an init container and expecting sidecar semantics; only `Always` is meaningful there.
  • Thinking the kubelet waits for the sidecar to be *ready* by default, when without a startupProbe it only waits for it to be started.
  • Assuming sidecar resource requests are folded in as a max like other init containers, rather than summed with the app containers.
  • Saying the Pod of a Job completes when *all* containers exit — with native sidecars, only the regular containers count.

context