skip to content

A Kubernetes Job's Pod runs a batch process plus a log-shipping helper container. The batch process exits 0, yet the Pod stays in `Running` and the Job never reports completion. What is happening, and what is the correct fix?

level: seniorimportance: should knowfreq 40%

answer

  1. Pod Succeeded only when ALL regular containers exit
  2. Log shipper never exits → Job hangs forever
  3. Fix: initContainers + restartPolicy: Always
  4. Sidecar terminated after app → tail flushed
  5. activeDeadlineSeconds turns a hang into a false failure

basics

~20 s

A Job's Pod is only complete when its regular containers have all exited, and the log shipper never exits. Fix it by declaring the helper as a native sidecar — an initContainers entry with restartPolicy: Always — so completion is judged on the batch container alone.

solid answer

~50 s

Job completion is evaluated on the Pod, and a Pod reaches `Succeeded` only when all of its regular containers have terminated. A log shipper is designed to run forever, so it pins the Pod in `Running` indefinitely. The Job's `completions` counter never advances; if `activeDeadlineSeconds` is set the Job eventually fails with `DeadlineExceeded`, and if it isn't, the Pod occupies its node resources forever. **Fix:** move the helper out of `containers` and into `initContainers` with `restartPolicy: Always`, making it a native sidecar. The kubelet then starts it before the batch container, judges completion on the regular containers only, and terminates the sidecar itself once they exit — after them, so the final log lines are still shipped. Before native sidecars, teams faked this with an `emptyDir` flag file the shipper polled, a `preStop` hook, or the main container calling the helper's shutdown endpoint. All of them coupled the app image to its helper; delete them once the cluster is on 1.29+.

code

bash · 10 lines
bash
kubectl get pods -l job-name=report
# NAME          READY   STATUS    RESTARTS   AGE
# report-x2k9d  1/2     Running   0          46m

kubectl get pod report-x2k9d -o jsonpath='{range .status.containerStatuses[*]}{.name}{"\t"}{.state}{"\n"}{end}'
# report	{"terminated":{"exitCode":0,"reason":"Completed"}}
# shipper	{"running":{"startedAt":"2026-08-12T08:02:11Z"}}

kubectl get job report -o jsonpath='{.status}'
# {"active":1,"startTime":"..."}   <- succeeded never increments

go deeper

for a junior

Recall that a Pod finishes only when all its regular containers exit, so a never-ending helper blocks the Job.

for a middle

Give the fix precisely — initContainers entry with restartPolicy: Always — and explain that completion is then judged on the regular containers.

for a senior

Diagnose from container statuses, cover the flush window (SIGTERM handling plus terminationGracePeriodSeconds), reject activeDeadlineSeconds as a fix, and name the legacy workarounds you would now delete.

for a principal

Decide sidecar-per-Job versus node-level collection on cost and blast radius, set the platform's minimum Kubernetes version for relying on native sidecars, and standardise grace-period and TTL defaults across all batch workloads.

## The rule that bites The Job controller does not look at containers. It looks at Pods, and counts a Pod toward `.status.succeeded` when the Pod reaches phase `Succeeded`. The kubelet only moves a Pod to `Succeeded` when **every regular container has terminated** and none failed in a way that the restart policy would retry. A helper such as a log shipper, a metrics pusher, a mesh proxy or a cloud-SQL auth proxy is written to run until killed. Put it in `containers` alongside a batch process and you have built a Pod that can never complete. The batch container sits in `Terminated (Completed)`, the helper sits in `Running`, and the Pod's phase stays `Running` forever. ## What that costs you - The Job's `completions` count never advances, so anything gated on Job success — a CI stage, a pipeline step, a downstream CronJob, an alert on "job did not finish" — hangs. - The Pod keeps holding its CPU/memory requests on a node. Run this hourly and you slowly starve the cluster. - `ttlSecondsAfterFinished` never fires, because the Job never finishes, so nothing is cleaned up. - If `activeDeadlineSeconds` is set, the Job is eventually killed with reason `DeadlineExceeded` and reported as **failed** — which is worse than hanging, because now a successful batch run pages someone. ## Diagnosis `kubectl get pod` shows `1/2` ready with status `Running` long after the work finished. `kubectl get pod -o jsonpath='{.status.containerStatuses[*].state}'` shows one `terminated` with `exitCode: 0` and one `running`. That pair — one completed, one running, phase Running — is the fingerprint. ## The fix Declare the helper as a **native sidecar**: move it into `initContainers` and set `restartPolicy: Always` on that entry. This changes three things at once: 1. **Startup order** — the kubelet starts the sidecar before the batch container, so no log lines from early startup are lost. 2. **Completion accounting** — Pod completion is evaluated on regular containers only; the running sidecar is ignored. 3. **Shutdown order** — when the batch container exits, the kubelet terminates the sidecar *after* it, giving the shipper a window (bounded by `terminationGracePeriodSeconds`) to flush its buffer. The sidecar also restarts independently if it crashes mid-run — even in a Job Pod whose `restartPolicy` is `Never`, where nothing else can restart in place. That matters: a crashed shipper no longer means a failed batch run. ## Flushing, properly Ordering alone does not guarantee the tail of the log is shipped. The sidecar must handle SIGTERM by flushing and exiting, and the Pod's `terminationGracePeriodSeconds` must exceed its worst-case flush time. If the shipper ignores SIGTERM it is SIGKILLed at the end of the grace period and you lose the buffer — the same data loss, moved later. Where the buffer can be large, a `preStop` hook that triggers an explicit flush endpoint is a reasonable belt-and-braces addition. ## The legacy workarounds (recognise, then delete) Before native sidecars, every team invented one of these: - **Flag file in a shared `emptyDir`** — the batch container `touch`es `/shared/done` on exit; the helper polls for it and exits. Requires modifying both images and leaks on abnormal termination. - **Wrapper entrypoint** — the batch container's script runs the work, then curls the helper's shutdown endpoint (Istio's `/quitquitquit` being the canonical example). Puts helper-specific knowledge into the application image. - **`preStop` hooks or a shared process supervisor** — helper run as a child of the main process, which defeats the point of separate containers. - **`activeDeadlineSeconds` as a timeout** — the worst one, because it converts every successful run into a failed Job. All are removable on Kubernetes 1.29+ (native sidecars beta and on by default; stable in 1.33). Being able to name both the workaround and the version at which it becomes dead code is what a senior answer sounds like. ## A related trap Sidecar resource requests are **summed** with the app containers when the scheduler computes the Pod's effective request, not folded in as a max like regular init containers. Attaching a shipper to thousands of short-lived Job Pods is a real capacity decision — which is often the argument for shipping logs with a node-level DaemonSet instead of a per-Pod sidecar, and reserving the sidecar for cases where the Job writes to a private `emptyDir` that a node agent cannot see.

  • Why is setting activeDeadlineSeconds a bad answer to this problem?
    It does stop the Pod from running forever, but by killing the Job and marking it `Failed` with reason `DeadlineExceeded`. A batch run that completed its work successfully is then reported as a failure, which corrupts your success metrics and alerting. It also caps legitimate long runs, so you end up tuning a timeout instead of fixing the completion semantics.
  • Even with a native sidecar, the last few log lines are still missing. What would you check?
    Ordering guarantees only that the sidecar is terminated after the batch container, not that it flushes in time. Check that the shipper handles SIGTERM by flushing and exiting rather than dying on SIGKILL, and that `terminationGracePeriodSeconds` exceeds its worst-case flush duration. A `preStop` hook that calls an explicit flush endpoint is a reasonable extra guard for large buffers.
  • When would you avoid a log-shipper sidecar on Job Pods altogether?
    When the workload is high-volume and short-lived. Every Pod pays the sidecar's image pull, startup latency and resource requests, and sidecar requests sum with the app container's for scheduling. A node-level DaemonSet collecting container stdout is usually cheaper and simpler; the sidecar earns its place only when the Job writes to a file in a private volume that a node agent cannot read.

saying these in an interview costs you the question

  • Blaming the batch container's exit code when its status clearly shows exit 0.
  • Using activeDeadlineSeconds as the fix, which reports successful runs as failed Jobs.
  • Believing the Job controller inspects individual containers rather than Pod phase.
  • Adding an emptyDir flag-file handshake on a cluster that already supports native sidecars.
  • Assuming a native sidecar automatically flushes buffers, without checking SIGTERM handling and the grace period.

context