skip to content

What is a Kubernetes Job, how does it differ from a Deployment, and which values may the `restartPolicy` field of a Job's Pod template take?

level: juniorimportance: must knowfreq 62%

answer

  1. Job = run to completion; Deployment = run forever
  2. restartPolicy: Never or OnFailure only — Always rejected
  3. OnFailure = restart container in place; Never = new Pod
  4. backoffLimit default 6, backoff 10s→6min cap
  5. At-least-once, not exactly-once

basics

~20 s

A Job runs Pods until a set number of them terminate successfully, then stops. A Deployment keeps a set of Pods running forever, restarting them whenever they exit. A Job's Pod template must use restartPolicy Never or OnFailure — Always is rejected.

solid answer

~50 s

A **Job** is a run-to-completion workload: it creates Pods and tracks them until the configured number of successful completions is reached, then stops creating Pods and records `Complete`. A **Deployment** is a run-forever workload: its ReplicaSet keeps exactly `replicas` Pods alive, and a container that exits — successfully or not — is restarted or replaced. So the difference is the terminal state. "Finished" is a valid, desirable outcome for a Job and an outage for a Deployment. A Job's Pod template must set `restartPolicy: Never` or `OnFailure`; `Always` is rejected by the API server, because it would mean the Pod could never terminate and the Job could never complete. - `OnFailure` — the kubelet restarts the failed **container in place** in the same Pod. - `Never` — the Pod is marked `Failed` and the Job controller creates a **new Pod**. Either way, failures count against `backoffLimit` (default 6). Use `Never` when you want each attempt's logs preserved as a separate Pod.

code

yaml · 14 lines
yaml
apiVersion: batch/v1
kind: Job
metadata:
  name: db-backup
spec:
  backoffLimit: 3
  ttlSecondsAfterFinished: 86400
  template:
    spec:
      restartPolicy: Never          # Always is rejected here
      containers:
        - name: backup
          image: registry.example.com/pg-backup:16
          args: ['--to=s3://backups/db']

go deeper

for a junior

State the core contrast — completes versus runs forever — and that only Never and OnFailure are valid restart policies for a Job.

for a middle

Explain what each restart policy does mechanically (container in place versus new Pod), how failures count against backoffLimit, and the retry backoff curve.

for a senior

Add the at-least-once guarantee and its consequences for idempotency, plus operational choices: Never for per-attempt logs, TTL for cleanup, label-based inspection of attempts.

for a principal

Frame Jobs as one execution substrate among several — decide when batch belongs in Kubernetes at all versus a workflow engine or queue-consumer Deployment, and define the platform defaults for retries, TTL and idempotency.

## Run-to-completion versus run-forever Kubernetes workload controllers differ mainly in what they consider a healthy steady state. A **Deployment** (through its ReplicaSet) has the steady state "exactly N Pods running". If a container exits, that state is violated and the platform repairs it — the kubelet restarts the container (`restartPolicy: Always`, the only value that makes sense there) or the ReplicaSet creates a replacement Pod. A Deployment is therefore structurally unable to finish; a batch script deployed this way runs, exits 0, and is restarted forever in a loop that looks like `CrashLoopBackOff`. A **Job** has the steady state "`.spec.completions` Pods have terminated successfully". The Job controller creates Pods, watches their terminal phase, increments `.status.succeeded` for each `Succeeded` Pod, and once the target is reached stops creating Pods and sets the `Complete` condition. The Job object and its finished Pods stick around afterwards so you can read logs and exit codes — deletion is governed by `ttlSecondsAfterFinished` or by you. ## Why restartPolicy is constrained The API server rejects `restartPolicy: Always` in a Job's Pod template. The reason is definitional: `Always` means the kubelet restarts every container whenever it exits, for any reason, so the Pod can never reach a terminal phase, so the Job can never count a completion. Only two values are legal: **`OnFailure`** — when a container exits non-zero the kubelet restarts *that container inside the same Pod*. The Pod object is reused; its `restartCount` climbs. This is cheaper (no rescheduling, no image re-pull, any downloaded state in an `emptyDir` survives) but it overwrites the container's previous logs — you need `kubectl logs --previous` to see the last attempt, and only the immediately previous one. **`Never`** — the kubelet never restarts the container. A non-zero exit puts the Pod into `Failed`, and the Job controller creates a brand-new Pod for the next attempt. Slower, but every attempt survives as its own Pod object with its own logs and events, which is why most production Jobs choose it for debuggability. In both cases the failures are counted against `.spec.backoffLimit` (default 6). With `Never` the counter tracks failed Pods; with `OnFailure` it tracks container restarts. When the limit is exceeded, the Job gets the `Failed` condition with reason `BackoffLimitExceeded` and stops retrying. Between retries the controller applies exponential backoff starting at 10 seconds and doubling, capped at 6 minutes. ## What the Job controller guarantees It guarantees *at-least-once* execution of your work, not exactly-once. A Pod can be counted as failed and retried after its container already did half the work — node crash, eviction, preemption, or a network partition between kubelet and API server all produce this. Job payloads must therefore be idempotent or externally guarded (an advisory lock, a claimed-row pattern, an idempotency key). ## The related shapes - **Single-Pod Job** — omit `completions` and `parallelism` (both default to 1). One Pod must succeed. - **Fixed completion count** — `completions: 20`, `parallelism: 4`. Twenty successes, at most four Pods at a time. - **Work queue** — leave `completions` unset with `parallelism: N`; the Pods coordinate through an external queue and the Job is complete when at least one Pod exits successfully and all Pods have terminated. - **CronJob** — a higher-level object that creates Job objects on a cron schedule. It does not run Pods itself; every CronJob execution is a Job, so all Job semantics above apply to each run. ## Reading Job status `kubectl get job` shows `COMPLETIONS` as `succeeded/desired`. `kubectl describe job` shows the `Complete` or `Failed` condition and the events for each Pod created. Pods carry the `job-name` label, so `kubectl get pods -l job-name=<job>` lists every attempt — which is exactly the workflow that `restartPolicy: Never` optimises for. ## Common mistakes Running batch work as a Deployment (endless restart loop), expecting a Job to re-run on a schedule (that is a CronJob), assuming a Job's Pods are cleaned up automatically (they are not without `ttlSecondsAfterFinished`), and assuming exactly-once execution.

  • Practically, when do you choose `Never` over `OnFailure` for a Job's Pod template?
    Choose `Never` when you want each attempt preserved as its own Pod, with its own logs, events and exit code — that is nearly always worth it in production, since diagnosing a batch failure after the fact is the common case. Choose `OnFailure` when restarts should be cheap and fast: no rescheduling, no image re-pull, and any state written to an `emptyDir` survives the retry. The cost of `OnFailure` is that logs from earlier attempts are lost beyond `--previous`.
  • What happens if you deploy a batch script as a Deployment instead of a Job?
    The script runs, exits 0, and the kubelet restarts it because a Deployment's Pod template uses `restartPolicy: Always`. You get an endless loop that eventually shows as `CrashLoopBackOff` with exit code 0, the work is repeated indefinitely, and there is no completion signal for anything downstream to wait on.
  • Does a Job guarantee your work runs exactly once?
    No — the guarantee is at-least-once. A Pod can be evicted, preempted, or lost to a node failure after doing part of its work, and the Job controller will start another attempt. Payloads must be idempotent or protected externally, for example by an advisory lock, a claimed-row pattern, or an idempotency key in the downstream system.

A Deployment is a shop that must be open every day; a Job is a delivery that must arrive once. Restarting a completed delivery forever would be absurd, which is why Always is not allowed.

saying these in an interview costs you the question

  • Setting `restartPolicy: Always` in a Job template and expecting it to be accepted.
  • Using a Deployment for batch work and being surprised by the restart loop.
  • Believing a Job re-runs on a schedule without a CronJob wrapping it.
  • Assuming finished Job Pods are garbage-collected automatically without `ttlSecondsAfterFinished`.
  • Treating Job execution as exactly-once and writing non-idempotent payloads.

context