skip to content

In a Kubernetes Job spec, how do the `completions` and `parallelism` fields interact, and what changes when `completionMode` is set to `Indexed`?

level: middleimportance: must knowfreq 55%

answer

  1. completions = successes needed; parallelism = concurrency cap
  2. Both default 1; controller throttles near the end
  3. Work queue = completions unset, one success + all terminated
  4. Indexed = fixed index 0..N-1, retried with the SAME index
  5. backoffLimitPerIndex / maxFailedIndexes need Indexed

basics

~20 s

completions is how many Pods must succeed; parallelism is how many may run at once. The Job keeps up to parallelism Pods running until completions successes accumulate. Indexed mode gives each Pod a fixed index 0..completions-1, so each does a distinct shard.

solid answer

~50 s

`completions` = total successful Pod terminations required. `parallelism` = maximum Pods running concurrently. Both default to 1. The controller keeps up to `parallelism` Pods in flight, throttling near the end so it does not overshoot the remaining work. Three standard shapes: - **Single** — both unset: one Pod must succeed. - **Fixed count** — `completions: 50, parallelism: 5`: fifty successes, five at a time. - **Work queue** — `completions` unset, `parallelism: N`: Pods pull from an external queue; the Job completes when at least one Pod succeeds and all have terminated. With the default `completionMode: NonIndexed`, Pods are interchangeable — nothing tells a Pod *which* item it owns, so you need an external queue. With **`completionMode: Indexed`**, each Pod gets a unique index `0..completions-1`, exposed as the `JOB_COMPLETION_INDEX` annotation and (via the downward API) an env var, and baked into the Pod hostname. Index *i* counts as complete only when a Pod with that index succeeds, so a static shard split needs no queue at all.

code

yaml · 24 lines
yaml
apiVersion: batch/v1
kind: Job
metadata:
  name: reindex
spec:
  completions: 12
  parallelism: 4
  completionMode: Indexed
  backoffLimitPerIndex: 2
  maxFailedIndexes: 1
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: worker
          image: registry.example.com/reindex:3.1
          env:
            - name: JOB_COMPLETION_INDEX
              valueFrom:
                fieldRef:
                  fieldPath: metadata.annotations['batch.kubernetes.io/job-completion-index']
            - name: SHARDS
              value: '12'
          command: ['/app/reindex', '--shard=$(JOB_COMPLETION_INDEX)', '--of=$(SHARDS)']

go deeper

for a junior

Recall the two definitions — successes required versus concurrent Pods — and that both default to 1.

for a middle

Cover all three shapes including the work-queue completion rule, and explain what Indexed mode gives each Pod and how retries keep the index.

for a senior

Discuss shard-split versus queue on straggler behaviour and load balance, per-index retry budgets, and live tuning of parallelism to control batch throughput.

for a principal

Position Job parallelism against workflow engines and stream consumers, and set platform defaults for partitioning strategy, retry budgets and how partial completion is surfaced to owners.

## The two knobs `.spec.completions` answers "how much work is there?" — the number of Pods that must terminate successfully before the Job is `Complete`. `.spec.parallelism` answers "how fast may we go?" — the ceiling on Pods running at the same time. Both default to 1. The Job controller reconciles toward those two numbers: it creates Pods until either `parallelism` are active or the remaining work is smaller than `parallelism`, and it stops once `.status.succeeded` reaches `completions`. Near the end it deliberately throttles — with 50 completions, 5 parallelism and 48 succeeded, it will run at most 2 Pods, not 5, because extra successes are wasted work. `parallelism: 0` is legal and pauses the Job; you can also patch either field on a live Job to speed it up or slow it down. ## The three canonical shapes **Single Pod.** Both fields unset (=1). One success finishes the Job. This is the shape a CronJob usually creates. **Fixed completion count.** `completions: N`, `parallelism: P`. The Job needs N successes and runs at most P Pods concurrently. Suits a known, fixed amount of work divided into N equal chunks. **Work queue.** `completions` left unset with `parallelism: P`. Now the *Pods* decide when the work is done, by consuming from an external queue (Redis, SQS, a database table). The completion rule changes: the Job is `Complete` when **at least one Pod exits successfully and all Pods have terminated**. The convention is that a worker that finds the queue empty exits 0, and once one worker has signalled "queue drained" the controller stops creating replacements and waits for the rest to finish. ## What Indexed mode adds In the default `completionMode: NonIndexed`, Pods are anonymous and interchangeable — the controller counts successes and nothing more. If your work splits into fifty distinct shards, nothing in the Job tells Pod #3 that it owns shard 3, so you must bring your own coordination (a queue, a leasing table). Setting `.spec.completionMode: Indexed` changes the accounting from a count to a *set*. Each Pod is assigned a unique index in `0..completions-1`, and index *i* is satisfied only when a Pod carrying index *i* succeeds. If index 7 fails, the controller creates a new Pod **with index 7**, not just "another Pod". The index is delivered three ways: - annotation `batch.kubernetes.io/job-completion-index` on the Pod, - as an environment variable via the downward API referencing that annotation (conventionally `JOB_COMPLETION_INDEX`), - in the Pod's hostname, as `<job-name>-<index>` — which, combined with a headless Service, gives stable per-index DNS, the mechanism behind Job-based distributed training and MPI-style workloads. This lets you do a purely static shard split: worker *i* processes rows where `id % completions == i`, or file *i* from a sorted list. No queue, no coordination service. The trade-off is that the split is fixed at submission time and stragglers cannot be rebalanced — an unevenly sized shard leaves the whole Job waiting on one Pod. ## Failure handling with indexes Indexed Jobs unlock per-index retry controls: - `backoffLimitPerIndex` caps retries for each index individually, instead of one global `backoffLimit` where one poison shard burns the whole Job's retry budget. - `maxFailedIndexes` lets the Job tolerate a bounded number of permanently failed indexes and still report the rest as done; failed indexes are listed in `.status.failedIndexes`. - `.status.completedIndexes` reports what succeeded, in compressed range form such as `0-4,7,9-11`, which is directly usable for a targeted re-run. Without indexes, none of this is expressible: a NonIndexed Job cannot tell you *which* unit of work failed, only how many did. ## Choosing between the shapes Use **Indexed** when the work partitions statically and evenly, when you want per-shard retry accounting, or when workers need stable identities and DNS. Use a **work queue** when items vary wildly in cost (dynamic pull naturally balances load), when the item count is unknown at submission, or when items arrive while the Job runs. Use a **single Pod** for anything that does not need to be parallel — most scheduled maintenance tasks. ## Operational notes `kubectl get job` prints `COMPLETIONS` as `succeeded/desired`. Raising `parallelism` on a running Job takes effect immediately, which is a useful lever when a batch is behind. Both fields are independent of `backoffLimit`, which counts failures, not successes — a Job can burn its retry budget long before it approaches its completion target.

  • What does it mean when `completions` is left unset but `parallelism` is 6?
    That is the work-queue shape. Six Pods run concurrently and coordinate through an external queue; the Job has no idea how many items exist. Completion is redefined: the Job is Complete when at least one Pod has exited successfully and all Pods have terminated. The convention is that a worker exits 0 when it finds the queue empty, which signals the controller to stop creating replacements.
  • When would you prefer a work queue over an Indexed Job?
    When item costs are uneven or the item count is unknown at submission time. A static index split fixes the partition up front, so one heavy shard leaves the whole Job waiting on a single straggler while other Pods idle. Dynamic pull from a queue self-balances and lets items arrive after the Job starts. The cost is the queue itself: extra infrastructure, visibility timeouts, and at-least-once redelivery semantics to handle.
  • Why can't a NonIndexed Job use `backoffLimitPerIndex`?
    Because there are no indexes to attribute failures to. A NonIndexed Job counts anonymous successes and failures, so the platform cannot tell whether ten failures were one poisoned work item retried ten times or ten different items failing once. Indexed mode makes failures addressable, which is what enables per-index retry budgets, `maxFailedIndexes`, and the `.status.failedIndexes` report.

saying these in an interview costs you the question

  • Thinking `parallelism` sets the total number of Pods rather than the concurrency ceiling.
  • Believing an ordinary (NonIndexed) Job tells each Pod which shard of the work it owns.
  • Expecting a failed index in an Indexed Job to be retried as "just another Pod" rather than with the same index.
  • Assuming `completions` and `backoffLimit` are related — one counts successes, the other failures.
  • Claiming a work-queue Job completes only when the queue is empty; the rule is one Pod succeeding plus all Pods terminated.

context