skip to content

In a Kubernetes Job spec, what do the `backoffLimit`, `activeDeadlineSeconds` and `ttlSecondsAfterFinished` fields control, and what happens when each of them is reached?

level: middleimportance: must knowfreq 50%

answer

  1. backoffLimit default 6; Pods (Never) or restarts (OnFailure)
  2. Backoff 10s doubling, cap 6 minutes
  3. activeDeadlineSeconds = wall clock, beats backoffLimit
  4. Reasons: BackoffLimitExceeded vs DeadlineExceeded
  5. TTL deletes Job + Pods after terminal state; 0 = immediate

basics

~20 s

backoffLimit (default 6) caps retries; exceeding it fails the Job with reason BackoffLimitExceeded. activeDeadlineSeconds caps wall-clock runtime; exceeding it kills running Pods and fails the Job with DeadlineExceeded. ttlSecondsAfterFinished deletes the finished Job and its Pods after that many seconds.

solid answer

~50 s

- **`backoffLimit`** — how many failures the Job tolerates before giving up. Default 6. With `restartPolicy: Never` it counts failed Pods; with `OnFailure` it counts container restarts. Retries back off exponentially from 10s, doubling, capped at 6 minutes. On exhaustion the Job gets condition `Failed`, reason `BackoffLimitExceeded`, and stops creating Pods. - **`activeDeadlineSeconds`** — wall-clock budget from the Job's start, counting queueing and backoff, not just execution. On expiry the controller terminates all running Pods and fails the Job with reason `DeadlineExceeded`. It **overrides** `backoffLimit`: the deadline fires even if retries remain. - **`ttlSecondsAfterFinished`** — cleanup. Once the Job reaches `Complete` or `Failed`, the TTL controller deletes it after that many seconds, cascading to its Pods. `0` deletes immediately. Without it, finished Jobs and their Pods accumulate until something else prunes them. Together: retry budget, time budget, and retention.

code

yaml · 14 lines
yaml
apiVersion: batch/v1
kind: Job
metadata:
  name: nightly-etl
spec:
  backoffLimit: 2                 # 3 attempts total
  activeDeadlineSeconds: 3600     # hard stop after 1h, retries included
  ttlSecondsAfterFinished: 43200  # keep 12h for debugging
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: etl
          image: registry.example.com/etl:7.3.0

go deeper

for a junior

Name each field's purpose in one line — retry cap, time cap, cleanup — and the default of 6 for backoffLimit.

for a middle

Explain what backoffLimit counts under each restart policy, the backoff curve, the two distinct failure reasons, and how TTL cascades to Pods.

for a senior

Reason about choosing values from observed runtime distributions, why the deadline is the reliable guard against hangs, and how TTL interacts with CronJob history limits and log retention.

for a principal

Set these as platform defaults with a rationale — retry budgets that surface real failures fast, deadlines derived from SLOs, and a retention policy that assumes centralised logging rather than Pod logs.

## backoffLimit — the retry budget `.spec.backoffLimit` bounds how many times a Job retries before declaring defeat. Default is 6. What gets counted depends on the Pod template's restart policy. With `restartPolicy: Never`, each failed **Pod** increments the counter and the controller creates a fresh Pod for the next attempt. With `restartPolicy: OnFailure`, the kubelet restarts the container inside the same Pod and each **container restart** increments the counter. Either way, once the count exceeds `backoffLimit`, the Job receives the `Failed` condition with reason `BackoffLimitExceeded`, no further Pods are created, and any running Pods are terminated. Retries are not immediate. The controller applies exponential backoff starting at 10 seconds and doubling — 10s, 20s, 40s, 80s … — capped at 6 minutes. That matters for expectations: the default budget of six retries can stretch over roughly twenty minutes of wall clock. A subtlety worth naming: a Job with a low `backoffLimit` and a container that fails fast can exhaust the budget in seconds; a Job with `backoffLimit: 0` gets exactly one attempt. For Indexed Jobs, `backoffLimitPerIndex` gives each index its own budget so one poisoned shard cannot burn the whole Job's retries, and `maxFailedIndexes` bounds how many indexes may fail permanently. ## activeDeadlineSeconds — the time budget `.spec.activeDeadlineSeconds` is a wall-clock limit measured from when the Job starts, and it counts everything: Pod scheduling delays, image pulls, execution, and the backoff gaps between retries. When it expires the controller terminates all active Pods and sets `Failed` with reason `DeadlineExceeded`. Two properties are commonly tested. First, it **takes precedence over `backoffLimit`** — the deadline is enforced regardless of remaining retries, which makes it the reliable guard against a Job that hangs rather than fails. Second, the *same field name* also exists on the Pod spec, where it limits that individual Pod's lifetime rather than the Job's total. Setting it in the Job's `spec.template.spec` is a different control from setting it in `spec`, and confusing the two is a classic error. Use it as a safety net for anything that can hang: a batch waiting on a lock, a network call with no timeout, a queue consumer that never sees an empty queue. Set it comfortably above the p99 runtime — too tight and you convert slow-but-successful runs into pages. ## ttlSecondsAfterFinished — the retention budget Jobs are deliberately not self-deleting: their Pods are the record of what happened, and `kubectl logs` on a finished Pod is the primary debugging tool. But without cleanup, a CronJob firing every five minutes leaves thousands of Job and Pod objects in etcd, slowing `kubectl` and, on large clusters, pressuring the API server. `.spec.ttlSecondsAfterFinished` hands that to the TTL-after-finished controller: once the Job reaches a terminal state (`Complete` or `Failed`), the controller deletes it after the given number of seconds, and the deletion cascades to its Pods via owner references. A value of `0` deletes immediately on completion — convenient but it destroys your logs, so it is only sensible when logs are shipped elsewhere. There is an important interaction: modifying the TTL of an already-finished Job is undefined behaviour, and clock skew between nodes can make deletions early or late by a few seconds. CronJobs have a second, independent retention mechanism: `successfulJobsHistoryLimit` (default 3) and `failedJobsHistoryLimit` (default 1) keep the last N Job objects per outcome. If both are configured, whichever prunes first wins; a common production setup uses TTL for time-based retention and the history limits for count-based retention. ## How they compose Think of them as three independent budgets: | Field | Budget | Failure reason | Overridden by | |---|---|---|---| | `backoffLimit` | attempts | `BackoffLimitExceeded` | `activeDeadlineSeconds` | | `activeDeadlineSeconds` | wall clock | `DeadlineExceeded` | nothing | | `ttlSecondsAfterFinished` | retention | n/a (post-terminal) | manual deletion | A production-grade Job usually sets all three: a small retry budget so genuine failures surface quickly rather than looping, a deadline generous enough for the slow tail but tight enough to catch hangs, and a TTL long enough to debug (hours to a day) but short enough to keep etcd tidy. ## Diagnosing `kubectl describe job` prints the terminal condition and reason, which immediately distinguishes "it failed repeatedly" (`BackoffLimitExceeded`) from "it hung" (`DeadlineExceeded`). If the Job object is already gone, TTL or history limits removed it — that is the moment teams discover they needed centralised logging rather than relying on Pod logs.

  • A Job has `backoffLimit: 10` and `activeDeadlineSeconds: 300`, and each attempt hangs for two minutes. What is the outcome?
    The Job fails after roughly five minutes with reason `DeadlineExceeded`, having made only two or three attempts. `activeDeadlineSeconds` is wall-clock and takes precedence over the retry budget, so the remaining retries are never used. Its running Pods are terminated when the deadline fires.
  • What is the difference between `activeDeadlineSeconds` on the Job spec and on the Pod template spec?
    On the Job spec it bounds the entire Job — all attempts, all backoff gaps — and failing it marks the Job `Failed` with `DeadlineExceeded`. On `spec.template.spec` it bounds each individual Pod's lifetime; a Pod that exceeds it is terminated and counts as one failure against `backoffLimit`, so the Job may still retry and eventually succeed. Use the Pod-level field to bound a single attempt and the Job-level field to bound the whole run.
  • Why is `ttlSecondsAfterFinished: 0` risky?
    The Job and its Pods are deleted the moment it reaches a terminal state, taking the logs and events with them. If the run failed, you have nothing to debug with. It is only acceptable when logs and exit metadata are shipped to an external system as they are produced; otherwise pick a TTL that covers at least one on-call cycle.

saying these in an interview costs you the question

  • Believing `backoffLimit` is a per-Pod container restart limit in all cases — with `restartPolicy: Never` it counts failed Pods.
  • Assuming remaining retries are honoured after `activeDeadlineSeconds` expires.
  • Confusing the Job-level and Pod-template-level `activeDeadlineSeconds`.
  • Thinking finished Jobs are garbage-collected by default without a TTL or CronJob history limit.
  • Expecting retries to be immediate rather than exponentially backed off up to a six-minute cap.

context