skip to content

Kubernetes documents that a CronJob may create a Job twice for one schedule point, or not at all. Given that guarantee, how would you design a nightly billing batch that must charge each customer exactly once?

level: principalimportance: nice to knowfreq 28%

answer

  1. At-least-once + possible zero; retries add duplicates
  2. Deterministic run key from the PERIOD, never now()
  3. Unique-constraint claim or fenced advisory lock in the DB
  4. Per-item idempotency key on external charges
  5. Missed run is silent → heartbeat + reconciliation pass

basics

~20 s

Treat the schedule as at-least-once and make exactly-once a property of the payload: derive a deterministic run key from the period, take a database lock or unique-constraint claim on it, make every charge idempotent with a per-invoice key, and alert on a heartbeat rather than on Job failures.

solid answer

~60 s

Kubernetes gives **at-least-once, sometimes zero** delivery, and Job retries add duplicate *executions* on top. Exactly-once must therefore live in the payload and the data model, not in the scheduler. Design: 1. **Deterministic run key** from the period — `billing-2026-08-12` — derived from the *logical* date, not `now()`, so a retry or a late catch-up run computes the same key. 2. **Mutual exclusion** — insert that key into a `runs` table with a unique constraint, or take a database advisory lock. A second concurrent copy exits 0 immediately. `concurrencyPolicy: Forbid` is a hint, not a guarantee. 3. **Idempotent per-item effects** — one row per (invoice, period) with a unique key, and an idempotency key on the payment provider call so a retry after a lost response does not double-charge. 4. **Resumability** — checkpoint per item so a killed Pod restarts where it stopped. 5. **Detect the zero case** — alert on the age of the last *successful business outcome*, not on Job status; a run that never started produces no failure signal. Also set `activeDeadlineSeconds`, a small `backoffLimit`, and an explicit `timeZone`.

code

sql · 21 lines
sql
-- one owner per logical run
CREATE TABLE batch_runs (
  run_key     text PRIMARY KEY,      -- 'billing-2026-08-12'
  started_at  timestamptz NOT NULL,
  finished_at timestamptz
);

-- one charge per customer per period, enforced by the database
CREATE TABLE charges (
  customer_id bigint NOT NULL,
  period      date   NOT NULL,
  amount_cents bigint NOT NULL,
  provider_idempotency_key uuid NOT NULL,
  PRIMARY KEY (customer_id, period)
);

-- second copy of the run loses this and exits 0
INSERT INTO batch_runs(run_key, started_at)
VALUES ('billing-2026-08-12', now())
ON CONFLICT (run_key) DO NOTHING
RETURNING run_key;

go deeper

for a junior

State the key fact — the schedule is at-least-once and may be zero — and that the job itself must be safe to run twice.

for a middle

Add the mechanics: a deterministic run key, a unique constraint or lock to claim the run, and idempotency keys on external side effects.

for a senior

Cover resumability and checkpointing, fencing against partitioned zombie Pods, activeDeadlineSeconds so a hang cannot suppress later runs, and heartbeat-based alerting for the silent-miss case.

for a principal

Argue where the guarantee should live at all — payload and data model versus a workflow engine or durable queue — quantify the blast radius of a duplicate charge, and define the reconciliation strategy that makes scheduling reliability a non-issue.

## Start by naming the guarantee Three independent sources of non-determinism stack up: 1. **The CronJob controller** reconciles rather than firing timers, so a controller restart, a clock adjustment, or a control-plane outage can cause a schedule point to be **missed** entirely, or (rarely) a duplicate Job to be created. The documentation states the guarantee as "about once per schedule" — not exactly once. 2. **The Job controller** retries failed Pods up to `backoffLimit`. A Pod that did half the work and then lost its node is retried from scratch. That is at-least-once *execution*. 3. **The kubelet and node** can evict, preempt or kill a Pod mid-write, and a network partition can hide a completed side effect from the caller. So the platform gives you at-least-once with a possibility of zero. Any "exactly once" property must be constructed above it. Saying this explicitly is the first thing an interviewer is listening for. ## Make the run addressable Everything hinges on a **deterministic key for the logical run**. Compute it from the period being billed — `billing-2026-08-12` — not from wall-clock `now()` inside the container. If the key is derived from the current time, a retry twenty minutes later and a catch-up run the next morning are different runs, and every duplicate-suppression mechanism downstream fails. Pass the logical date in explicitly, or derive it from the schedule with a documented rule for late starts. ## Enforce single execution where the truth lives `concurrencyPolicy: Forbid` prevents a second *Job* while one is active. It does not prevent a duplicate Job created across a controller restart, a manually triggered run, a run in a second cluster during a failover, or a zombie Pod whose node was partitioned but is still writing. Mutual exclusion must be enforced by the system of record: - **Unique-constraint claim** — `INSERT INTO batch_runs(run_key, started_at) VALUES (...)`; a duplicate key violation means "someone else owns this run", and the second copy exits 0. Durable and visible. - **Advisory lock** with a lease and heartbeat — good for long runs, but requires careful expiry so a partitioned holder cannot resurrect and keep writing (fence with a monotonically increasing token). Whichever you pick, the fencing token or claim row must be checked on **every write**, not only at startup, if the run is long enough for a partition to occur mid-flight. ## Make each item's effect idempotent Single execution of the *run* is not enough, because a run can die halfway and be retried. Push idempotency down to the item: - One row per `(customer, period)` with a unique constraint, so re-processing is a no-op or an upsert. - An **idempotency key** on every external call — most payment providers accept one and deduplicate server-side, which is the only way to be safe when the response to a charge is lost in the network. - Checkpoints: record each item as done in the same transaction as its effect, so a resumed run skips it. If the effect is external, use the outbox pattern — write intent transactionally, then have a separate process deliver it with retries and an idempotency key. At this point the number of times the Job runs stops mattering, which is exactly the property you want: correctness no longer depends on a scheduling guarantee you do not have. ## Handle the zero case, which is the harder half Duplicates are loud; a missed run is silent. Nothing fails, no Job exists, no alert fires. Defences: - **Heartbeat from the payload** — the job writes "billing for period X completed" to a monitoring system; alert when that record is older than a period plus a margin. This survives the CronJob being suspended, deleted, misconfigured or stuck. - **Bounded catch-up** — a sane `startingDeadlineSeconds` (a few intervals) so a brief outage recovers automatically, without an unbounded backlog and without tripping the controller's 100-missed-schedules guard, which sticks and stops the schedule permanently. - **A reconciliation pass** — a second, later job that finds unbilled periods and processes them. With idempotent items this is safe to run always, and it converts "scheduler reliability" into "eventual consistency", which is a much cheaper property to guarantee. ## Set the Job's own budgets - `activeDeadlineSeconds` — a hang would otherwise hold the lock and, with `Forbid`, suppress every subsequent run indefinitely. - A small `backoffLimit` — with idempotent items, retries are safe but should not mask a systematic failure; fail loudly after two or three. - `restartPolicy: Never` so each attempt keeps its own Pod and logs, plus `ttlSecondsAfterFinished` and centralised log shipping so evidence outlives cleanup. - Explicit `timeZone`, and a documented decision about DST — a schedule near a transition can be skipped or repeated, which for a *daily* billing run is a correctness question, not a cosmetic one. ## Know when to leave CronJob behind CronJob's honest scope is "periodically kick off a task whose correctness does not depend on the kick". Once you need dependency graphs between steps, guaranteed backlog execution, per-run audit trails, human approval gates or exactly-once orchestration, the right answer is a workflow engine (Argo Workflows, Temporal, Airflow) or an event-driven consumer with a durable queue — while frequently still executing the steps as Kubernetes Pods. A principal-level answer states that boundary rather than bending CronJob toward guarantees it does not offer.

  • Isn't `concurrencyPolicy: Forbid` enough to guarantee a single run?
    No. It only stops the controller creating a second Job while one is Active in that cluster. It does not cover a duplicate Job created across a controller restart, a manual trigger, a run started in a second cluster during failover, or a partitioned Pod that is still writing while Kubernetes considers it gone. Treat it as a useful hint that reduces noise, and enforce real mutual exclusion in the system of record with a claim row or fenced lock.
  • How do you detect that a nightly run never happened at all?
    Not from Kubernetes — no Job means no failure event, and the CronJob object still looks healthy. Have the payload emit a heartbeat recording the business outcome ("billing for 2026-08-12 completed, 41,203 invoices") and alert when that record is older than one period plus a margin. Back it with a reconciliation job that scans for unbilled periods, which turns a missed run into a delay rather than a data loss.
  • At what point would you move this off CronJob entirely?
    When correctness starts to depend on orchestration rather than on the payload: multi-step dependencies, guaranteed execution of a backlog, per-run audit trails, retries with different policies per step, or human approval gates. A workflow engine such as Argo Workflows or Temporal, or an event-driven consumer on a durable queue, provides those primitives directly — usually still executing each step as a Kubernetes Pod, so you keep the packaging and lose only the scheduling semantics you had outgrown.

Treat the schedule like a doorbell, not a contract: it may ring twice or not at all, so the ledger inside the house — not the bell — decides whether the customer was charged.

saying these in an interview costs you the question

  • Treating `concurrencyPolicy: Forbid` as an exactly-once guarantee.
  • Deriving the run key from `now()` inside the container instead of from the logical period.
  • Assuming Job retries are safe without per-item idempotency or checkpoints.
  • Alerting only on Job failure, which never fires when the run was never created.
  • Claiming Kubernetes CronJob offers exactly-once scheduling.

context