What makes a scheduled batch task idempotent, and why do orchestrators require it?
answer
- think about what a retry leaves behind
- same window in, same end state out
- scope from the interval, replace don't append
- side effects have no partition to overwrite
basics
~20 sA task is idempotent when running it again for the same input window leaves the target in the same state as one successful run. Orchestrators retry, replay and re-run tasks constantly, so anything else duplicates data.
solid answer
~50 sIdempotence means the observable end state is the same whether the task ran once or five times for the same input window. Three things make it hold: the task derives **which slice it reads and writes** from its run parameters (the interval), not from the wall clock; the write **replaces** that slice — a partition overwrite or a merge on a natural key — instead of appending to it; and any external side effect (email, webhook, file drop) is either suppressed on replay or keyed so the receiver ignores duplicates. It matters because every orchestrator re-runs tasks: automatic retries after transient failures, a worker that died after writing but before reporting success, replaying missed intervals, an operator clearing a run, or a deliberate backfill after a logic change. If the task is idempotent all of those are routine; if not, each one is a silent data-corruption incident.
code
python · 23 lines# not idempotent: scope from the clock, write appends
def load(conn):
conn.execute("""
INSERT INTO sales_daily
SELECT order_id, amount, order_ts::date AS dt
FROM raw_orders
WHERE order_ts >= now() - interval '1 day'
""")
# idempotent: scope from the run interval, write replaces
def load(conn, interval_start, interval_end):
with conn.begin():
conn.execute(
"DELETE FROM sales_daily WHERE dt >= %s AND dt < %s",
(interval_start, interval_end),
)
conn.execute(
"""INSERT INTO sales_daily
SELECT order_id, amount, order_ts::date
FROM raw_orders
WHERE order_ts >= %s AND order_ts < %s""",
(interval_start, interval_end),
)go deeper
Be ready to define idempotence in one sentence and name the everyday cause: a retry. Know that appending rows on every run is the classic non-idempotent shape.
Explain the mechanics — deterministic scope from run parameters, a replacing write such as delete-and-insert in one transaction or a merge on a key, and guarded side effects — and identify which of the three a given task violates.
Show you have operated this: talk about a task that died after writing but before reporting success, about half-written partitions, and about how you verify idempotence with row counts and checksums rather than trusting the code review.
Own the platform angle — make idempotence a property the framework enforces rather than a convention each team remembers, since one non-idempotent task poisons every downstream consumer that trusts a replayable pipeline.
## What idempotence means for a batch task An operation is idempotent when applying it more than once has the same observable effect as applying it once. For a pipeline task the observable effect is the state of the target: the rows in a table, the files under a partition prefix, the message that left the building. The task is idempotent if running it twice for the same input window leaves the target exactly as one successful run would have. Note what this is *not* a claim about. It is not a claim that the second run does less work — it may read the same source, burn the same compute and write the same rows again. It is a claim about the end state only. It is also not the same as determinism: a task can be idempotent (two runs today produce one copy of the data) while still producing a different answer than it did last year, because the source table has been updated in place since. ## Why an orchestrator forces the issue Any scheduler will execute your task more than once, whether you designed for it or not: - **Automatic retries.** A transient failure — a network blip, a killed worker, a throttled API — triggers attempt two, and attempt one may have written half its output already. - **False failures.** The task finished writing, then lost its heartbeat before reporting success. The orchestrator believes it failed; the data says otherwise. - **Replaying missed intervals.** A pipeline enabled with a start date in the past causes the scheduler to enumerate every window between then and now and queue a run for each. - **Explicit backfills.** You changed the transformation and want the last 90 days recomputed. - **A human clearing a run** because a downstream number looked wrong. - **Duplicate scheduling** — two schedulers, or a redeploy, briefly launching the same window twice. With an idempotent task all six are boring operations. Without one, each is a corruption event, and the characteristic symptom is that nobody notices for a week, until a revenue figure is exactly double for one day. ## The three properties that make it hold **1. Deterministic scope.** The task computes which input rows it reads and which slice of output it owns from its run parameters — the interval start and end handed to it — rather than from `now()` or from "whatever is newest". A filter like `WHERE event_ts >= now() - interval '1 day'` produces a different window on every execution, so a replay of an old interval reads the wrong data. **2. Deterministic, replacing write.** The task must be able to overwrite the slice it owns. Two common shapes: delete the interval's rows and insert the recomputed ones inside a single transaction, or `MERGE`/upsert on a stable natural key. Writing to a fresh partition path and atomically swapping it in is the object-store equivalent. Plain `INSERT` is the shape that breaks, because an append has no way to undo the previous attempt. **3. Guarded side effects.** Anything the task does outside the target store — sending mail, calling a webhook, dropping a file to a partner, incrementing a counter, triggering downstream consumers — has no partition to overwrite. Those steps must be suppressed during replay or made safe by a key the receiver can deduplicate on. ## Non-idempotent shapes you will meet in real pipelines - `INSERT INTO fact SELECT ... WHERE ts >= now() - interval '1 day'` — both the scope and the write are wrong. - Surrogate keys drawn from a sequence at load time, so a re-run assigns different keys and previously-built joins drift. - "Append now, deduplicate later": the nightly dedupe becomes load-bearing, and it usually has an edge case. - Accumulators: `UPDATE totals SET amount = amount + :delta`. - Consuming from a queue and deleting on read. The re-run finds nothing and quietly writes an empty partition — the same corruption, but silent. ## Idempotent versus reproducible A backfill often wants something stronger than idempotence: it wants the run to produce the answer the *original* run would have produced. That only holds if the input the task reads is immutable — an append-only raw landing zone partitioned by arrival or event time, rather than a source table that gets updated in place. This is why teams keep an untouched raw copy: idempotence makes replay safe, an immutable source makes replay meaningful. ## How to prove it In a test environment, freeze the source, run the task twice and assert the target is row-identical — a row count plus a checksum over the interval is usually enough. Then test the interrupted case: kill the task midway and re-run it. A delete-then-insert that is not wrapped in one transaction leaves a hole in the data between the two statements, which is exactly the failure that a naive rerun test never catches.
- Is a task that is idempotent also guaranteed to reproduce what the original run produced?No. Idempotence says two runs today leave the same state as one run today. If the source table is updated in place, replaying a two-year-old window reads today's version of the source and can legitimately produce a different answer. Reproducibility needs an immutable raw copy — append-only landing data partitioned by event or arrival time — not just an idempotent write.
- Where does a load that assigns surrogate keys from a sequence break idempotence?On re-run the sequence hands out new key values, so the same business rows get different identifiers. Downstream tables that stored the old keys now point at nothing, and dimension joins silently drop rows. The fix is a deterministic key — a hash of the natural business key, or a lookup that resolves an existing key before generating a new one.
- How would you test that a task really is idempotent before shipping it?Freeze the source, run the task twice against the same interval, and assert the target is row-identical: a row count plus a checksum over the interval. Then kill a run midway and re-run it — that catches writes where the delete and the insert are not in one transaction, which a clean double-run test never exposes.
An idempotent task is a switch labelled ON rather than a switch labelled TOGGLE: pressing it a second time changes nothing, so nobody has to remember whether it was already pressed.
saying these in an interview costs you the question
- Says the scheduler will not re-run an interval that already succeeded
- Claims retries are safe because failed tasks never write anything
- Plans to append rows and deduplicate in a later cleanup job
- Confuses idempotence with the task being fast or read-only
- Treats emails and webhooks as covered by an idempotent table write