skip to content

Idempotency, Backfills and Catchup

Making a pipeline safe to run twice, and then using that property to replay history when logic changes or a source arrives late. Interviewers ask you to design a task that can be rerun for any past day without duplicating rows — it is the single most common data-engineering design question.

on this pageshow

questions

6

What makes a scheduled batch task idempotent, and why do orchestrators require it?

level: juniorimportance: must knowfreq 75%

answer

  1. think about what a retry leaves behind
  2. same window in, same end state out
  3. scope from the interval, replace don't append
  4. side effects have no partition to overwrite

basics

~20 s

A task is idempotent when running it again for the same input window leaves the target in the same state as one successful run. Orchestrators retry, replay and re-run tasks constantly, so anything else duplicates data.

solid answer

~50 s

Idempotence means the observable end state is the same whether the task ran once or five times for the same input window. Three things make it hold: the task derives **which slice it reads and writes** from its run parameters (the interval), not from the wall clock; the write **replaces** that slice — a partition overwrite or a merge on a natural key — instead of appending to it; and any external side effect (email, webhook, file drop) is either suppressed on replay or keyed so the receiver ignores duplicates. It matters because every orchestrator re-runs tasks: automatic retries after transient failures, a worker that died after writing but before reporting success, replaying missed intervals, an operator clearing a run, or a deliberate backfill after a logic change. If the task is idempotent all of those are routine; if not, each one is a silent data-corruption incident.

code

python · 23 lines
python
# not idempotent: scope from the clock, write appends
def load(conn):
    conn.execute("""
        INSERT INTO sales_daily
        SELECT order_id, amount, order_ts::date AS dt
        FROM   raw_orders
        WHERE  order_ts >= now() - interval '1 day'
    """)

# idempotent: scope from the run interval, write replaces
def load(conn, interval_start, interval_end):
    with conn.begin():
        conn.execute(
            "DELETE FROM sales_daily WHERE dt >= %s AND dt < %s",
            (interval_start, interval_end),
        )
        conn.execute(
            """INSERT INTO sales_daily
               SELECT order_id, amount, order_ts::date
               FROM   raw_orders
               WHERE  order_ts >= %s AND order_ts < %s""",
            (interval_start, interval_end),
        )

go deeper

for a junior

Be ready to define idempotence in one sentence and name the everyday cause: a retry. Know that appending rows on every run is the classic non-idempotent shape.

for a middle

Explain the mechanics — deterministic scope from run parameters, a replacing write such as delete-and-insert in one transaction or a merge on a key, and guarded side effects — and identify which of the three a given task violates.

for a senior

Show you have operated this: talk about a task that died after writing but before reporting success, about half-written partitions, and about how you verify idempotence with row counts and checksums rather than trusting the code review.

for a principal

Own the platform angle — make idempotence a property the framework enforces rather than a convention each team remembers, since one non-idempotent task poisons every downstream consumer that trusts a replayable pipeline.

## What idempotence means for a batch task An operation is idempotent when applying it more than once has the same observable effect as applying it once. For a pipeline task the observable effect is the state of the target: the rows in a table, the files under a partition prefix, the message that left the building. The task is idempotent if running it twice for the same input window leaves the target exactly as one successful run would have. Note what this is *not* a claim about. It is not a claim that the second run does less work — it may read the same source, burn the same compute and write the same rows again. It is a claim about the end state only. It is also not the same as determinism: a task can be idempotent (two runs today produce one copy of the data) while still producing a different answer than it did last year, because the source table has been updated in place since. ## Why an orchestrator forces the issue Any scheduler will execute your task more than once, whether you designed for it or not: - **Automatic retries.** A transient failure — a network blip, a killed worker, a throttled API — triggers attempt two, and attempt one may have written half its output already. - **False failures.** The task finished writing, then lost its heartbeat before reporting success. The orchestrator believes it failed; the data says otherwise. - **Replaying missed intervals.** A pipeline enabled with a start date in the past causes the scheduler to enumerate every window between then and now and queue a run for each. - **Explicit backfills.** You changed the transformation and want the last 90 days recomputed. - **A human clearing a run** because a downstream number looked wrong. - **Duplicate scheduling** — two schedulers, or a redeploy, briefly launching the same window twice. With an idempotent task all six are boring operations. Without one, each is a corruption event, and the characteristic symptom is that nobody notices for a week, until a revenue figure is exactly double for one day. ## The three properties that make it hold **1. Deterministic scope.** The task computes which input rows it reads and which slice of output it owns from its run parameters — the interval start and end handed to it — rather than from `now()` or from "whatever is newest". A filter like `WHERE event_ts >= now() - interval '1 day'` produces a different window on every execution, so a replay of an old interval reads the wrong data. **2. Deterministic, replacing write.** The task must be able to overwrite the slice it owns. Two common shapes: delete the interval's rows and insert the recomputed ones inside a single transaction, or `MERGE`/upsert on a stable natural key. Writing to a fresh partition path and atomically swapping it in is the object-store equivalent. Plain `INSERT` is the shape that breaks, because an append has no way to undo the previous attempt. **3. Guarded side effects.** Anything the task does outside the target store — sending mail, calling a webhook, dropping a file to a partner, incrementing a counter, triggering downstream consumers — has no partition to overwrite. Those steps must be suppressed during replay or made safe by a key the receiver can deduplicate on. ## Non-idempotent shapes you will meet in real pipelines - `INSERT INTO fact SELECT ... WHERE ts >= now() - interval '1 day'` — both the scope and the write are wrong. - Surrogate keys drawn from a sequence at load time, so a re-run assigns different keys and previously-built joins drift. - "Append now, deduplicate later": the nightly dedupe becomes load-bearing, and it usually has an edge case. - Accumulators: `UPDATE totals SET amount = amount + :delta`. - Consuming from a queue and deleting on read. The re-run finds nothing and quietly writes an empty partition — the same corruption, but silent. ## Idempotent versus reproducible A backfill often wants something stronger than idempotence: it wants the run to produce the answer the *original* run would have produced. That only holds if the input the task reads is immutable — an append-only raw landing zone partitioned by arrival or event time, rather than a source table that gets updated in place. This is why teams keep an untouched raw copy: idempotence makes replay safe, an immutable source makes replay meaningful. ## How to prove it In a test environment, freeze the source, run the task twice and assert the target is row-identical — a row count plus a checksum over the interval is usually enough. Then test the interrupted case: kill the task midway and re-run it. A delete-then-insert that is not wrapped in one transaction leaves a hole in the data between the two statements, which is exactly the failure that a naive rerun test never catches.

  • Is a task that is idempotent also guaranteed to reproduce what the original run produced?
    No. Idempotence says two runs today leave the same state as one run today. If the source table is updated in place, replaying a two-year-old window reads today's version of the source and can legitimately produce a different answer. Reproducibility needs an immutable raw copy — append-only landing data partitioned by event or arrival time — not just an idempotent write.
  • Where does a load that assigns surrogate keys from a sequence break idempotence?
    On re-run the sequence hands out new key values, so the same business rows get different identifiers. Downstream tables that stored the old keys now point at nothing, and dimension joins silently drop rows. The fix is a deterministic key — a hash of the natural business key, or a lookup that resolves an existing key before generating a new one.
  • How would you test that a task really is idempotent before shipping it?
    Freeze the source, run the task twice against the same interval, and assert the target is row-identical: a row count plus a checksum over the interval. Then kill a run midway and re-run it — that catches writes where the delete and the insert are not in one transaction, which a clean double-run test never exposes.

An idempotent task is a switch labelled ON rather than a switch labelled TOGGLE: pressing it a second time changes nothing, so nobody has to remember whether it was already pressed.

saying these in an interview costs you the question

  • Says the scheduler will not re-run an interval that already succeeded
  • Claims retries are safe because failed tasks never write anything
  • Plans to append rows and deduplicate in a later cleanup job
  • Confuses idempotence with the task being fast or read-only
  • Treats emails and webhooks as covered by an idempotent table write

context

open as a page

Why is a task that appends its window's rows unsafe to re-run, and what write pattern fixes it?

level: middleimportance: must knowfreq 68%

basics

~20 s

An append has no way to remove what a previous attempt wrote, so every re-run adds another copy. Fix it by making the write replace the slice: delete-then-insert for that window in one transaction, a merge on a natural key, or an atomic partition swap.

open as a page

How does a scheduler's automatic catchup of missed windows differ from an operator-triggered backfill?

level: middleimportance: should knowfreq 62%

basics

~20 s

Catchup is the scheduler filling in windows it never ran — every interval between the pipeline's start date and now — automatically as part of normal scheduling. A backfill is a human explicitly asking for a stated range to be re-run, usually because the logic or the source changed.

open as a page

Why should a batch task derive its target partition from the run's interval rather than the current date?

level: middleimportance: should knowfreq 58%

basics

~20 s

Because the wall clock moves and the interval does not. A task scoped by the current date reads and writes whatever today happens to be, so retries after midnight, replays and backfills all land in the wrong slice and collide with each other.

open as a page

A backfill re-sent last month's customer emails — how do you make pipeline side effects replay-safe?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Separate computing data from delivering it. Put external effects — email, webhooks, partner file drops — in a step that only fires for current runs, gated on a backfill flag or on the window being recent, so a replay recomputes tables without re-notifying anyone.

open as a page

How would you plan a three-year backfill so it neither starves the daily schedule nor explodes cost?

level: principalimportance: should knowfreq 38%

basics

~20 s

Size it first: windows times per-run cost and duration. Then chunk into coarser windows where the transform allows, run it on isolated compute with a hard concurrency cap, checkpoint progress so it resumes, build into a shadow target and swap after validation.

open as a page