skip to content

How does a scheduler's automatic catchup of missed windows differ from an operator-triggered backfill?

level: middleimportance: should knowfreq 62%

answer

  1. who asked for the runs
  2. one closes gaps the scheduler noticed itself
  3. the other names an explicit range of windows
  4. a past start date can queue hundreds of runs at once

basics

~20 s

Catchup is the scheduler filling in windows it never ran — every interval between the pipeline's start date and now — automatically as part of normal scheduling. A backfill is a human explicitly asking for a stated range to be re-run, usually because the logic or the source changed.

solid answer

~50 s

They produce similar-looking runs for different reasons. **Catchup** is the scheduler's own bookkeeping: when a pipeline is enabled with a start date in the past, or has been paused, the scheduler enumerates every window it has no run for and queues them. It is automatic, unbounded by anything but the start date, and it fires the moment you deploy. **Backfill** is out-of-band and deliberate: an operator names a range — "re-run 1 March to 30 April" — usually after changing a transformation, fixing a bug, or having a source corrected, and it re-runs windows that already succeeded. In practice the distinction matters operationally: catchup surprises people (deploy a pipeline with a year-old start date and hundreds of runs launch at once, hammering the source and blowing through compute budget), whereas a backfill is planned and can be given its own concurrency, priority and isolation. Both are only safe on idempotent, interval-scoped tasks.

code

text · 11 lines
text
start_date = 2024-01-01, cadence = daily, today = 2025-02-04

CATCHUP (scheduler-initiated, on enable)
  windows with no run: 2024-01-01 .. 2025-02-03   -> 400 runs queued now
  bounded by: start_date only
  targets:    windows that have never produced data

BACKFILL (operator-initiated, after a logic change)
  requested range: 2024-03-01 .. 2024-04-30       -> 61 runs queued
  bounded by: the range the operator names
  targets:    windows that ALREADY hold data -> the write must replace

go deeper

for a junior

Know the difference in one line: the scheduler fills windows it never ran, a person asks for a named range to be run again. Know that a past start date can queue a lot of runs at once.

for a middle

Explain what actually happens on enable — the scheduler enumerating every window without a run — and the mitigations: disable catchup, move the start date, cap active runs, and load history deliberately instead.

for a senior

Demonstrate the operational instincts: the stampede on the source system, notifications firing once per window, downstream triggers fanning out, and checking that the source can still answer for the range before replaying it.

for a principal

Own the default posture for the platform — catchup off, recent start dates, a documented and well-worn backfill path — because a team that finds backfilling scary stops changing transformations at all.

## Two ways the same window gets run again Orchestrators distinguish between windows they *should have run and never did*, and windows an operator *wants run again*. Filling the first set is catchup; requesting the second is a backfill. The runs look identical in the logs — same code, same interval parameters — but who initiates them, how they are bounded, and what they imply about existing data are different. ## Catchup: the scheduler closing its own gaps A scheduled pipeline has a start date and a cadence. The scheduler's job is to ensure a run exists for every window between the start date and now. When you enable a pipeline whose start date is 400 days in the past, there are 400 windows with no run, so it creates 400 runs. Same thing after a pause: unpause a pipeline that was off for a week and the week's windows appear. The consequences are the ones people meet the hard way: - **The stampede.** Hundreds of runs queue instantly. They compete for the same workers, hammer the source system — which may be an operational database that cannot take 200 parallel extracts — and consume real money on a metered warehouse. - **Side effects fire per window.** Every one of those runs executes the whole task graph, including the step that emails a summary or posts to a channel. - **Downstream fan-out.** If downstream pipelines trigger on this one's completion, the stampede propagates. - **Ordering is not guaranteed to be sequential** unless you constrain it, which matters when a window's computation depends on the previous window's output. That is why disabling catchup is such a common default: a newly deployed pipeline should usually start from the next window, with any history loaded through a deliberate, controlled backfill instead. The other common mitigations are setting the start date near the present, and capping how many runs of the pipeline may be active at once. ## Backfill: a bounded, deliberate request A backfill is initiated by a person or an automated job with an explicit range. The typical triggers: - the transformation logic changed and history must be recomputed under the new rules; - a source system was corrected and the old output is now wrong; - a new column or table was added and needs history populated; - a window failed silently and was only noticed later. What distinguishes it operationally is that it is *planned*. You know the range, so you can estimate the cost and duration up front, cap concurrency, route it to isolated compute, order it deliberately, and disable notification steps for the duration. Crucially, a backfill usually targets windows that **already have data**, which is exactly why the task's write must replace rather than append. ## What both require Neither is safe unless the task is interval-scoped and its write replaces the slice it owns. Catchup on an appending task multiplies rows by the number of windows re-run; catchup on a clock-scoped task makes every run resolve to the same output partition and stampede over one target. That is the pairing this whole topic exists to teach: catchup and backfill are only features if idempotence already holds — otherwise they are bulk corruption tools. ## The questions to ask before either happens 1. **How many windows?** Multiply by per-run duration and cost. Three years of hourly windows is over 26,000 runs. 2. **Does window N depend on window N-1?** If a run reads the previous window's output — a running balance, a state table — the runs must be sequential, and parallel catchup will produce garbage. 3. **What side effects fire per run?** Anything leaving the system needs suppressing. 4. **Who consumes the target while this happens?** Rebuilding a published table in place means consumers read a half-rebuilt mart. 5. **Is the source still able to answer for old windows?** Retention, deleted partitions and mutated rows all mean an old window may no longer be reproducible, and replaying it can *overwrite* good historical output with an empty or wrong result — the worst outcome in this whole topic. ## A practical default Deploy pipelines with catchup off and a recent start date, so turning something on is never an incident. Load history explicitly, in controlled chunks, with isolated compute and notifications disabled — and keep that backfill path well-worn and documented, because a team that finds backfilling painful will avoid it, and a team that avoids backfilling ships transformations it is afraid to change.

  • You deploy a pipeline with a start date a year in the past and hundreds of runs launch. What do you do first?
    Pause the pipeline to stop the queue growing, then decide whether that history is actually wanted. If it is, restart the pipeline from a recent window with catchup disabled and load history as a controlled backfill with capped concurrency and notifications off. Also check what already ran — an appending task may have duplicated whatever it touched.
  • When must catchup or backfill runs be executed sequentially rather than in parallel?
    When a window's computation reads the previous window's output — running balances, state carried forward, slowly changing dimensions built incrementally. Parallel execution then reads a not-yet-written or half-written predecessor and produces wrong results silently. Constrain the pipeline to one active run at a time, or restructure the task so each window is computable from source alone.
  • Why can replaying a very old window be more dangerous than not replaying it?
    Because the source may no longer be able to answer for it: raw data aged out of retention, rows were mutated in place, or a system was decommissioned. The replay then computes an empty or wrong result and — because the write replaces the slice — overwrites good historical output with it. Check source availability for the range before backfilling, and prefer building aside and swapping.

saying these in an interview costs you the question

  • Treats catchup and backfill as the same thing with different buttons
  • Assumes catchup runs windows one at a time in order
  • Deploys a pipeline with an old start date without capping concurrency
  • Backfills a range without checking whether the source still covers it
  • Believes the scheduler will skip windows that already have data

context