skip to content

After a multi-step data pipeline fails midway, when is restarting only the failed steps unsafe?

level: seniorimportance: should knowfreq 50%

answer

  1. what did the dead step already write
  2. the source has moved since step two
  3. green upstream is not the same as correct upstream
  4. temp space may not have survived
  5. if every step replaces its own window, restart is free

basics

~20 s

Restarting from the failure point assumes the completed steps' outputs are still valid and the failed step left nothing behind. It is unsafe when the failed step wrote partially, when earlier steps captured a source snapshot that has since moved, or when a shared staging area was cleared.

solid answer

~50 s

Restart-from-failed is cheap and usually right, but it rests on two assumptions. First, **the failed step left no partial state** — if it appended half its rows before dying, resuming duplicates them unless the step overwrites its target window or merges on a key. Second, **the completed steps' outputs are still the right ones** — if step 2 extracted a source snapshot at 02:00 and you restart at 14:00, later steps now combine a stale extract with whatever the restarted steps read fresh, producing a mixed-vintage result that no single point in time would have produced. Restart is also wrong when the failure is a *symptom* of an earlier step producing bad-but-successful output, and when steps share a per-run scratch area that has been cleaned. A full rerun costs compute but restores one consistent view. The durable fix is making every step atomic and self-overwriting, so restart is always safe.

code

sql · 8 lines
sql
-- NOT restartable: a partial run leaves rows behind, restart duplicates them
insert into analytics.orders
select * from staging.orders where order_date = :run_date;

-- restartable: the step owns and replaces exactly its own window
delete from analytics.orders where order_date = :run_date;
insert into analytics.orders
select * from staging.orders where order_date = :run_date;

go deeper

for a junior

Know that restarting from the failed step skips the completed ones entirely, and that anything the failed step already wrote is still there unless the step replaces its own output.

for a middle

Explain the two preconditions — no partial state left behind, completed outputs still valid — and give the append-versus-overwrite example that shows why one duplicates and the other does not.

for a senior

Diagnose live: check what the dead step wrote, whether the source has moved since the completed steps ran, whether scratch space survived, and whether the failure was really a symptom of a green-but-wrong upstream step.

for a principal

Push the decision out of the incident entirely by mandating atomic, window-scoped, snapshot-pinned steps, so restart-from-failed becomes a cost choice rather than a correctness gamble under time pressure.

## What restart-from-failed actually promises When a run fails at step 4 of 6, the attractive option is to resume at step 4 and leave steps 1–3 alone. Every orchestrator offers some version of this, and for well-built pipelines it is the right default: it saves the expensive extract, it shortens recovery, and it keeps the deadline in reach. But it is an *assertion*, not a convenience. You are asserting that the world is in exactly the state a fresh run at step 4 expects. ## Failure mode 1: the failed step's partial writes A step that dies partway has usually already done something. If it appends rows, the restarted attempt appends them again and the target silently double-counts — the worst kind of failure, because the run then goes green and the number is wrong. Restart is only safe here if the step's write is **all-or-nothing or self-replacing**: write to a scratch location and swap at the end, overwrite the whole target window for the run, or merge on a key so re-applying is a no-op. If the step is a blind append into a table with no natural key and no window boundary, restarting it is a data-corruption decision, and the correct recovery is to delete the run's partial output first, or to rerun from a step that rebuilds the target wholesale. ```sql -- restartable: the step replaces exactly its own window delete from analytics.orders where order_date = :run_date; insert into analytics.orders select * from staging.orders where order_date = :run_date; ``` ## Failure mode 2: the moving source This is the subtle one and the reason seniors get asked it. Suppose step 2 read a mutable source table at 02:00 and landed an extract. The run failed at step 4. You restart at 14:00. Steps 4–6 now run against the 02:00 extract, but if any of them also read the source directly — a lookup, a dimension join, a late reconciliation — they see the 14:00 state. The output is a blend of two points in time that never coexisted: an order that did not exist at 02:00 joined to a customer record that changed at 11:00, or worse, orders from the extract whose keys are missing from a dimension the restarted step read later. Nothing errors. The result is simply not a coherent snapshot of anything. The test: **does any step downstream of the restart point read state that has changed since the completed steps ran?** If yes, either pin that read to the same snapshot (a version, a timestamp bound, a frozen extract) or rerun from the step that established the snapshot. ## Failure mode 3: the vanished scratch space Many pipelines write intermediate results to per-run temp locations, and many platforms clean those on failure, on a lifecycle policy, or when the worker went away. A restart at step 4 that expects step 3's temp output to still exist fails immediately in the best case, and reads a stale leftover from an earlier run in the worst. ## Failure mode 4: the failure was a symptom Step 4 crashed on a null it should never have seen. Step 2 succeeded — but it succeeded at loading a truncated file. Restarting at step 4 either fails again or, after you "fix" step 4 to tolerate nulls, produces confidently wrong output from bad input. Before choosing a restart point, establish that everything upstream of it is not merely *green* but *right*. Green is a claim about exit codes. ## Failure mode 5: parameters recomputed at restart If a step derives its target window from the wall clock at execution time rather than from the run's own pinned parameters, restarting it twelve hours later processes a different window than the completed steps did. The pipeline must carry its window as run-scoped parameters that every attempt and every restart resolves identically. ## When a full rerun is the right call Rerun the whole thing when the source is mutable and you need one coherent snapshot, when the failed step is not self-replacing and you cannot cleanly remove its partial output, when you have changed transformation logic that earlier steps also embody, and whenever the cost of being subtly wrong exceeds the compute cost of doing it again. For most warehouse pipelines the compute is cheap relative to a wrong published number; the real constraint on a full rerun is usually the *deadline*, not the money. ## Designing so the question stops being interesting The mature answer is to make restart-from-failed unconditionally safe: - every step writes atomically — scratch then swap, or a single window-replacing statement; - every step is scoped to a window derived from run parameters, not from the clock at execution time; - source reads inside a run are pinned to one snapshot (an immutable landed extract, a version, a timestamp predicate); - intermediate outputs are durable and addressed by the run's parameters, so a restart can find them; - steps validate their inputs, so an upstream truncation fails at step 2 rather than corrupting step 5. When those hold, restart-from-failed is a cost optimisation with no correctness argument attached — which is exactly where you want the decision to be at 03:00.

  • What single design property makes restart-from-failed safe for any step?
    Atomic, self-replacing writes scoped to run parameters. If each step either fully replaces its own output window or merges on a key, and derives that window from the run rather than the clock, then re-executing it from a clean start is indistinguishable from having run it once. Partial writes leave nothing to duplicate.
  • The failure was caused by an upstream step that succeeded on truncated input. Where do you restart?
    From the step that produced the bad input, not from the step that crashed. A green exit code only says the step did not error; it says nothing about completeness. Add an input validation gate at that boundary so truncation fails loudly at the source rather than surfacing as a mysterious crash several steps later.

saying these in an interview costs you the question

  • Assuming the orchestrator undoes a failed step's writes before restart
  • Restarting hours later against a source that has since changed
  • Treating a green upstream step as proof its output is correct
  • Recomputing the processing window from the clock when restarting
  • Always doing a full rerun because reasoning about restart feels hard

context