skip to content

Why is a nightly retraining run for store demand forecasts built as a graph of declared steps rather than one long script?

level: juniorimportance: should knowfreq 50%

answer

  1. granularity of failure
  2. each step declares what it reads
  3. completed outputs are restart boundaries
  4. unchanged branches can be skipped
  5. one task per store region

basics

~20 s

A graph makes each step's inputs and outputs explicit, so the run gains restart points: a failed step reruns alone, unchanged branches are skipped, independent steps fan out in parallel, and progress is visible per step instead of per run.

solid answer

~40 s

In a script, the unit of success is the whole run: a failure in the evaluation stage at 05:00 throws away the aggregation and the fit that already succeeded. Expressing the same work as steps that each declare what they read and what they write turns every completed output into a **restart boundary** — the rerun starts at the failed step, reuses the outputs already on disk, and leaves untouched branches alone. The declarations also let independent work fan out (one aggregation task per store region rather than a loop), let each step ask for the resources it actually needs, and give per-step status instead of one red run. The cost is discipline: a step that reads something it never declared is invisible to the graph.

go deeper

for a junior

Be able to say what a step declares — its inputs, its outputs, its resources — and why that turns a completed output into a place the run can resume from after a failure.

for a middle

Explain how the edges are derived from declarations rather than written by hand, how fan-out per partition falls out of disjoint outputs, and why an undeclared read is invisible to the scheduler.

for a senior

Show where you would draw the boundaries in a real nightly run, and justify the ones you refused to split because scheduling and durable-write overhead would cost more than the restartability bought.

for a principal

Argue the operational contract the shape has to guarantee before anyone relies on it: atomic publication, durable completion records, and a per-step status that an on-call engineer can act on at 05:00.

## What a declared step actually is A nightly retraining run for per-store, per-item demand forecasts is not one job; it is a set of **steps** whose edges come from data. Each step declares three things: - the **inputs** it reads — a date range of sales events, a store dimension table, yesterday's aggregates; - the **outputs** it writes — one partition of daily units per store-region, a fitted model artifact, an evaluation report; - the **resources** it needs — a few slots on a shared batch engine for an aggregation, one large reservation for the fit. The edges follow from those declarations: a step that reads what another step writes runs after it. The whole structure is a **directed acyclic graph**; acyclic because a node that could feed itself has no well-defined restart point. A long script does the same work with none of the declarations. Its dependencies live in statement order, its intermediate results live in memory, and its unit of success is the exit code of the whole process. ## What the declarations buy - **Restart granularity.** A completed, durably written output is a point the run never has to reproduce. Failing at evaluation costs you evaluation, not the six hours of fitting behind it. - **Skipping unchanged work.** If a branch's declared inputs did not change since the last run, its outputs are still valid and the graph can reuse them instead of recomputing. - **Parallel fan-out.** A loop over four hundred store regions inside a script is sequential by construction. The same work declared as four hundred tasks with disjoint outputs runs as wide as the engine allows. - **Failure isolation.** With one task per region, a corrupt source file for one region fails one task. In a script, the same file ends the run. - **Right-sized resources.** The fit wants a large, long reservation; the aggregations want many small short ones. Only a graph can ask for both. - **Honest status.** "The run failed" is not an operational signal. "Aggregation is green, fit is green, publish failed" tells the on-call engineer what to do at 05:00. - **Targeted reruns.** When a defect is found in one transformation, you rerun that step and its descendants rather than the whole night. ## Script against graph | Concern | One long script | Graph of declared steps | |---|---|---| | Unit of failure | The entire run | One step, often one partition | | Cost of a late failure | Everything repeats | Only the failed step and its descendants | | Dependencies | Implicit in statement order | Declared, and checkable before anything runs | | Parallelism | Whatever the process does itself | Scheduler fans out independent tasks | | Reuse between runs | None; every run starts empty | Unchanged branches can be skipped | | Resource request | One shape for all the work | Per step, sized to that step | | Status | One exit code | Per-step state, with per-step timing | ## Where the boundaries belong Steps are not split by code aesthetics. The rule is: **a boundary goes where a durable output is produced that another step reads.** 1. **A durable artifact is written.** Aggregates land in a partition, features land in a table, the model lands as an artifact. Each is a natural node. 2. **The resource shape changes.** Wide cheap fan-out and one heavy long fit do not belong in the same node; merging them forces the big reservation to be held for the small work. 3. **The failure mode changes.** Reading external source files fails differently from arithmetic on already-validated rows; separate nodes let you retry one and not the other. 4. **Data-parallel work exists.** Where the computation is independent per region or per day, one task per partition buys isolation and throughput at once. Over-splitting has a real cost. Every node carries scheduling latency, a durable write of its output and a completion record, so decomposing a two-second transformation into six steps can cost more in overhead than it ever saves in restarts. ## What the shape demands in return The graph can only reason about what it was told. Two obligations follow, and both are routinely broken: - **Declare everything the step reads.** A reference table read inside the step body but absent from its declaration is a hidden edge: the graph will happily reuse a stale output because, as far as it knows, nothing changed. - **Publish outputs atomically.** A completion record means "this output exists in full". If a step can be recorded complete while its output is half-written, restartability is a fiction — the next run resumes on top of a partial file. With those two held, the graph gives a usable operational promise: any interrupted night can be resumed, and resuming costs the work that had not finished rather than the work that had.

  • How fine should the steps be — one per stage, or one per store region?
    Split where a durable output is produced and where work is genuinely data-parallel. One task per region buys isolation and throughput when regions are independent. Stop splitting when the scheduling latency and the durable write of a step's output cost more than the restart it would have saved.
  • What does the graph need beyond edges before it can actually resume a failed night?
    Durable per-step completion records keyed to that run, and outputs published atomically so a completion record truthfully means the output exists in full. Without both, a resume can skip a step whose output is half-written, which is worse than rerunning from the top.

saying these in an interview costs you the question

  • Treats the graph as a dashboard picture, not as restart state
  • Says rerunning the whole script from the top is equivalent
  • Splits steps by code file length instead of by durable output
  • Assumes step order alone is the dependency, with nothing declared
  • Believes more, smaller steps always make a run faster
  • Calls a step complete while its output is still half-written