skip to content

A retraining graph rerun after a pricing fix reused a cached aggregate step, so the fix never reached the model — why?

level: seniorimportance: should knowfreq 44%

answer

  1. a skip is a claim about declarations
  2. the key covers declared inputs and code
  3. an undeclared read is a hidden edge
  4. the skip propagates down the branch
  5. green run, stale model

basics

~20 s

Step caching skips a step whose declared inputs and code are unchanged. The aggregate step read the pricing table without declaring it, so the fix was invisible to the cache key and the branch was reused instead of recomputed.

solid answer

~40 s

A graph skips work by asking a narrow question: did anything this step **declared** change? The key covers the declared inputs and the step's own code and parameters. Here the aggregate step read a pricing reference table inside its body without listing it as an input, which makes it a **hidden edge** — the fix changed data the graph was never told about, the key came out identical, and the cached output was reused. The model therefore trained on pre-fix aggregates and no step reported an error. The fix is to make the declaration complete: every reference table a step reads is an input, and the step's code and parameter version belong in the key too. The mirror-image failure is over-declaring something that changes every run, which turns every skip into a miss.

go deeper

for a junior

Know that a graph can skip a step it has already run, and that it decides by looking only at what the step declared it reads — not at everything the step's code actually touches.

for a middle

Explain what goes into the key — declared input identity plus the step's code and parameters — and why an undeclared reference table produces a matching key and a reused, stale output.

for a senior

Diagnose this from provenance: the reused output predates the fix, the skip propagated down the branch, and a no-cache rerun of the same dates produces different data. Then close it by completing the declaration.

for a principal

Treat declaration completeness as a platform guarantee: never-skip marking for steps with unavoidable ambient reads, pointers resolved once per run, and a scheduled full rebuild whose diff against the incremental path is the standing audit.

## How a step cache decides to skip In the nightly demand-forecast graph, most branches do not change between runs — the store dimension, the holiday calendar, the price list. Recomputing them nightly is waste, so the graph asks, for each step: > *Do I already hold an output produced from exactly these declared inputs, by exactly this step definition?* If yes, the step is skipped and its previous output stands in. The key behind that question covers two things, and only two: - the identity of every **declared input**; - the step's own **definition** — its code version and its parameters. The cache is therefore only as truthful as the declarations. It is not watching the file system, and it does not know what the step's body actually touched. ## The hidden edge The pricing fix went into a reference table that the aggregate step reads to convert units into price-adjusted demand. The step's declaration listed the sales events and the store dimension; the pricing table was read directly in the body and never declared. So: 1. The fix landed in the pricing table. 2. The graph was rerun for the affected dates. 3. The aggregate step's declared inputs were unchanged and its code was unchanged, so the key matched the previous run. 4. The step was skipped and its stale output reused. 5. The fit step's own inputs were therefore unchanged too, so it was skipped as well — the skip propagated down the branch. 6. The published model was byte-identical to yesterday's, and every step was green. That last property is what makes this class hard: **there is no failure**. The run is fast and successful, which looks like the caching working perfectly. ## Confirming it 1. **Compare the output's provenance to the fix.** The reused partition was produced by an earlier run than the one that carried the fix. A step whose output predates the change it was meant to absorb is the whole diagnosis. 2. **Rerun the branch with caching disabled** for the same dates and diff the outputs. If the no-cache output differs, the cache key is missing an input. 3. **List what the step body reads** against what it declares. Reference tables, configuration blobs and lookup files are the usual omissions, because they feel like environment rather than data. ## The two ways a key goes wrong | Symptom | Cause | Consequence | |---|---|---| | Steps skip when they should rerun | The key omits something the step reads — an undeclared reference table, an unpinned code or parameter version | Silent staleness: the fix never lands, the run is green | | Steps never skip | The key includes something that changes every run — a wall-clock read, a per-run temporary path, a mutable pointer resolved inside the step | No reuse at all: full-cost nights and a backfill that cannot be throttled | Both are declaration defects, and they pull in opposite directions, which is why teams tend to overcorrect from one straight into the other. ## Making declarations honest - **Every read is an input.** Reference tables, price lists, holiday calendars, category mappings. If the step's output would change when that data changes, it belongs in the declaration. - **The step definition is part of the key.** A change to the aggregation logic or to a parameter alters the output for identical data; if it is outside the key, a genuine code fix is skipped exactly like the pricing fix was. - **Resolve mutable pointers once, at the top of the run.** A step that resolves a `latest` pointer inside its own body has an input that is not stable across attempts; resolve it once and pass the resolved identity down, so every step in the run sees the same thing. - **Keep ambient reads out of the body.** Wall-clock reads, environment lookups and per-attempt temporary paths make a step either uncacheable or wrongly cached; pass them in as declared parameters instead. - **Make caching opt-out per step.** A step that genuinely cannot declare everything it reads should be marked never-skip rather than quietly wrong. - **Rebuild without cache periodically.** A scheduled full recomputation compared against the incremental one is the audit that catches a hidden edge before a stakeholder does. Platforms differ in how much of the surrounding environment they fold into the key — some include the step's runtime image, some only the declared data — so the safe habit is to pin the step's version explicitly rather than assume the platform noticed it. ## The operational lesson Step caching is a claim about **completeness of declarations**, not about data freshness. The moment a step reads something it did not declare, the graph's skip logic becomes a bug generator, and it generates bugs that look like successful runs. When a fix demonstrably fails to reach the model while everything is green, the undeclared read is the first place to look.

  • Same graph, opposite symptom: no step is ever skipped. What causes that?
    A step declares something that changes on every run — a wall-clock read, a per-run temporary path, or a mutable pointer resolved inside the body. Its key never matches, so it reruns and drags its whole downstream branch with it. Resolve such values once at the start of the run and pass the fixed identity into the step.
  • Should the step's code version be part of the cache key?
    Yes. A change to the aggregation logic or a parameter changes the output for identical data, so leaving it out means a genuine code fix is skipped the same way the pricing fix was. Platforms differ in how much of the runtime they fold in, so pin the step's version explicitly rather than relying on the default.
  • How would you stop this class of defect reaching a published model again?
    Add a completeness check between the graph and the publish step: the published model's provenance must list an input version at or after the change it was supposed to absorb. A periodic no-cache rebuild, diffed against the incremental outputs, catches the remaining hidden edges before anyone downstream does.

saying these in an interview costs you the question

  • Assumes the cache watches files and notices any change
  • Thinks a cached output expires on a timer
  • Leaves reference tables out of the declaration as environment
  • Excludes the step's code and parameters from the key
  • Resolves a mutable latest pointer inside each step body
  • Disables caching everywhere instead of fixing the declarations