A nightly report pipeline stamps the same date on every run — how do you make each run recompute it?
answer
- who computes it, and when
- hand over a recipe, not a result
- a function the run invokes
- factory called once per execution
- per run, not per element
basics
~20 sReplace the baked-in value with a factory the pipeline calls when a run starts: a step that accepts a function returning a source defers construction to execution, so each run computes its own date instead of reusing the one captured during assembly.
solid answer
~50 sThe date is wrong because it was computed at assembly and captured into a description that is then reused unchanged every night. Nothing in a run revisits the builder, so the captured value can never refresh. The fix is to move the construction into a **factory** — a function that builds the source, handed to a step that invokes it once per execution. Each run then calls the factory, computes its own date, and builds its own stages around that value. Compute it once inside the factory and let the stages capture it there, so every stage of one run agrees on one date; computing it inside a per-value body instead would give each row a different date. The cost is that the enclosed construction repeats on every run, so it must be cheap and safe to repeat.
code
pseudocode · 11 lines// frozen: the date is computed once, while describing the chain
reportDate = today()
pipeline = sourceOfRowsFor(reportDate)
.transformEach(row -> stamp(row, reportDate))
// deferred: the factory is invoked once per execution
pipeline = deferPerRun(function()
reportDate = today() // one value per run
return sourceOfRowsFor(reportDate)
.transformEach(row -> stamp(row, reportDate))
end)go deeper
Remember that a value computed while building the chain is fixed from then on. To get a fresh one, the pipeline has to be given a function it can call later.
Explain the mechanism: a step that takes a source-producing function invokes it once per execution, so construction and any value inside it belong to that run.
Diagnose the stale report from the symptom, place the computation at run scope rather than element scope, and name what the deferral now repeats on every run.
Weigh one shared description with an internal factory against rebuilding per run, and set the rule for which pipelines in your platform may hold assembly-time state at all.
## Why the value is stuck An assembled chain is a description, and a description is a value. When the builder computes today's date and hands it to the stages, that date stops being a computation and becomes a constant inside the description. The nightly run executes the description; it does not re-enter the builder, and nothing in the chain knows that the constant was once the result of a call. The pipeline is behaving correctly and reporting a date from the night the service last started. This is the general shape of the defect: **anything evaluated at assembly is frozen for the lifetime of the chain.** A date, a batch identifier, a snapshot of configuration, a cursor into a data set — each one is correct on the first run and stale forever after. ## The fix: build per execution The remedy is to hand the pipeline something it will *call* rather than something it has already *computed*. A step exists in every ecosystem for this: it accepts a function that returns a source, and it invokes that function once per execution, building a fresh source for that run. What changes: - The date is computed **inside** the factory, so its value belongs to a single execution. - The stages built in that factory capture that execution's date, so every stage of one run agrees on one date. - The outer chain still exists as a stable value that can be stored, shared and run repeatedly — the deferral is internal to it. What it costs: - The construction inside the factory repeats on every run, so it must be cheap. - It must also be **safe to repeat**: a factory that mutates shared state or performs an external write does so once per run, which is usually not what a builder's author intended. - Two overlapping runs invoke the factory twice and must not share mutable state through it, or the isolation the deferral bought is lost again. ## Three placements, three different frequencies | Where the date is computed | How often it is evaluated | Result | |---|---|---| | In the builder, at assembly | once, ever | every run reports the build date | | In a factory invoked per execution | once per run | correct: one date per report | | In a per-value body | once per element | rows in one report disagree with each other | The third row is the trap that catches people who have understood the first. Moving the call "somewhere later" is not the fix; moving it to the *right frequency* is. A report needs run-scoped values, so run-scoped is where the computation goes. ## Choosing between deferral and rebuilding There are two honest ways to get a per-run value, and an interviewer will usually push for the comparison: 1. **Keep one assembled chain and defer its variable part.** The chain stays a long-lived value; only the inner factory re-runs. Assembly cost is paid once, and the shared description can be handed to schedulers, tests and other callers. 2. **Rebuild the chain on every run.** The builder is called per run, so everything in it is fresh by construction and no deferral step is needed. It is simpler to read and pays full assembly cost per run, and it gives up the single shared description — anything holding the old chain keeps the old behaviour. Rebuilding is a fine answer for a pipeline run once a night, where assembly cost is irrelevant. It is a poor answer when the chain is published to other components, or when runs are frequent enough that rebuilding is measurable. The deferral step is what lets one stable description still produce per-run values, which is why it exists. ## Verifying the fix The test is behavioural, not structural: run the assembled pipeline twice with a controllable clock and assert that the two runs disagree on the value and that within one run every stage agrees. A test that only checks the first run passes on the broken version too, which is precisely why the defect survives to production and shows up as a report dated the day of the last deployment.
- Where should the per-run value be computed if several stages need the same one?Once, at the top of the factory, with the stages built beneath it capturing that value. Every stage of a run then agrees on it. Computing it separately in each stage risks disagreement, and computing it per element guarantees it.
- What does deferring construction cost?The enclosed build work repeats on every execution, so it must be cheap and safe to repeat. A factory that writes externally or mutates shared state performs that effect once per run, which is rarely what a builder's author expected.
- How does this behave when the same chain is run twice concurrently?Each execution invokes the factory separately and gets its own values and stages, so the two runs do not share the deferred value. That isolation is lost if the factory reads or writes mutable state defined outside it.
- When is rebuilding the whole chain per run the better answer?When assembly is cheap relative to the run and the chain is not shared. It needs no deferral step and is fresh by construction. It is the wrong answer when the description is published to other components, which would keep running the old one.
A printed menu lists dishes; the kitchen cooks only when an order is placed. Print tonight's date on the menu at the print shop and it is wrong forever; write it when the order is taken and it is right every night.
saying these in an interview costs you the question
- Thinks re-running an assembled chain re-evaluates its captured values
- Moves the computation into a per-value body, dating each row differently
- Wraps the value in a step that still evaluates it during assembly
- Believes the only fix is rebuilding the entire chain every run
- Claims a stored chain is unsafe to reuse and must be copied per run
- Ignores that the deferred factory now repeats its side effects per run