skip to content

A stream pipeline assembled once at start-up writes a log line immediately, before any run — why?

level: middleimportance: must knowfreq 62%

answer

  1. describing is still running code
  2. arguments evaluate where they are written
  3. bodies wait, expressions do not
  4. one evaluation at build, many per run
  5. eager statement inside a deferred chain

basics

~20 s

Building the chain runs ordinary code. The expressions handed to each step are evaluated as the chain is described, so a log line or a computed value written there fires once at assembly, not on each later run of the pipeline.

solid answer

~40 s

Deferral applies to the *step bodies*, not to the code that builds the chain. Assembling a pipeline is ordinary program execution: every statement in the builder runs, and every argument expression handed to a step is evaluated right there, before anything is run. What gets deferred is the function value a step is given — it is created at assembly and invoked later, once per value or once per execution. So `sourceOf(loadRows())` performs the load while the chain is being described, and a log statement written in the builder prints exactly once, at start-up, no matter how often the pipeline later runs. The rule of thumb is: if it is an expression, it has already happened; if it is a body the pipeline will call, it has not.

code

pseudocode · 9 lines
pseudocode
function buildPipeline()
    log("building report")          // runs now, during assembly
    startedAt = currentTime()       // evaluated now, exactly once
    return sourceOfRows()
             .transformEach(row -> stamp(row, startedAt))
end

pipeline = buildPipeline()          // "building report" already printed
runEveryNight(pipeline)             // only the per-value body runs here

go deeper

for a junior

Remember that building a chain is normal code that runs immediately. Only the functions you hand to steps are kept for later.

for a middle

Explain the split precisely: argument expressions are evaluated at assembly, while the bodies given to steps are created then and invoked during a run.

for a senior

Show how it presents in production: work charged to start-up instead of the run, a frozen timestamp, or an initialisation failure that is really a pipeline defect.

for a principal

Treat it as a review standard. Decide what builders are allowed to do at assembly, and require anything that must be fresh or costly to be expressed as a body the pipeline invokes.

## Assembly is ordinary execution The headline claim about declarative pipelines — that describing work is not doing work — is true of the **chain**, and routinely misread as being true of the **code that builds the chain**. Building is ordinary, eager program execution. Statements in the builder run in order, calls in it are made, and arguments are evaluated before the step they are handed to is applied. Assembly produces a description; producing it is itself work that happens now. That is why a log line in the builder appears in the start-up logs of a service whose pipeline will not run until midnight. Nothing malfunctioned: the line is a statement, and statements in a function that is being called execute. ## What is deferred and what is not The distinction that actually matters is between a **value** handed to a step and a **body** handed to a step. | What the step receives | When it is produced | When its effect happens | |---|---|---| | An expression yielding a value | at assembly | at assembly | | A source produced by calling a function | at assembly | at assembly | | A function the step calls per value | created at assembly | once per value, during a run | | A function the step calls per execution | created at assembly | once per run | Creating a function value is cheap and has no effect; invoking it is where the work is. So in a chain built from a transforming step and a per-value body, the body's *text* is fixed at assembly and its *execution* waits. Anything that is not inside such a body has already happened by the time the chain exists. This is a property of the host language's evaluation order rather than of the pipeline. Where a language evaluates arguments before the call — which most do — the argument's effects land at assembly. A language whose default is to defer arguments until demanded would behave differently, which is exactly why a pipeline library cannot rely on that and provides explicit steps that take functions. ## The consequences a service actually sees - **An eager side effect at the wrong time.** A read, a write or a network call written in the builder happens at start-up, once, whether or not the pipeline is ever run. - **A value frozen at build time.** A timestamp, an identifier or a configuration snapshot computed in the builder is baked into the description and is the same on every later run. - **A cost paid where nobody is measuring.** Expensive preparation inside the builder shows up in start-up time, not in the run's own timings, and so goes unattributed. - **A failure at the wrong moment.** If the eager call throws, it throws during assembly — the chain is never produced, and the error looks like a start-up fault rather than a pipeline fault. - **A count that surprises.** Build the chain twice and the eager work happens twice, regardless of how many runs follow. ## Reading a builder correctly 1. Split the builder into statements that compute something and steps that receive a body. The first group has already run by the time the chain exists. 2. For each step, ask what it was actually handed. A ready-made source or value means assembly-time work; a function means deferred work. 3. For each deferred body, ask how often the pipeline invokes it — per value or per execution. Those are different frequencies and they produce different bugs. 4. Confirm empirically by counting: an assembly-time line logs once per build, a per-execution body once per run, a per-value body once per element. ## Why the design is this way, not a wart Something must be eager, or the chain itself could never be constructed. Assembly is the phase where the shape of the work is decided — which steps exist, in what order, with what parameters — and it is deliberately cheap and repeatable so that the resulting description can be stored, passed around and run many times. Anything that must be fresh on each run has to be expressed as a body the pipeline invokes, because that is the only thing the pipeline controls the timing of. An engineer who holds that one line — *the pipeline controls when it calls the bodies you gave it, and nothing else* — can predict when any given line of a builder will execute.

  • Which part of a step genuinely waits for the run?
    The function the step was handed. It is created while the chain is being built, but the pipeline decides when to invoke it — once per element for a per-value body, once per execution for a body that produces a source. Nothing outside such a body is under the pipeline's control.
  • How would you establish, in a running service, whether a line executes at assembly or per run?
    Count it. Emit a record with a monotonic counter or a run identifier and compare the count against the number of builds and the number of runs. An assembly-time line appears once per build and never again, however many runs follow.
  • What happens if the eager code in the builder throws?
    The chain is never produced. The failure surfaces wherever assembly happened, typically start-up, and looks like an initialisation fault rather than a pipeline fault, since no run has begun and no step has been invoked.

saying these in an interview costs you the question

  • Believes nothing at all executes until the pipeline is run
  • Thinks the whole builder function is deferred along with its steps
  • Expects a value computed at build time to refresh on each run
  • Calls the early log line a library bug rather than ordinary evaluation
  • Assumes putting a call inside the chain makes it deferred by itself