A per-record step body reads the wall clock and a random draw — what does that cost the job's repeatability?
answer
- the engine cannot see inside the body
- a value from when, not from what
- a recomputed piece reads a later clock
- call count is not in the program text
- fix it once, or derive it from the record
basics
~20 sThe output becomes a function of when and where the body ran, not of the input. Two executions disagree, and one execution can disagree with itself when a lost piece is recomputed and stamped with later values than its first attempt produced.
solid answer
~50 sA body handed to the engine as ordinary code is an **opaque step**: the plan rewriter — the component that edits the declared graph into a cheaper equivalent — cannot look inside it, so it knows only that the function runs, never that it reads a clock or an entropy source. That has three consequences. Two executions over the same input produce different values. A single execution can become internally inconsistent, because a piece lost to a dead worker is recomputed and the second attempt reads the clock later. And how many times the body runs is not visible in the program text: a slow piece may be run a second time in parallel with the first, and a reused branch may be recomputed rather than held. The repair is to make the value a function of the record — fix one timestamp in the coordinating process and pass it in, or derive a seed from a stable field of the record.
go deeper
Recall that a step body reading the current time or a random number makes the output depend on when it ran, so two executions over the same input will not match.
Explain why the engine cannot help: the body is ordinary code it can only call, so it cannot know the body is unstable, and it never promised a fixed number of calls.
Demonstrate the within-one-run failure — a recomputed piece stamped later than its first attempt — and give the repair as a fixed constant passed in or a seed derived from the record.
Treat purity of step bodies as a platform rule with a review check behind it, and decide where run-level values such as an as-of timestamp are injected so every job gets them the same way.
## The body the engine cannot see An engine reasons about steps it understands. On a **declared-operator surface** — where the author names operations the engine already knows, such as filter, project, group and join — the engine can reorder, narrow, fuse or skip work because it knows what each step means. A body handed over as ordinary code is the opposite: an **opaque step**, which the engine can only call, never read. It has no idea whether the body is a pure function of its argument or a call into the operating system's clock. That is the whole of the problem. A pure body makes the output a function of the input. A body that reads the wall clock, draws from an entropy source, generates a fresh identifier, or looks up a system that is changing underneath it makes the output a function of **when and where it ran**. ## Three different failures, not one 1. **Two executions disagree.** Re-run tomorrow and every timestamp is a day later, every draw is different. This is the failure everyone predicts. 2. **One execution disagrees with itself.** A worker dies; the piece it held is recomputed on another machine; the recomputed records carry later timestamps and different draws than the ones the first attempt produced. Half the output is from 09:14 and half from 09:31, with nothing in the data saying so. 3. **A later reprocessing produces numbers the original run never produced.** Replaying history through the same code stamps it with today's clock. Whether a historical replay is *supposed* to reproduce the original numbers is a separate subject with its own owner; what matters here is that an unstable source inside a body makes the choice for you. The second one is the one candidates miss, and it is the one that shows up in production. ## How many times a body runs is not in the program text A reader counts one call per record. The engine makes no such promise: - a piece lost with its worker is **recomputed**, so its records pass through the body twice in total; - a piece running far slower than its peers (a **straggler**) may be run a second time in parallel on another machine, whichever copy finishes first being kept — so some records pass through the body twice in the same run, deliberately; - a branch used by two later steps may be **recomputed** rather than held, depending on whether the engine holds intermediate results and whether that hold survived; - and in the other direction, an engine that narrows a read to the columns later steps actually use may skip work entirely — engines differ on whether they will drop a derived column nothing reads when the body producing it is opaque, so the number of calls is not something the program text tells you. With a pure body, none of this is observable. With an unstable one, all of it is. ## Making the value a function of the record | unstable source | why it varies | repair | |---|---|---| | wall clock inside the body | reads the moment the piece ran, which differs per attempt | fix one timestamp in the coordinating process and pass it into the step as a constant, or assign it when the record is ingested so it is part of the input | | random draw from ambient entropy | new value on every call and every attempt | derive the seed from a stable field of the record, so the draw is a pure function of that record | | freshly generated identifier | a re-run invents a new identity for the same record | derive the identifier from the record's own content or natural key | | lookup into a system that changes | answers differ by the minute | snapshot the reference data and join against the snapshot, so the run reads a fixed thing | The shape of every repair is the same: move the varying value **out of the body and into the input**, either as a constant fixed once for the whole run or as something computed from the record itself. ## Where engines differ - A **finite job re-run from scratch** reads the clock at the new run's moment — the whole output shifts together, which is at least consistent. - A **continuous record-at-a-time model**, where one fixed graph stays running and each record passes through as it arrives, resumed from a saved picture of its progress (a **checkpoint**), reads the clock at resume time for everything it reprocesses, so the output has a seam in it. - A **repeated-small-batch model**, which runs continuous work as a succession of small finite jobs, reads the clock once per small job, so values step forward in blocks rather than smoothly. None of those is more correct than the others; the point is that a candidate who says "the timestamp will be the time of the run" is describing one of the three. ## The interview answer in one line An opaque body with an unstable source inside it turns the job from a function of its input into a function of its execution — and because the engine cannot see inside, nothing will warn you.
- Why is a per-worker seed, fixed at process start-up, not enough?Because the sequence a worker produces is consumed in the order records reach it, and which records land in which piece — and in what order — is not fixed between executions. The same record can draw a different value on a second run, or on a recomputed attempt, even though the seed was stable. Tie the seed to the record, not to the process.
- Can holding a computed result stop the body running twice?Not reliably. Holding an intermediate result is a best-effort optimisation: if the held copy is lost with its worker, or was never fully retained, the branch is recomputed from its inputs, and the body runs again with a new clock reading. It reduces recomputation; it does not make an unstable body safe.
- What is the cheapest way to give a whole job one consistent as-of time?Read the clock once in the coordinating process, before the graph is handed to the cluster, and pass that value into the steps as an ordinary constant. Every worker, every attempt and every recomputed piece then sees the same instant, and the run's output records which instant it was.
saying these in an interview costs you the question
- Assumes each record passes through the body exactly once.
- Says the engine will detect and warn about a clock read.
- Seeds a generator once per worker and calls it reproducible.
- Thinks holding a computed result guarantees no recomputation.
- Believes only a whole re-run, never a retry, changes the values.
- Treats a freshly generated identifier as stable across runs.