Why is an input read twice when one distributed collection feeds two separate calls that each demand an answer, and which engine models never re-read it?
answer
- the handle holds no records
- two demands, two independent runs
- shared prefix walked back to the sources
- bytes read roughly doubles
- not true of every engine model
basics
~20 sBecause a declared step holds no result of its own: each demand runs the graph again from its sources. Models that write every intermediate to storage, and models that keep one graph running and fan records out to both consumers, do not re-read.
solid answer
~50 sIn an engine that records declared steps and runs none of them until a demand arrives, the handle you kept is a description, not a result. Two demands over it are two independent runs of the same graph, each starting at the sources. The symptoms are visible without any special tooling: the run's reported bytes read is roughly double the input, the source step appears twice across the two runs, and total wall time is close to twice the single-branch case. This is not universal. A model that writes each intermediate result to storage before the next step begins has already materialised the shared prefix, so the second consumer reads a file. A model that keeps one fixed graph running evaluates the shared prefix once and sends each record to both downstream branches. Holding a computed result deliberately so it is not recomputed is a separate subject.
go deeper
Recall that holding a chain of declared steps in a variable holds a description, not records, so asking for two answers from it means the work behind it happens twice.
Explain the mechanism: each demand walks back to the sources of that sub-graph and runs it. Be able to say which numbers in the run's own reporting would show it.
Diagnose it from evidence rather than suspicion — bytes read, the source appearing twice, roughly additive wall time, the bill — and state the condition under which your claim holds rather than presenting one engine's behaviour as the rule.
Weigh the platform consequence: a model that retains nothing makes duplicated work cheap to write and invisible until the invoice, while a model that materialises everything pays for writes nobody reads. The choice sets which failure mode your teams will keep hitting.
## A declared step is a description, not a result The variable you are holding after a chain of **assembly calls** — calls that only add a step to the graph and compute nothing — refers to a position in the **step graph**, the ordered set of steps the engine derives from the program, each naming the steps whose output it reads. It does not refer to any records, because none have been produced. So when two separate **demanding calls** — calls that ask for an answer or for output to be written — are made against that same position, the engine has no stored result to serve the second one from. It does what it did the first time: walks back to the sources of that position's sub-graph and runs the whole thing. Reading the input, applying the filter, performing whatever grouping the shared prefix contained — all of it happens twice, in two independent runs. The trap is that the program *looks* like it computed something once and used it twice. In the source text there is one chain and one name. In execution there are two runs that happen to share a shape. ## How you notice it This is a diagnosis question as much as a mechanism question, and the signals are ordinary: - **Bytes read.** The run reporting shows roughly the input's full size for each of the two runs. If the two demands together read twice what the input holds, the shared prefix ran twice. - **The source step appears twice.** Across the two runs, the same read of the same input is listed twice, with similar record counts. - **Wall time is roughly additive.** Two demands over a shared prefix take about twice as long as one, where a reader expecting reuse would predict a small increment for the second. - **Cost.** On storage charged per byte scanned, or on capacity rented by the minute, the duplicate shows up as a line on the bill before it shows up in a conversation. - **The two answers can disagree.** If anything inside the shared prefix is unstable — a value drawn from a clock, a random draw, a source being written to while you read it — the two runs saw different things. The general subject of when two runs of the same job agree is a separate one; here it is only a symptom that tells you the branch really did run twice. ## Where it does not happen Stating "a reused branch is always recomputed" as the model of distributed processing is the error this question is built to catch. It depends entirely on the engine model. | engine model | what a second consumer of the same branch does | why | |---|---|---| | defers assembly, finite job, nothing retained | recomputes the prefix from the sources | the position holds a description, not records | | runs one grouping step at a time, writing every intermediate to storage | reads the intermediate back from storage | the prefix was materialised by construction, not by choice | | keeps one fixed graph running, record at a time | evaluates the prefix once and sends each record down both branches | there is one running graph, not one run per answer | | runs continuous work as a succession of small finite jobs | recomputes within each small job unless something retains it | each small job is a complete run of the same graph | That table is the whole answer to the second half of the question, and it is also the reason the first half must be stated conditionally. In a model where every intermediate hits storage, duplicated computation is not the risk — the unavoidable write of every intermediate is. In a running graph, a branch consumed twice costs an extra downstream path, not an extra read of the source. ## What this question is not about Deliberately pinning a computed result so that downstream steps stop recomputing it — deciding what to pin, where it is held, what it costs in memory or local disk, and when it is a mistake — is a separate subject with its own material. The point here stops at: the recomputation happens, here is why, here is how you would see it, and here is where the model differs. In an interview, naming the remedy in one clause and moving on is the right proportion; explaining it at length usually means the candidate has substituted the remedy for the mechanism. ## Answering it in an interview Three beats. First the mechanism: the handle is a description, each demand runs the sub-graph from its sources, so the shared prefix runs twice. Second the evidence: bytes read roughly doubling and the source appearing twice across the runs, which is how you would prove it rather than assume it. Third the condition: this is the behaviour of a model that defers assembly and retains nothing, and it is *not* the behaviour of a model that writes every intermediate to storage or of one that keeps a single graph running. A candidate who supplies the third beat unprompted is telling you they have used more than one engine in this class.
- How would you prove the branch really ran twice rather than assuming it?Look at what the engine reports for the two runs. Bytes read close to twice the input's size, the same source step listed once per run with similar record counts, and total wall time near twice the single-demand case together make the case. If the source is charged per byte scanned, the bill is a second, independent witness.
- Why is the behaviour different in a model that writes every intermediate to storage?Because the shared prefix was never held in flight: each grouping step's output was written out before the next began. A second consumer of that position reads a file that already exists. Nothing was retained on purpose; materialisation is how that model works, which is also why it pays a write for every intermediate whether or not anything reuses it.
- Does a job that keeps one graph running have this problem at all?Not in this form. There is one running graph and one pass of each record through the shared prefix; a record produced by that prefix is sent down both downstream branches. The cost of the second consumer is the extra downstream work and any extra network hop, not a second read of the source.
- Two demands over the same branch returned different numbers. What does that tell you?That the branch ran twice and saw different input on the two runs — an unstable value inside a body, or a source still being written to while you read it. It is evidence of the recomputation rather than a separate fault. Whether two runs of a job are expected to agree at all, and what has to be pinned before comparing them, is its own subject.
saying these in an interview costs you the question
- Assumes the first demand leaves a result behind for the second
- Says every engine recomputes a branch consumed twice
- Blames the duplicate read on the cluster being misconfigured
- Cannot name any observable evidence that the prefix ran twice
- Claims a running graph re-reads its source once per downstream branch
- Treats differing answers from two demands as an engine defect