Two later steps each read the same computed intermediate — what does pinning that result change about the work done?
answer
- count the readers first
- nothing is kept by default
- each reader walks the branch again
- one reader buys you nothing
- kept copy sits in retained-result memory
basics
~20 sPinning tells the engine to keep the computed intermediate after the first reader finishes, so the second reader reads the kept copy instead of re-running every step that produced it. Without a pin, that branch is computed twice.
solid answer
~50 sIn most engines of this class a named intermediate is a description, not a stored thing: records flow through the operators and are dropped as they go. So each later step that asks for the intermediate walks the whole branch above it again, back to the source, and repeats every step in it. Pinning a result tells the engine to keep the computed intermediate — each worker keeps its own share — so the second and later readers read the kept copy instead. The saving is `n - 1` walks of the branch where `n` later steps read it, and exactly nothing where only one does. What "kept" means varies: some engines hold it as ordinary objects of the worker's language, some as a packed byte layout, some overflow it to disk attached to the worker, and on several a pin is a request the engine may drop under memory pressure rather than a guarantee.
go deeper
Recall that a named intermediate is normally recomputed every time a later step reads it, and that pinning keeps the computed copy so the second reader skips that work. Say out loud that one reader means no saving.
Explain where the kept copy lives — the worker's retained-result share, next to the operator working memory the next step needs — and that engines differ over whether a pin is a guarantee or a request they may drop.
Show that you decide by counting readers and by reading the run's own numbers afterwards, and that you check whether the steps after the pin started spilling, because that is where a pin turns into a loss.
The call you own is whether authors may pin at all on a shared platform, what the default form is, and who is accountable for releasing pins — because an unreleased pin costs every other job on the same worker.
## What "a computed result" means here A **computed result** is a named intermediate your program describes — the output of a chain of steps over some input — that later steps read. In most engines of this class the program is a *description* first: you name the steps, and work begins when something demands an answer (a count, a write, a sample pulled back). When that demand arrives, the engine walks back up the **branch above** it — each step pulling from the step before, up to the source — and runs the whole chain. Nothing in that walk is kept by default. Records flow through the operators, are consumed, and are dropped. That is deliberate, not an oversight: the alternative is holding every intermediate of every step, and no worker has room for that. The consequence is the one this question turns on — a **second** demand that reads the same named intermediate walks the same branch again and re-runs every step in it. ## What pinning does **Pinning a result** means telling the engine to keep a computed intermediate so later steps read it instead of recomputing the branch that produced it. The work is distributed, so the copy is too: each **worker** — one operating-system process on one machine that runs part of the job and owns a fixed amount of memory nothing else can borrow — keeps its own share of the result, next to the units of work running inside it. The copy is held in **retained-result memory**: the part of a worker's budget holding results the job was told to keep, as distinct from **operator working memory**, the part operators borrow while sorting, building a grouping table or building the held side of a join and give back when the step ends. What "kept" physically means is not one thing: | where the copy can live | cost to read it | footprint | |---|---|---| | as host-runtime objects — records as ordinary objects of the worker's language | cheapest read | largest; each object carries the runtime's own bookkeeping | | as a packed byte layout the engine lays out and interprets itself | a decode step per read | markedly smaller | | on local scratch disk attached to the worker | a disk read per use | bounded by disk rather than memory | | nowhere, because the engine dropped it under memory pressure | a full recompute of the branch | none | That last row is the one beginners miss. On several engines a pin is a **request**, not a promise: when operator work needs room the engine may perform *the dropping of a pinned result to make room for operator work*, and the next reader quietly recomputes the branch. ## The one thing that justifies a pin Reuse, and only reuse. Count the later steps that actually read the intermediate: 1. **One reader** — the branch runs exactly once either way. The pin buys nothing and still holds memory. 2. **Two or more readers** — the branch runs once instead of `n` times. The pin buys `n - 1` walks of it. 3. **Pinned and never read** — pure loss: memory held for a result nobody asks for. The honest cases are concrete: two outputs written from the same filtered and cleaned set; one prepared set feeding both a join and an aggregate; an iterative computation that reads the same prepared set every round. ## What varies between engines - Engines built around finite input generally expose an explicit pin. A **record-at-a-time continuous runtime** largely does not: it has no finite computed result to hold, and what it remembers between records is long-lived keyed state, which is a different mechanism with a store of its own. - The oldest two-phase disk-to-disk model in this family writes its intermediate out between phases whether or not you asked, so "reuse" there means reading a written output again rather than pinning anything. - An engine that runs continuous work as a rapid succession of small finite jobs sits in between: a pin can hold within one of those small jobs and not across them. ## How you check it worked Settle it from **the run's reported numbers** — whatever the engine reports per unit of work after a run — rather than from arithmetic: - the second reader's steps should show the kept copy being read, not the branch's steps appearing a second time; - the held bytes should be visible somewhere, in a form you can reason about; - spill figures on the steps that follow should not have grown. If they have, the pinned copy is displacing operator working memory, and the pin is a trade rather than a free win. ## One thing a pin is not A pin is not a durability device. A consistent picture written so a crashed job can resume is a **recovery snapshot** — a different mechanism, written to durable shared storage, whose purpose is surviving failure rather than saving recomputation. Treating one as the other is the most common confusion on this subject.
- What happens when a pinned result does not fit in the worker's retained-result memory?It depends on the engine and on the form you asked for. Some keep what fits and recompute the rest of the branch on demand. Some overflow the surplus to local scratch disk attached to the worker, trading a disk read for a recompute. Some hold it in a packed byte layout that is several times smaller than ordinary objects, at a decode cost per read. The failure mode to watch for is paying twice: holding bytes and still recomputing.
- Why does pinning an intermediate that only one later step reads still cost something?Because the memory the copy occupies comes out of the same worker process that operators are drawing on. Retained-result memory and operator working memory are shares of one fixed budget, so a copy nobody reads twice still shrinks the room available for the next sort, grouping table or held join side, and can push that step into writing part of its working set out to local disk.
- Does pinning change the answer the job produces?It should not, and where it does you have found a real problem: a branch whose result differs between two walks. That happens when the branch reads a source that changes under it, or contains something time-dependent or random. In that situation the pin is not an optimisation — it is quietly choosing one of the two answers and freezing it, which is worth noticing before you rely on it.
saying these in an interview costs you the question
- Pinning always makes a job faster, so pin every intermediate
- A pinned copy is free when the cluster has spare capacity
- Pinning a branch that only one step reads still saves that step time
- Once pinned, a result is guaranteed to stay in memory
- Pinning is how a job survives a worker dying