What does a recorded plan buy on one machine, where there is no network round trip to avoid?
answer
- the engine sees the end first
- unreferenced columns cost nothing
- condition applied during the read
- one traversal, no intermediate buffer
basics
~20 sThree local savings: columns the plan never references are never read or decoded, a row condition can be applied during the read instead of after it, and consecutive element-wise steps can collapse into one traversal with no intermediate buffer.
solid answer
~50 sRecording the steps first means the engine sees the end of the chain before it reads the first byte, and on one machine that knowledge is worth three concrete things. First, **columns the recorded steps never reference are never read or decoded** — a wide source costs what the query needs, not what the file holds. Second, **a row condition written several steps in can be applied while the input is read**, so rows that will be discarded are never materialised. Third, **consecutive element-wise steps can be fused into one traversal**: under eager evaluation each operator allocates a full-length buffer, so a chain of four allocates four; a recorded plan can produce the same result with one output and no per-operator temporary. How much of this any given design actually performs varies — deferral creates the opportunity rather than guaranteeing the rewrite.
code
pseudocode · 12 lines# eager: each step runs on its own line and returns a complete object
loaded = read_all(source) # every column materialised
filtered = keep_rows(loaded, amount > 100) # a second full object
withfee = add_column(filtered, fee = amount * 0.02)
answer = mean_of(withfee, "fee") # 4 traversals, 3 intermediates
# plan-then-run: the same four steps recorded first, then triggered once
p = describe(source)
p = keep_rows(p, amount > 100)
p = add_column(p, fee = amount * 0.02)
p = mean_of(p, "fee")
answer = trigger(p) # only now is anything readgo deeper
Recall the headline: if the recorded steps never mention a column, the engine has no reason to read it. That one idea is most of why deferral pays on a single machine.
Explain all three savings and, for the fused chain, name the evaluation model in the sentence — a per-operator temporary under eager evaluation against one output buffer under a recorded plan.
Show that you measure rather than assume: which of the three a given engine actually performs is a property of that engine, and the way to find out is one chain timed and sized against its immediate equivalent.
The angle is commitment. Choosing a deferred entry point for a pipeline buys these opportunities and changes how the team debugs, times and reports every stage of it; that is a standing cost to weigh, not a free win.
Deferral is usually explained with a distributed system in the background: hold the work, ship less of it over the wire. That explanation is incomplete, because the three largest wins are local and have nothing to do with a network. ## The property everything else follows from When steps are recorded rather than run, the engine holds the **whole chain** at the moment it is asked for an answer. It knows the last thing wanted before it opens the first byte of the source. An eager step cannot know that: when it runs, the steps after it have not been written yet as far as it is concerned, so it must produce a complete, general result that any later step could use. Everything below is that one property cashed out. ## 1. Columns the plan never names If the recorded steps reference three of a source's two hundred columns, the other hundred and ninety-seven are dead. The engine can decline to read them and decline to decode them. Decode is often the larger half of that saving: bytes on disk have to be turned into the in-memory representation of the column, and for text columns that is allocation, not just copying. An eager read cannot make this call. It is handed a source and asked for a table, so it produces the table, and the discard happens afterwards — after the bytes were read, decoded and made resident. ## 2. The condition moved into the read A row condition written three steps after the read still selects rows of the source. A planner that holds the whole chain can evaluate it *while* the input is being consumed, so rows that will not survive are never materialised into the in-memory representation at all. Under eager evaluation the read produces every row first, and the condition then allocates a second object holding the survivors — both are resident at the same moment, which is what sets the peak. The saving is proportional to how selective the condition is. A condition that keeps 95% of rows saves almost nothing; one that keeps 2% changes the footprint by more than an order of magnitude. ## 3. Element-wise steps collapsed into one traversal This is the one people most often get backwards, so state the evaluation model in the sentence. | | eager evaluation | recorded plan, fused | |---|---|---| | a chain of 4 element-wise operators | 4 full-length buffers allocated in turn | 1 output buffer | | traversals of the column | 4 | 1 | | peak set by | the widest pair of live buffers | the input column plus one output | Under eager evaluation each operator returns a complete column before the next is applied, so a chain of four allocates four full-length temporaries and peak grows with the length of the chain. Where the same expression is recorded, it can be executed as one traversal that computes the final value per element, costing one output and no per-operator temporary. Both costs come from the same author-visible expression, which is exactly why the model has to be named rather than assumed. ## What it does not buy - It does not make an individual operation faster. Adding two columns costs what it costs; deferral changes how many times you do it and how much is resident while you do. - It does not reduce peak footprint by a fixed factor. How much a step adds to peak depends on what that step touched and what the design does with the columns it did not touch: where untouched column buffers are shared with the result, peak grows only by the columns actually rewritten; where every step copies every column, it grows by the whole table. - It does not help a chain whose very first step needs every column and every row. ## What varies between designs Some deferred designs rewrite the recorded chain aggressively — reordering what they can, fusing element-wise runs, narrowing the read. Others record faithfully and execute close to as written, so the handle defers *when* the work happens without changing *what* it is. Some are able to fuse only within a run of steps whose behaviour they can reason about, and stop at the first step they cannot. So the honest statement is that recording the chain creates these three opportunities; which of them a particular engine takes is a property of that engine, and measuring one chain against its immediate equivalent is how you find out. The interview answer worth giving is the mechanism, not the list: the engine is allowed to make decisions about reading and allocation that an eager step cannot make, because it knows the end of the chain before it starts.
- Which of these three savings survives when the source is an uncompressed text file rather than a columnar one?Fusing element-wise steps and applying the condition during the read both survive — they are about what is materialised, not about the file. Skipping unreferenced columns survives only in part: the bytes of a row-oriented text file still have to be scanned past, but the columns that are never referenced need not be parsed into their in-memory representation, and parsing is usually the expensive half.
- Why does the selectivity of the condition decide how much the second saving is worth?Because the saving is the rows never materialised. A condition that keeps most rows leaves the resident footprint near where it would have been anyway; one that keeps a small fraction means the in-memory table is built at that fraction of the size, and every later step traverses that much less. Before optimising, find out roughly what fraction survives.
- Does a recorded plan reduce peak memory by a predictable factor?No. The change depends on what each step touches and on what the design does with the columns a step did not touch. Where untouched column buffers are shared with the result, peak grows only by the columns actually rewritten; where every step reallocates every column, it grows by roughly the whole table. Measure it rather than quoting a multiple.
Handing a shopper the whole list before they walk in, instead of one item at a time. Knowing the full list, they can skip the aisles nothing on it comes from, pick things up on the way in rather than doubling back, and carry one basket instead of setting down a full one at the end of every aisle. The shop did not change; the route did.
saying these in an interview costs you the question
- Recording the steps only pays off once a network is involved
- Every deferred design fuses consecutive element-wise steps
- A recorded plan makes each individual operation faster
- Deferring the work reliably halves peak memory
- Chained arithmetic always allocates one temporary per operator