A recorded plan is triggered twice for two different outputs and the second run takes as long as the first — why?
answer
- a recipe, not a result
- the second trigger replays the chain
- nothing was retained in between
- one plan, two outputs, one traversal
basics
~20 sNothing was retained. A recorded plan describes work rather than holding a result, so a second trigger replays the whole chain from the source unless the shared part was explicitly kept — held in memory, or written out and read back.
solid answer
~50 sThe handle you triggered twice holds a description of steps, not the rows they produce. The first trigger executed the chain and gave you an answer; it did not put anything back into the handle. So the second trigger starts again from the source, reads the input again, and re-does every step — which is why the two runs cost the same. Three ways out, depending on the shape of the work: **retain the shared prefix once** by materialising it and deriving both outputs from that object; **express both answers in one plan**, where designs that can produce several outputs from one chain traverse the shared part once; or **write the shared intermediate out** and read it back for each consumer. Each trades the repeated traversal for resident bytes or for disk, so the choice depends on which resource you have.
go deeper
Recall the distinction: the handle holds instructions, not rows. Triggering it a second time runs the instructions a second time, from the source.
Explain why retaining nothing is a defensible default — intermediates are often larger than the answer — and name at least two ways to pay for the shared prefix only once.
Demonstrate the diagnosis: two nearly equal run times for two overlapping answers, confirmed by bytes read from the source per trigger, and then a deliberate choice between resident bytes and a local intermediate file.
The framing is a resource trade the team makes repeatedly: deferral buys traversals with memory, retention buys memory with traversals. Decide which resource your machines are short of and make that the default the pipeline is written against.
This is the second price of plan-then-run execution, and it catches people precisely because the handle *feels* like data. ## A plan is a recipe, not a meal The object you are holding is a description: *from this source, keep these rows, add this derived column, group and total*. Triggering it hands that description to the engine, which reads the input and produces an answer. The answer is returned to you. Nothing is stored back into the description — it is exactly what it was before, a set of instructions. So the second trigger is not a resumption and not a cache hit. It is a fresh execution of the same recipe, from the source, and it costs what the first one cost. If the two triggers differed only in the last step, every step before that last one was performed twice. ## Why a design would behave this way Retaining results by default would be the wrong default for the workload deferral is aimed at. The whole point of recording steps is that the intermediate results are often much larger than the answer, and frequently larger than memory; a design that silently kept every intermediate would defeat the footprint saving that made deferral attractive. So the contract is: you get the answer, and you decide what is worth keeping. ## Three ways out 1. **Retain the shared prefix.** Trigger the chain once up to the point the two outputs diverge, hold that result as an ordinary in-memory object, and derive both outputs from it. Cost: the shared result is now resident for as long as you hold it, which is the footprint you deferred in order to avoid. 2. **Express both outputs in one plan.** Where the design can produce several results from one recorded chain, the shared prefix is traversed once and both branches are computed from it in the same run. Cost: none in traversals, but the shared part still has to be held while the second branch consumes it, and not every design offers this. 3. **Write the shared intermediate out, then read it back.** Persist the result of the shared prefix to a local file and point each consumer at that file. Cost: one write and two reads instead of two full runs — usually the right trade when the shared prefix is expensive to compute and cheap to store, and the only option when it does not fit in memory. | approach | source read | shared work done | extra resident bytes | |---|---|---|---| | trigger twice | twice | twice | none | | retain the shared prefix | once | once | the shared result, held | | one plan, two outputs | once | once | the shared result, while consumed | | write out and read back | once, plus two reads of the intermediate | once | none, but disk used | ## Diagnosing it The symptom is distinctive: two answers that overlap heavily in their derivation, and two run times that are nearly equal rather than one long and one short. Compare the elapsed time of the second trigger against a run of the shared prefix alone. If they are close, the prefix is being redone. Reading the operating system's counters for bytes read from the source over each run makes it unambiguous — the source is being read once per trigger. Watch for the same effect arriving by accident. Asking a deferred handle for a row count, printing the first few rows to check the transform, and then triggering the real answer is three runs of the chain, not one plus two cheap peeks. Interactive work is where this cost is most often paid without noticing, because each individual wait feels tolerable. ## What varies between designs Some designs offer an explicit way to say "keep this result and reuse it for later triggers", which turns the recipe into something closer to a held value. Some can plan several outputs together. Some will reuse work within a single trigger when the same sub-chain appears twice in one plan, but not across triggers. And some retain nothing at all under any circumstances. The behaviour to assume in an interview is the conservative one — a trigger executes the chain and retains nothing — and then to say that particular designs offer a way to opt into retention, which is the thing to look for before hand-rolling a write-out-and-read-back. The underlying judgment is the one that generalises: deferral trades resident bytes for traversals. Retaining results trades the other way. Neither is free, and the shape of the work decides which side you want to be on.
- Why would a design choose not to retain results between triggers by default?Because intermediates in this kind of work are often far larger than the answer and sometimes larger than memory. Silently keeping each one would undo the footprint saving that made deferral worth using, and would make peak memory depend on how many handles happen to be alive. Retention is therefore something you opt into for a result you know is worth the bytes.
- How does this cost show up during interactive exploratory work?Invisibly, one tolerable wait at a time. Asking the handle for a row count, printing a few rows to sanity-check a transform, then triggering the real answer is three complete executions of the chain, not one run plus two cheap peeks. Over an afternoon this is usually the largest single waste, and retaining one small sample fixes most of it.
- When is writing the shared intermediate to a local file the better option than holding it in memory?When the shared result does not fit comfortably in memory, when several separate processes or sessions need it, or when you want it to survive a crash or a restart. The trade is a write plus one read per consumer against a resident copy, so it wins whenever the result is expensive to compute and large relative to what you can hold.
A recipe card and a cooked meal. Handing the card to the kitchen twice gets the dish made twice; the card is unchanged either way. If you want the sauce from the shared first half only once, you either keep a pot of it or write down that you now want two dishes from one session.
saying these in an interview costs you the question
- The library caches the result of every trigger automatically
- Holding the plan handle means the data is already in memory
- The operating system's file cache makes the second run free
- A second trigger reuses work because the recorded steps are identical
- Retain every intermediate, then blame deferral for the memory