Each of three steps on a large table returns instantly, then one later call runs for four minutes — why?
answer
- nothing ran on the first three
- steps recorded, not executed
- one call carries the whole chain
- time the trigger, not the step
basics
~20 sPlan-then-run execution: the three steps only recorded a description of the work, so they returned immediately. The later call is the trigger, where the whole recorded plan finally executes and every byte of work is paid at once.
solid answer
~40 sThe three steps did not compute anything. Under **plan-then-run execution** the library records each step as a description — read this source, keep these rows, add this derived column — and hands back a handle to the recorded steps rather than data. Nothing has been read yet, and appending a node costs the same whether the source holds a thousand rows or a billion. The fourth call is **the trigger**: the point at which the recorded steps are handed to the engine, the input is finally opened, and an answer appears. All four minutes belong to the whole chain, not to that one call. The practical consequence is that timing an individual step tells you nothing about the data, and the profile you actually need is of the trigger.
go deeper
Recall that some libraries record a step instead of running it, and that the work then happens at one later call. A step returning instantly on a huge input is the tell.
Explain what the handle actually holds — a description of steps, not rows — and why appending to it costs the same at any input size. Be able to say where the wall-clock time really belongs.
Show how you would profile a chain whose per-line timings are meaningless, and say out loud that each prefix you time re-reads the source, so the measurement itself is not free.
Frame it as a team-visible property: deferred entry points change what a timing, a stack trace and a log line mean, so the conventions for measuring and reporting pipeline cost have to be set alongside the choice.
The timing you observed is the clearest symptom of an execution model, and naming that model is most of the answer. ## Two models behind the same-looking code Under **eager evaluation**, every step runs on the line it is written. "Keep the rows where the amount exceeds 100" opens the input, evaluates the condition, allocates an object holding the survivors, and hands it back. The next step begins from that object. Time is spent where it is written, so a profiler's per-line numbers mean what they look like they mean. Under **plan-then-run execution** — deferred evaluation, where the steps are recorded rather than executed as they are written — those same three lines touch no data at all. Each appends a node to a description the library holds: *from this source; keep the rows matching this condition; add this derived column*. What comes back is a handle to the recorded steps, not a table of values. Appending a node is arithmetic over a handful of small objects, so it returns in microseconds whether the source holds a thousand rows or a billion. That is why your first three calls looked free: they were free, because they did nothing to the data. The fourth call is **the trigger** — the point at which the recorded steps are handed to the engine and an answer is demanded. Only then is the input opened. The four minutes are the cost of the whole chain, collected on one line. ## Reading the timing correctly | | eager | plan-then-run | |---|---|---| | when a step's work happens | on its own line | at the trigger | | what a step returns | a materialised object | a handle to recorded steps | | where wall-clock time appears | spread across the lines | concentrated at the trigger | | what a per-step timing means | the step's real cost | the cost of appending a node | | where a failure is reported | at the step that was wrong | at the trigger | The practical consequence: timing individual steps in a deferred chain measures nothing about the data. If you need to know which part of the work is expensive, use whatever the design offers for inspecting or instrumenting the recorded steps, or trigger progressively longer prefixes of the chain and diff the times — accepting that each prefix re-reads the input, because the earlier trigger retained nothing. ## Why a design would wait at all Deferral is not delay for its own sake. Holding the steps until an answer is demanded means the engine sees the end of the chain before it reads the first byte, and can use that knowledge in ways a step running on its own line cannot: - columns the recorded steps never reference need never be read or decoded; - a row condition written three steps in can be applied while the input is being read, so fewer rows are ever materialised; - consecutive element-wise steps can be collapsed into a single traversal, allocating one output rather than one full-length buffer per operator. None of that needs a network or a second machine. It follows from knowing what is wanted before starting. ## What the wait costs Two prices come with the same property, and both are visible in the timing you just saw. 1. **The failure arrives at the trigger, not at the step that was wrong.** The step never ran where it was written, so an error is reported at the line that executed the plan. How much that hurts depends on the design: some planners resolve the schema while the plan is being built and reject an unknown column immediately, deferring the work without deferring the check. 2. **A plan is a description, not a result.** Triggering the same handle again replays the chain from the source unless something was explicitly retained, so two answers drawn from one chain can cost two full runs. ## What varies between designs Not every library that defers does the same amount with the delay. Some record the steps and then execute them close to as written; others rewrite the chain substantially before running it. Some resolve column names and types at build time; others discover them at the trigger. Some offer an immediate and a deferred entry point over the same data model, so the identical expression has both timing profiles depending on which one you opened with. The answer an interviewer wants names the model first — *these three returned instantly because they were recorded, not run* — and then says which payoffs and which prices apply on the design in front of you, rather than asserting one tool's behaviour as the rule. The diagnostic itself is simple. If three heavy-looking steps return instantly and one later call carries all the time, you are in a plan-then-run model. If the time spreads across the lines roughly in proportion to the work each one describes, you are in an eager one.
- How would you find out which part of a deferred chain is actually expensive?Use whatever the design exposes for inspecting or instrumenting the recorded steps first, since that costs nothing. Failing that, trigger progressively longer prefixes of the chain and diff the elapsed times — remembering that each prefix re-reads the input from the source, so the measurements are cumulative rather than independent, and the total cost of the exercise is several full runs.
- If the steps returned instantly, does that mean the deferred version is doing less total work?Not by itself. Instant returns only mean nothing had run yet. The chain may genuinely do less work — unreferenced columns never read, fewer rows materialised, fewer intermediate buffers — but that comes from what the engine does with the whole chain at the trigger, not from the steps returning quickly. Measure the trigger against an immediate run to know.
- What tells you, without reading any documentation, which model a given entry point uses?Point it at a large source and time a step that must touch every row. Under an immediate model that step's own line takes proportional time. Under plan-then-run it returns in microseconds and the time appears later at the trigger. A second tell: ask for the row count or print the object and watch whether work starts.
saying these in an interview costs you the question
- The last call must be doing the expensive work itself
- The first three steps were fast because the data was cached
- Each of the three steps already produced a table in memory
- Per-step timings from a deferred chain are a usable profile
- The four minutes prove deferring is slower than running each step