A job's middle step is ordinary author-written code the engine cannot look inside. What does the engine still know about that step?
answer
- a call, not a computation
- position known, contents unknown
- no field list, no determinism promise
- cost of a call is invisible
- callable, never reasoned about
basics
~20 sThe engine knows only that this step must be called once per record, where the author put it, along with what it was handed and the shape of its result. It cannot see which fields the body reads, whether repeating it is safe, or what it costs.
solid answer
~50 sThe engine derives a **step graph** from the program — the ordered set of steps it will run, each naming the steps whose output it reads. For most steps it holds a meaning it can manipulate: keep these rows, produce these columns, group by this key. An **opaque step** is the exception: its body is ordinary code the **plan rewriter** (the component that edits the declared graph into an equivalent, cheaper graph before running it) cannot look inside. It knows the step's position, its arity and declared result type, that it runs once per record, and whether it was handed the whole record or named fields. It does not know which fields the body reads, whether the body is deterministic, whether it has side effects, or what a call costs. So it can run the step and never reason about it.
go deeper
Recall the one-line version: a step whose body is ordinary code can only be called, once per record, exactly where it was placed. The engine knows the position and the result shape, not what happens inside.
Explain the mechanics: each rewrite needs a proof — which columns are needed, whether calls may be dropped, whether repeating one is safe — and none of those proofs can be derived from an unreadable body.
Show that you act on it. Ask what the body was handed and how often it is called, because both are author-controlled, and say which engine model you are talking about before claiming anything was optimised.
Frame it as a platform choice: how much of your organisation's logic lives as code the engine cannot see, and what declaring determinism or field requirements would buy back across every job at once.
## The step the engine can only call Before any of a distributed job runs, the engine derives a **step graph** from the program — the ordered set of steps it will run, each step naming the steps whose output it reads. For most of those steps the engine holds a *meaning*. "Keep the rows where the amount exceeds ten." "Produce these four columns." "Group by this key and sum that one." A meaning can be manipulated: two conditions can be merged, a select-these-columns step can be pushed toward the input, a grouping can be split into a local part and a global part. An **opaque step** is the exception: a step whose body is ordinary code the plan rewriter cannot look inside, so it can only be called, never reasoned about. The **plan rewriter** is the engine component that edits the declared graph into an equivalent, cheaper graph before running it. Where it meets an opaque body it holds a reference to something callable plus a small contract about its shape, and nothing else. ## What the engine still knows 1. **Position** — where the step sits in the graph as written, and therefore which steps feed it and which read its output. 2. **Arity and declared result** — how many arguments it takes and what type it claims to return. This is what lets the rest of the graph be checked at all. 3. **That it is called per record** — the engine placed the call, so it knows it happens once for each record arriving at that point. 4. **What it was handed** — the whole record, or a named list of fields. This is the most consequential item on the list, because it is the *only* visibility the rewriter has into what data the body could possibly need. 5. **Roughly how many records will reach it**, from whatever it knows or has measured about the input. ## What the engine does not know - Which fields of the record the body actually reads. - Whether the body is **deterministic** — whether two calls on the same input must agree. - Whether it has **side effects**: writing somewhere, calling a remote service, updating a counter. - What it **costs**. A body that returns a constant and a body that makes a remote call per record look identical to the rewriter. - Whether it **fails** on some inputs, which is why running it on extra rows is not obviously harmless. Every rewrite is licensed by a proof. "Delete this step" needs a proof that nothing depends on it, including nothing outside the job. "Run it on fewer records" needs a proof that the calls that no longer happen did not matter. "Read only these columns" needs the set of columns that anything later will touch. None of those proofs is available around an opaque body, so the engine falls back on the only universally safe assumption: call it exactly as written, on everything it was going to be handed. ## Where engines differ | engine model | what an opaque body costs there | |---|---| | an engine with a rewriter over a **declared-operator surface** — a surface where the author names operations the engine already understands, such as filter, project, group and join | the loss is real and local to the step: rewriting that depended on seeing through it stops there | | the **two-phase disk-handoff model** — an engine model that runs one grouping step at a time and writes every intermediate result to disk before the next begins | almost no rewriting is lost, because that model rewrites almost nothing for anybody; the body's own work still costs what it costs | | the **continuous record-at-a-time model** — one fixed graph of steps stays running and each record passes through it as it arrives | the body is usually collapsed into one pass with its neighbours, but the rewriter is exactly as blind to what it does | Several engines also let the author *declare* what the engine cannot infer: that a body is deterministic, that it returns nothing surprising for an absent value, that it needs only these fields. Where such a declaration exists it restores one specific rewrite, and it is a promise the author is trusted on rather than a fact the engine verified. ## What an interviewer is listening for That the candidate does not say "the engine optimises my code". An engine optimises what the author declared in a vocabulary it already understands; anything handed over as code is executed where it was put. A candidate who has internalised that will ask, unprompted, two questions about any such step: what was it handed, and how often is it called. Those are the two things an author can change without touching the body at all.
- Does the engine know whether calling the body twice on the same record is safe?No. Nothing in a plain body announces that it is deterministic and free of side effects, so an engine that might otherwise recompute a branch rather than keep its result has to assume a second call could differ or could do something visible outside the job. Some engines accept an author's declaration to that effect, which is trusted, not checked.
- Does the engine know how expensive the body is?Not from the body. A step that returns a constant and one that makes a remote call per record are the same reference to the rewriter. Some engines let an author attach a cost hint, and some measure elapsed time per step while the job runs, but neither is knowledge derived from the code.
A courier handed a sealed parcel for desk five. They know where it must go and roughly how heavy it is, so they can plan the route around it — but because they cannot open it, they can never decide it is unnecessary, combine it with another parcel, or leave out the part nobody needs.
saying these in an interview costs you the question
- Thinks the engine reads the body and optimises the code inside it
- Assumes the step is skipped whenever nothing later reads its output
- Believes an author-written body is reordered as freely as a named filter
- Says every engine of this class rewrites plans, so it will handle it
- Claims the engine estimates the body's cost from the code