A project and a filter are written after a step whose body the engine cannot read. Why does neither reach the read?
answer
- rewrites are deductions, not magic
- no field list means assume all columns
- a value cannot be tested before it exists
- whole record assembled per call
- narrowing stops at the unreadable body
basics
~20 sBoth rewrites need to know what the unreadable body reads and produces. A condition on a value the body computed cannot be evaluated before the body has run, and a read cannot be narrowed to fields nobody can prove the body ignores, so whole records still arrive.
solid answer
~50 sTwo rewrites are at stake. **Read-time column narrowing** tells the read to produce only the columns some later step actually uses; **filter at the read** moves a condition down to the point of reading so rejected rows are never produced. Both are computed from what later steps demonstrably need. A body handed the whole record tells the engine nothing about which fields it touches, so narrowing has to assume every column may be needed. A condition testing a value the body produced cannot possibly run before it, so every row is read and passed through. The result: the read still produces full records, and where records are held in a compact internal layout, one must be assembled for each call. What varies — if the body is attached to named fields, narrowing can still see those; and a condition on columns the body neither reads nor produces may still be moved earlier on some engines.
go deeper
Recall that an engine can tell a read to produce fewer columns and to reject rows early, and that both depend on the engine knowing what later steps need.
Explain the mechanism both ways: no field list means assume every column, and a condition on a value the body computed cannot be evaluated before the call. Name what still gets narrowed when the body takes named fields.
Reach for the measurement and the fix: the read producing every column on a wide input, the call count against the final row count, and the hand-written narrowing placed before the step.
Consider the aggregate: how much of the platform's read volume is explained by unreadable bodies sitting above the conditions, and whether a convention about what such bodies may be handed is worth enforcing.
## The two rewrites at stake This question is about an engine that has a **plan rewriter** — the component that edits the declared graph of steps into an equivalent, cheaper graph before running it. Two of its rewrites carry most of the benefit on a wide input: - **read-time column narrowing** — the rewrite that tells the read to produce only the columns some later step actually uses; - **filter at the read** — the rewrite that moves a condition down to the point of reading, so the rows it rejects are never produced at all. Both are derived the same way: the rewriter walks the graph and collects what the later steps demonstrably need, then pushes that requirement toward the input. Neither is magic; each is a deduction, and a deduction needs facts. ## Where the deduction fails An **opaque step** is a step whose body is ordinary code the rewriter cannot look inside, so it can only be called, never reasoned about. Put one between the read and the later steps and the deduction runs into two different walls. 1. **The column set becomes unknowable.** If the body was handed the whole record, the rewriter cannot say which fields it reads. Narrowing to the fields named *elsewhere* in the graph would risk producing a record with the body's own inputs missing, so the safe assumption is that every column may be needed. The read produces all of them. 2. **A condition on the body's output cannot move below the body.** The value being tested does not exist until the call has been made. There is no rewriting that changes this; it is arithmetic, not policy. Every row must be read, must reach the body, and must be passed through it before anything can be discarded. The visible consequence is the one an interviewer wants named: the job reads far more than it uses, and the unreadable body is invoked on far more records than the final result contains. ## The record that has to be assembled There is a second, quieter cost when the body takes a whole record rather than named fields. The engine has to hand it a record — a complete one, including the fields the body ignores — because it cannot know which of them will be touched. Where an engine holds records between steps in a compact internal layout, one has to be turned into something the body can read for every call; where an engine already holds ordinary language objects, the call itself is cheaper and the cost sits elsewhere. Which representation an engine chooses, and what that trade costs across the job, belongs to Record Representation; what matters here is only that the choice cannot be avoided by narrowing, because the narrowing is exactly what the opaque body blocked. ## What varies | situation | what still gets narrowed | |---|---| | the body is handed the whole record | nothing above it: the read must produce every column | | the body is handed named fields | the rewriter can narrow the read to those fields plus whatever the rest of the graph needs | | the condition tests a value the body produced | nothing: it cannot be evaluated before the call, on any engine | | the condition tests a column the body neither reads nor produces | some engines move it earlier anyway; others will not, because moving it changes how many times a body with unknown effects is called | | the engine has no rewriter — the model that runs one grouping step at a time, writing every intermediate to disk | the question does not arise: nothing was going to be narrowed for you | That last row is why the neutral statement of this material is "rewrites that depend on knowing what the body reads or produces stop at it", not "nothing crosses an opaque step". ## The shape of a good answer A strong candidate says three things in order. First, the rewrites are deductions from what later steps need. Second, an unreadable body supplies no such facts, and a condition on its output is not merely hard to move but impossible. Third — the part that separates a middle answer from a recitation — the author can supply by hand what the engine could not deduce: apply the conditions and select the fields *before* the step, and hand the body named fields rather than the whole record. That restructuring is this leaf's other half; the point here is why the engine will never do it for you.
- Is a condition written after the body ever moved below it?Sometimes. If it tests a column the body neither reads nor produces, moving it earlier is semantically harmless and some engines do it. Others refuse, because reducing the rows changes how often a body with unknown effects is called. A condition on the body's own output is never moved — the value does not exist yet.
- Why does the read still produce a column that only the body could possibly use?Because nobody can show that it does not use it. Narrowing collects the columns later steps demonstrably need; an unreadable body handed the whole record contributes the whole record to that set. The safe answer and the wasteful answer are the same answer.
- Does writing the same job as a query string instead of chained calls change this?No. Both spellings reach the same rewriter and produce the same graph, so an unreadable body blocks the same deductions either way. What changes the outcome is what the body is handed and where it sits, not which notation the surrounding operators were written in.
saying these in an interview costs you the question
- Says the engine pushes every filter to the read regardless
- Thinks the engine infers the body's columns by parsing it
- Believes a condition on the body's output can run before the body
- Claims writing the job as a query string restores the narrowing
- Asserts nothing at all can be rewritten around an unreadable body