skip to content

A program reads a hundred-column input, keeps two columns and filters on a date - which two edits does the engine make before running it?

level: juniorimportance: must knowfreq 68%

answer

  1. two edits, both at the bottom
  2. fewer columns produced by the read
  3. the condition moves down to the read
  4. only where operations are named, not coded

basics

~20 s

Two: read-time column narrowing, which tells the read to produce only the two columns later steps use, and filter at the read, which moves the date condition down so rejected rows are never produced. Both shrink what enters the job.

solid answer

~50 s

Before running anything, the engine holds the program as a step graph - the ordered set of steps it derives from the program, each naming the steps whose output it reads. A plan rewriter, the component that edits that graph into an equivalent but cheaper one, makes two edits at the bottom of it. Read-time column narrowing walks the graph to see which fields any later step references and tells the read to produce only those; the other ninety-eight columns never enter the job. Filter at the read moves the date condition down to the point of reading, so rows it rejects are never produced at all. Both are only available where the author named operations the engine understands. If the steps are author-supplied per-record bodies, the engine knows only that a function runs and can narrow nothing.

go deeper

for a junior

Recall the two names and what each one does to the read: produce fewer columns, and produce fewer rows. The key idea is that the program you wrote is a description the engine is allowed to edit before running it.

for a middle

Explain the mechanism: the rewriter collects field references up the graph to trim the read, and moves the condition down only when doing so cannot change the answer. Separate an input that evaluates the condition from one where the engine applies it just above the read.

for a senior

Show that you check whether the rewrites actually happened rather than assuming them, and that you know what disables them - a body the rewriter cannot read, or a surface where operations are not named at all. Restructure the program so the narrowing survives.

for a principal

The angle here is which authoring surface a platform standardises on. Naming operations buys the engine the sight it needs for these edits; handing over code buys expressiveness and gives that sight up. Decide which default your teams get, and what the exception costs.

## The program is a description, not the execution On a **declared-operator surface** - a surface where the author names operations the engine already understands, such as filter, project, group and join, rather than handing over ordinary code - a call computes nothing. It adds a step to the **step graph**: the ordered set of steps the engine derives from the program, each step naming the steps whose output it reads. Something else, a **demanding call** that asks for an answer or for output to be written, is what makes the graph run. That gap is what the **plan rewriter** exists for: the engine component that edits the declared graph into an equivalent, cheaper graph before running it. *Equivalent* is the binding constraint - the rewritten graph must produce the same answer as the one the author wrote. Two of its edits sit at the very bottom of the graph, at the read, and they are the two an interviewer expects you to name first. ## Read-time column narrowing The rewrite that tells the read to produce only the columns some later step actually uses. The mechanism is a walk up the graph, collecting the field references of every step - the two columns kept, plus any column named in a condition, a grouping key or a join key - and then trimming the read's output list to that set. In the hundred-column example, ninety-eight columns are simply not produced: not decoded into records, not carried between steps, not written out if a later step needs to move records between workers. Two things are worth keeping straight: - **It is driven by use, not by the author's select call.** A column the author never selected but a later condition mentions survives the narrowing; a column the author selected and then dropped again does not. - **How much it saves below the engine depends on the input's layout.** Some layouts can produce a subset of columns without touching the rest; others must parse each record whole and hand back only the requested fields. That is a storage-layer property and belongs to the storage owner - the plan-level fact is only that the read was told what to produce. ## Filter at the read The rewrite that moves a condition down to the point of reading, so the rows it rejects are never produced at all. Two outcomes hide under the one name, and a good answer separates them: 1. **The input can evaluate the condition itself.** The condition is handed to the read, and rejected rows never become records. 2. **The input cannot.** The engine instead applies the condition immediately above the read. Rows are still produced, then dropped at once. Everything above the read sees fewer records, but the read itself costs the same. Either way, the author's program is unchanged and the answer is unchanged; only the position of the work moved. Whether an input can also skip whole files or blocks on stored metadata before reading them is a property of the table and storage layers, not of this rewrite - but the condition has to reach the read before any of that can happen at all. ## What each edit changes | edit | what the read produces afterwards | what it saves | what stops it | |---|---|---|---| | read-time column narrowing | only fields some later step references | decoding, memory per record, bytes moved between workers | a step whose body the rewriter cannot read, so field use is unknown | | filter at the read | fewer rows, or the same rows dropped one step later | per-record work in every step above | an input that cannot take the condition, or a condition the rewriter cannot prove safe to move | ## What varies between engines This is the part candidates skip. **Only what is declared in named operators can be rewritten.** Where the author hands over a **per-record function surface** - a surface where the engine receives a function to call once per record and knows only that it runs, never what it does - there is nothing for the rewriter to analyse, and neither edit happens. And the oldest model in this family, a **two-phase disk-handoff model** that runs one grouping step at a time and writes every intermediate result to disk before the next begins, rewrites essentially nothing: the author's written order is the executed order, and narrowing and early filtering are things the author does by hand. One distinction that is *not* real: writing the same declared job as a query string rather than as chained calls. Both are spellings of the same declared-operator surface, both reach the same rewriter, and both get the same two edits.

  • What does filter at the read achieve when the input cannot evaluate the condition itself?
    The engine applies the condition immediately above the read instead. Rows are produced and then dropped at once, so the read costs the same but every step above it handles fewer records - less per-record work, less memory, and fewer bytes if a later step has to move records between workers. The answer is identical either way; only the position of the work changed.
  • The author already selected two columns, so what is left for read-time column narrowing to do?
    The select declares the output shape, not the read's. Without the rewrite, the read produces whole records and the narrowing happens one step later, after every column has been decoded and carried. The rewrite pushes that decision into the read itself. It is also driven by use rather than by the select: a column named only inside a condition or a grouping key survives, and a column selected then dropped does not.
  • Does writing the job as a query string instead of chained calls change which rewrites happen?
    No. Both are spellings of the same declared-operator surface and reach the same rewriter, so they get the same edits. The distinction that does matter is a different one: whether the step is a named operation the engine understands at all, or an author-supplied body it can only call. That choice, not the spelling, decides how much the engine can change.

saying these in an interview costs you the question

  • Says the engine always reads every column and discards the rest afterwards
  • Believes the rewrite can change which rows or values the job produces
  • Claims these edits happen however the job is written, including per-record bodies
  • Confuses narrowing columns with reading fewer rows
  • Assumes every input can evaluate a condition handed down to it