skip to content

In an engine that runs no declared step until something demands an answer, what distinguishes a describing call from a demanding one?

level: juniorimportance: must knowfreq 74%

answer

  1. two kinds of call, one runs
  2. describing a step versus asking for an answer
  3. graph recorded first, submitted on demand
  4. counting, sampling, writing: these demand
  5. assembly still resolves names and types

basics

~20 s

A describing call only adds a step to the graph the engine will run and computes nothing. A demanding call asks for an answer or for output to be written, and that demand is what makes the recorded steps execute.

solid answer

~50 s

Calls against a **distributed collection** — the engine's handle on records spread across the cluster — fall into two groups. An **assembly call** (filter, choose columns, group, join, hand over a per-record function) only records a step in the **step graph**: the ordered set of steps the engine derives from the program, each naming the steps whose output it reads. Nothing is read and nothing is computed. A **demanding call** — a row count, a sample brought back to the process that assembled the graph, a write to cluster-visible storage — asks for an answer, and that is what submits the graph and runs it. The split exists so the engine sees the whole shape before committing to physical work. It is not total: where the schema is known, a missing column or a type error is rejected during assembly.

go deeper

for a junior

Sort calls into two buckets: those that only describe a step, and those that ask for an answer or for output. Say clearly that only the second causes any record to be read.

for a middle

Explain what the delay buys — the engine sees the whole declared shape before choosing how to run it — and name the work that does still happen while assembling, such as resolving column names and types against a known schema.

for a senior

Show where the model bends. A continuously running job is started once rather than demanded repeatedly, and a model that writes every intermediate to storage has almost nothing to defer. State the condition your claim depends on.

for a principal

Weigh what the split costs a platform. Deferral buys rewriting and cheap composition, but it moves failure, wall time and spend onto a line that names none of the steps responsible, which shapes every debugging, benchmarking and charge-back conversation the team will have.

## Two kinds of call, and only one of them works An engine of this class hands the author a **distributed collection** — the engine's handle on a set of records spread across the cluster, the thing steps are declared against. The calls made against that handle divide cleanly in two. - An **assembly call** adds a step to the graph and computes nothing. Naming a filter, choosing a subset of columns, grouping by a key, joining two collections, or handing over a function to be called once per record are all assembly calls. - A **demanding call** asks for an answer, or for output to be written: how many rows there are, a small sample brought back to the process that assembled the graph, or a write to cluster-visible storage. What the assembly calls build is the **step graph**: the ordered set of steps the engine derives from the program, each step naming the steps whose output it reads. Until a demand arrives, that graph is the only thing that exists. No worker has been given anything to do and no input record has been read. ## Why an engine would wait Deferral is not laziness in the ordinary sense; it is how the engine buys the right to change your program. Once the whole declared shape is visible, the **plan rewriter** — the component that edits the declared graph into an equivalent, cheaper graph before running it — can act on facts that are simply not available line by line: - it can tell the read to produce only the columns some later step actually uses, because it can see every later step; - it can move a condition down to the point of reading, so rejected rows are never produced at all; - it can collapse several adjacent per-record steps into one pass over each record, so no intermediate collection exists between them. None of those decisions can be made while the third line of a twenty-line program is executing, because the twentieth line has not been written yet. Which rewrites are available, and how far they go, is a separate subject; the point here is only that the delay is what makes any of them possible. ## What does happen during assembly "Nothing runs" is about records, not about the engine being idle. During assembly an engine typically does some or all of the following: 1. **Resolves names and types** against a schema it already knows, so a reference to a column that does not exist can be rejected immediately. 2. **Consults catalog or metadata services** for the shape and location of the inputs, and in some designs lists the files that will be read. 3. **Builds and validates the graph itself**, rejecting combinations of steps that cannot be expressed. What it does not do is read or process a single record. That is the line that matters, and it is the line that explains everything else about this subject: where a failure surfaces, what a stopwatch around those calls measured, and why a branch consumed twice is computed twice. ## Where engine models draw the line differently This split is real, but it is not universal, and a candidate who states it as the model of all distributed processing is over-claiming. | engine model | what the author assembles | what starts the work | a demand per answer? | |---|---|---|---| | defers assembly over a finite input | a step graph | a call asking for an answer or writing output | yes — one run per demand | | runs continuous work as a fast succession of small finite jobs | the same step graph | one call starts the continuous query | no — one start, then it runs on | | keeps one fixed graph running, each record passing through as it arrives | the same step graph | one call submits the graph | no — it runs until cancelled | | runs one grouping step at a time, writing each intermediate to storage | a job description | submitting the job | there is almost nothing to defer | So the honest statement is: *in an engine that defers assembly, the graph is inert until a demand arrives.* In a continuously running job the deferral happens once — the graph is built and nothing runs while you build it — but afterwards there is no repeated per-answer demand at all, and asking "which call ran it" has one answer for the life of the job. ## What the split costs the author Three consequences follow only from the delay, and an interviewer usually walks straight into one of them: - a failure caused by one declared step surfaces at the demand, which names a different line entirely; - a clock placed around the assembly calls measures graph construction and reports a duration that describes no work; - a branch consumed by two separate demands is computed twice, because a declared step holds no result of its own. ## Answering it in an interview Say the two categories, give one concrete example of each, and state the boundary precisely: assembly records steps and reads no records; the demand runs the graph. Then add the caveat — that some validation still happens while you assemble, and that a continuously running job is started once rather than demanded repeatedly. The caveat is what separates a candidate who has read the phrase from one who has used more than one engine.

  • Does anything at all happen while the steps are only being declared?
    Usually some metadata work: resolving column names and types against a schema the engine already knows, consulting a catalog, and in some designs listing the input files. No records are read or processed. That is exactly why a reference to a missing column can fail immediately while a malformed value in row nine million cannot.
  • Is the same split present in an engine that keeps one graph running?
    A form of it. The author builds the graph and nothing runs while it is built, but a single call submits it and it then runs until it is cancelled or fails, so there is no repeated per-answer demand. In a model that writes every intermediate result to storage before the next step begins, the unit submitted is the whole job and there is little left to defer.
  • Why is writing output counted as a demand when it returns nothing to the author?
    Because it asks for the result to exist somewhere. The test is not whether a value comes back to the process that assembled the graph, but whether an answer is required: a write requires every record to be produced, so the graph must run. A call that returns no value can still be the most expensive line in the program.

Writing a shopping list costs nothing and touches no shelves, however long the list gets. The trip to the shop is what turns the list into groceries — and it is the trip, not the writing, that can fail, cost money and take an hour.

saying these in an interview costs you the question

  • Thinks each declared step runs at the moment its line executes
  • Believes the first step starts reading input as soon as it is declared
  • Says absolutely nothing happens during assembly, including name and type checks
  • Assumes every engine model has a separate per-answer demanding call
  • Treats writing output as not a demand because it returns no value
  • Cannot name a single call that actually forces the graph to run