skip to content

A deferred pipeline fails at the trigger with an error about a step written twenty lines earlier — why there, and what would have caught it sooner?

level: seniorimportance: should knowfreq 52%

answer

  1. the step never ran there
  2. the trigger is where it executed
  3. deferring work is not deferring checks
  4. schema resolved at build, or not
  5. reproduce on a tiny prefix first

basics

~20 s

The step never ran where it was written — it ran at the trigger, so the report points there. How avoidable that is depends on the design: some planners resolve the schema while building and reject an unknown column immediately.

solid answer

~60 s

The line that failed is the line that executed the plan, because that is where the bad step finally ran. But deferring the *work* and deferring the *checks* are two different decisions, and conflating them is the usual mistake. Errors split into two classes. **Schema-resolvable** problems — a column name that does not exist, an operation invalid for a column's type, a mismatch between two chained steps — can in principle be caught while the plan is built, and some planners do exactly that: they resolve the schema step by step, reject the bad step at once, and still defer every byte of work. **Data-dependent** problems — a value out of range, a division by a zero that appears only in row 4 million — cannot be known before the data is read, and will always surface at the trigger. What catches more, sooner: ask the planner to describe the resulting columns and types before triggering, trigger the chain on a small prefix of the input as a smoke test, and split a long chain so each segment is checked where it was written.

go deeper

for a junior

Recall that the failing step ran at the trigger, not on the line where it was written, so the reported location is about execution rather than authorship.

for a middle

Explain the split between deferring work and deferring checks, and give one example each of a schema-resolvable failure and a data-dependent one that no planner could catch early.

for a senior

Show the diagnostic route you would actually take: ask the planner to describe its output, reproduce on a small prefix, bisect against that prefix rather than the full input, and assert expected shape at segment boundaries.

for a principal

The tradeoff to own is where the chain gets cut. Each boundary buys error locality and pays in retained bytes and lost cross-step rewriting, so the placement rule belongs in the team's pipeline conventions rather than to whoever is debugging that day.

This is the first price of recording steps instead of running them, and the answer an interviewer is listening for separates two things that get bundled together. ## Why the report points at the trigger When the bad step was written, it appended a node to a description. Appending succeeded. The step itself — the thing that could not be done — only executed when the whole chain was handed to the engine, at the trigger. A failure is reported where execution was, which is the trigger, and a stack trace from inside the engine describes the engine's own machinery rather than your twenty-lines-earlier line. Good designs attach the originating step to the message; weaker ones leave you to work it out. ## Two different things are being deferred | error class | example | knowable before the data is read? | where it can surface | |---|---|---|---| | schema-resolvable | a column name the source does not have | yes | plan build, if the design resolves schemas | | schema-resolvable | arithmetic on a column whose type forbids it | yes | plan build, if the design resolves schemas | | schema-resolvable | a step consuming a column an earlier step dropped | yes | plan build, if the design resolves schemas | | data-dependent | a value outside the range a later step assumes | no | the trigger, unavoidably | | data-dependent | a group key whose cardinality exhausts memory | no | the trigger, unavoidably | | environmental | the source moved or was truncated | no | the trigger, unavoidably | The blanket claim — *with deferral nothing is evaluated until you ask, so a mistake only surfaces at the end* — is too strong. A planner can resolve the schema as the plan is built: it knows the source's column names and types, it applies each recorded step to that schema, and it can reject an unknown column or an invalid operation at the moment that step is appended, while still deferring every byte of real work. Designs differ on whether they do this. Some resolve eagerly and fail fast; some resolve lazily and everything lands at the trigger; some resolve partially, checking names but not every type interaction. Find out which one you are on before you build a debugging strategy around it. ## The chain does not always reach the trigger intact There is a second reason a failure can appear before the final trigger, and it surprises people. Some operations cannot be described without being partly run: a step that needs the row count, one that needs the rows in sorted order, one that needs the set of distinct labels to decide the shape of its own output. When such a step is appended, a planner may have to execute the chain up to that point right there. The practical effects are that part of the work happens on a line you did not think of as the trigger, and a failure can be reported at that line rather than at the end. ## Narrowing it down when the design will not help 1. **Ask the planner to describe its output.** Most deferred designs can report the columns and types the chain would produce without running it. If that call fails, the failure is schema-resolvable and it tells you which step broke the schema — often naming the step directly. 2. **Trigger on a tiny prefix of the input.** Running the identical chain over a few thousand rows exercises every step for a fraction of the cost and reproduces every schema-resolvable failure and most type failures. It will not reproduce a value that only appears at row 4 million. 3. **Bisect the chain.** Trigger the first half, then three quarters, and so on. Each trigger re-reads the source, so do this against the small prefix rather than the full input. 4. **Split the pipeline at deliberate boundaries.** A long chain triggered once has one place for errors to land. Three segments, each triggered and checked, have three. The cost is the retained intermediates, which is a real footprint cost and the reason not to do it everywhere. ## What to build in Assert the shape at segment boundaries — the columns you expect to exist and the types you expect them to carry — rather than relying on the engine to complain in a way you can read. An assertion written where the step was written reports at that line, which is precisely the property deferral took away, and it costs nothing at plan-build time if the design can answer the schema question without running. The summary worth saying out loud: the work is deferred by construction; whether the *mistake* is deferred too is a property of the planner, not of deferral itself. Name the class of error, say whether the design resolves schemas at build time, and then reach for the small-prefix trigger — it answers most of these in under a minute.

  • Which failures will surface at the trigger no matter how eagerly the planner resolves schemas?
    Anything that depends on the values rather than the shape: a number outside an assumed range, a division by a zero that appears in one row, a grouping key whose cardinality exhausts memory, a source that moved or was truncated. Schema resolution knows column names, types and the shape each step produces; it knows nothing about what is in the rows.
  • Why can part of a deferred chain run before you reach the final trigger?
    Because some steps cannot be described without being partly executed — one needing the row count, one needing the rows sorted, one needing the set of distinct labels to fix the shape of its own output. Appending such a step can force the chain so far to run there and then, which is both a hidden cost and an extra place a failure can be reported.
  • What is the cheapest habit that recovers most of the lost error locality?
    Run the identical chain over a small prefix of the input as a routine step. It exercises every recorded step, reproduces every schema-resolvable failure and most type failures, and costs seconds. Pair it with an explicit assertion of expected columns and types at each segment boundary, so a broken shape is reported on the line that broke it.
  • Is splitting one long chain into several triggered segments always an improvement?
    No. It buys error locality and costs resident bytes: each boundary holds a materialised intermediate, and the engine loses the ability to fuse or narrow across the seam. Put boundaries where the shape genuinely needs checking or where a result is reused, not uniformly along the chain.

saying these in an interview costs you the question

  • Deferring the work makes every mistake undiscoverable until the end
  • The trigger line is the buggy line
  • No deferred design can validate a column name before running
  • Deferred execution is therefore unusable for production pipelines
  • A value out of range should have been caught while building the plan
  • Splitting the chain into segments costs nothing