Why is a job's parse error reported at the call that writes output rather than at the parse step, and how do you localise it?
answer
- the failing line is not the guilty line
- no record met the step until demanded
- the error travels back from a worker
- shape errors early, value errors late
- bisect by demanding shorter prefixes
basics
~20 sBecause the parse step only recorded itself; no record reached it until the write demanded an answer, so the failure surfaces there, carried back from a worker. Localise it by demanding answers from shorter prefixes of the graph over a small sample.
solid answer
~50 sIn an engine that records declared steps and runs none of them until a demand arrives, the parse step never executed when its line executed. The write is the first call that requires records to exist, so it is the call that runs the graph — and it is the call whose frame sits at the top of the failure. The body itself ran on a worker, so the message usually carries two contexts: the demand, locally, and the failing body, remotely. Mistakes visible in the declared shape, such as a missing column or an illegal type, are still rejected during assembly; only failures that depend on a record's value can wait this long. To localise, cut the input down and demand an answer after each prefix of the graph until the failure appears, then inspect the record that breaks the body.
go deeper
Recall that the failure is reported where the answer was demanded, not where the step was written, and that this is expected rather than a bug in the engine.
Explain the split between mistakes provable from the declared shape, which fail during assembly, and mistakes that depend on a record's value, which cannot surface until the run. Then describe how you would narrow it down.
Demonstrate a routine that works on production-sized input: shrink, bisect by demanding shorter prefixes, read the frames carried back from the worker, capture the offending records to a side output rather than crashing on them.
Consider where the team should catch this class of failure at all: a contract at the edge of the platform, a bad-record budget with counters and alerts, and a policy on whether a pipeline should ever fail whole on one unparseable value.
## Why the reported line is not the guilty line When an engine records declared steps and runs none of them until something demands an answer, the program text and the execution order come apart. Writing the parse step created an entry in the **step graph** — the ordered set of steps the engine derives from the program, each naming the steps whose output it reads. It did not parse anything. The write is a **demanding call**: it requires every output record to exist, so it is the call that submits the graph, and therefore the call that is on the stack when a worker somewhere fails on a record. The result is a failure whose top frame names a line that is, by construction, innocent. Nothing was lost — the engine did not swallow the cause — but the line number you are shown identifies the *demand*, not the declared step that produced the bad value. ## Two classes of mistake, two moments | mistake | when it surfaces | why | |---|---|---| | a column that the schema does not contain | during assembly | visible in the declared shape; no record needed | | an illegal combination of steps, or an unparseable query string | during assembly | the graph itself cannot be built | | a type that cannot be coerced, where the schema is known | during assembly | the declared shape is enough to prove it | | a value that will not parse in row nine million | at the demand | depends on a record, and no record has been read | | a per-record body dividing by zero on rare input | at the demand | the body runs only when records flow through it | | an input that disappeared after the graph was built | usually at the demand | discovered when the read is actually attempted | The rule of thumb: anything provable from the declared shape alone can fail early, and an engine that knows the schema will usually make it fail early. Anything that depends on the contents of a record cannot be proved from the declared shape, so it waits for the run. ## Reading the failure you were given A failure from a demand usually has layers, and each layer answers a different question: - the **outermost frames** are local: they name the demanding call and the process that assembled the graph. This tells you *which run* failed, not which step. - the **carried-back frames** come from the worker that executed the body. These name the actual function that threw, and they are the frames worth reading first. - many engines also name the **operator or step in the plan** that was executing. That label is the bridge from the failure back to the line you wrote, and it is usually the fastest route. - the **record or piece identifier**, where one is included, tells you which unit of input to fetch and inspect by hand. ## Localising it without rerunning the world A practical order, cheapest first: 1. **Shrink the input.** Point the same graph at a small subset. If the failure survives, every later experiment costs seconds instead of an hour. 2. **Bisect the graph.** Demand an answer after each prefix — declare the first two steps and count the rows, then the first three, and so on. The prefix that first fails contains the guilty step. This works precisely *because* a demand forces the prefix to run. 3. **Read the carried-back frames and the plan label**, not the local line number, to name the step inside that prefix. 4. **Capture the offending record.** Make the body tolerant for one diagnostic run — route unparseable records to a side output or count them — so the run completes and you can look at what actually arrived. 5. **Decide whether the record or the body is wrong**, and fix the one that is. A bad record class usually deserves a guard and a counter, not a crash. ## What varies across engine models Do not state this as the behaviour of all distributed processing: - In a model that defers assembly over a finite input, everything above holds and the demand really is the reporting point. - In a model that keeps one fixed graph running, the graph was submitted once and the call that submitted it returned long ago. A record that breaks a body fails a running task, and what you see is the job's own failure reporting and its restart behaviour, not an exception raised at a call site in your program. - In a model that runs one grouping step at a time, writing every intermediate to storage, the unit submitted is the whole job and the failure is attributed to the phase that was running, which is a coarser but more direct pointer. ## Answering it in an interview Lead with the mechanism — the declared step only recorded itself, the demand ran it — then separate the two classes of mistake, then give the bisect-on-a-small-sample routine. The detail that marks experience is knowing that the useful frames are the ones carried back from the worker, and that a continuously running job does not report this way at all.
- Which mistakes still fail immediately, before any demand?Anything provable from the declared shape: a column the schema does not contain, a type that cannot be coerced, an unparseable query string, an illegal combination of steps. Where the engine knows the schema it will reject these while you are still assembling. Anything that depends on a record's contents cannot be proved from the shape alone, so it waits for the run.
- Why does the message often carry a second set of frames from elsewhere?Because the body ran on a worker, not in the process that assembled the graph. The worker's failure is carried back and re-raised at the demanding call, so you get two contexts stacked: where the answer was asked for, locally, and where it actually broke, remotely. The remote frames are the ones that name the guilty function.
- Does adding logging inside the body help?It helps, but the output lands wherever that worker's output goes, not in the process you are watching, and one line per record over a large input is unusable. Prefer counting bad records and routing a bounded sample of them to a side output, then reading those — that survives a large run and gives you the actual values.
- How does this look in a job that keeps running?There is no call left to attach the error to: the submission returned when the graph started. A record that breaks a body fails the running task, and you see it in the job's own failure reporting, with its restart behaviour deciding what happens next. Which is why operators of such jobs guard bodies against bad input rather than relying on a stack trace at a call site.
saying these in an interview costs you the question
- Blames the write step because the trace names it
- Thinks the engine lost or replaced the original stack trace
- Expects every mistake, including bad values, to surface during assembly
- Debugs by rerunning the full job on the full input each time
- Assumes the reported line number identifies the failing declared step
- Ignores the frames carried back from the worker as noise