skip to content

A run fails on one step of the job graph and the error names no record. How do you narrow it to a single input piece?

level: seniorimportance: must knowfreq 60%

answer

  1. the error names a step
  2. same unit failing every attempt
  3. map the unit back to input range
  4. one process, one record in hand
  5. a local pass rules out logic only

basics

~20 s

Start from what the failure names: the step and the failed unit of work. If the same unit fails each attempt, map it back to its input range, then run the step's function over that range in one process.

solid answer

~50 s

The error names a **step** — one transformation in the job's graph — because that is the granularity the engine schedules; the record is inside a function it treats as opaque. So work backwards. First ask whether the *same* unit of work failed on every attempt: if it did, the constant is the data it was reading, and you are hunting a record; if a different unit fails each time, the cause is on the machine side, a different investigation. Then map the failing unit back to a range of the input — a file and offset range, or a partition and offset range. How much of that mapping the engine publishes varies, so where it publishes little, carry the source identity on the record and re-run once. Finally read only that range in an ordinary single process and apply the step's function record by record until it throws.

go deeper

for a junior

Understand that the failure names the transformation, not the record, because the engine cannot see inside the function you supplied. Finding the record is work you do, not a lookup.

for a middle

Explain the mapping from a failed unit of work to a range of the input, and why re-reading only that range and calling the function in a loop is the cheapest way to reach one record.

for a senior

Use the engine's own retry as a free experiment to split data causes from machine causes first, and state plainly what a single-process reproduction can never show.

for a principal

Insist the mapping is designed in: carry the source identity and offset on the record so any future failure is one query from its input, instead of a day of reconstruction.

## What the failure already tells you An engine reports at the granularity it schedules. It hands a worker a **unit of work** — the smallest thing it dispatches and the smallest thing it retries by itself — and that unit runs a **step**, one transformation in the job's graph applied to every piece of the input. When the user-supplied function inside that step throws, the engine sees a failed unit on a named step. It has no idea which of the ten million records was in flight, because the function is opaque to it. So the failure hands you three things and withholds the fourth: - the step, always; - the identity of the failing unit, and how many times it was attempted; - usually, the machine the attempts ran on; - **not** the record. How much of the unit-to-input mapping is published varies across engines and across sources. Some name the file and byte range the unit was reading; some name only a unit number you must map yourself. Do not build a method that assumes the generous case. ## The first fork: is it the data or the machine? The engine retries a failed unit inside the same run, normally on some other machine, and that retry is a free experiment. Read it: | What you observe | Most likely | Next move | |---|---|---| | The same unit fails on every attempt, elsewhere too | Something in the records that unit reads | Narrow to the piece, below | | A different unit fails each time | Machine or capacity, not the data | A different investigation entirely | | One unit fails, others on the same machine pass | Still the data, or that unit's size | Compare the piece's size with its neighbours | | Everything on one machine fails | That host | Take it out and re-run | Only the first row belongs to this method. The other three route elsewhere, and saying so out loud is half of what an interviewer is listening for. ## From a failing unit back to a range of input 1. **Take the unit's identity from the run's figures** and, where the engine gives it, the input range it was assigned. 2. **Where it does not, reconstruct the mapping.** The input was cut into pieces by a rule you can apply yourself: files in a listing, byte ranges within a file, partitions of a source with an offset range per unit. Ordering the same way the reader did usually reproduces the assignment. 3. **Where even that is unavailable, instrument for one more run.** Add a counter per reason, or carry the source file name and offset through the record as an extra field. One extra run bought with the right field costs far less than a day of guessing. 4. **Re-read that range alone.** Now you have a piece of input, typically some hundreds of megabytes, sitting where you can read it. ## Getting one record in hand The step's function is ordinary code. Whether the whole program can be executed in one process varies — some engines offer that and some do not — but you never needed the program, only the function: - Read the range, iterate it, call the function per record, and stop at the first throw. Print the record's identifying fields there, where you are allowed to print. - If the piece is too large or the throw is slow to reach, bisect: halve the range and repeat. Four or five halvings turn a million records into a handful. - Keep the guard narrow. Catching every exception to keep going turns a diagnosis into a survey; you want the first offender, intact. ## What a single process cannot show you This is the boundary, and stating it is what separates a senior answer from a confident one. | Reproducible in one process | Not reproducible in one process | |---|---| | A malformed value, a bad cast, a null the code did not expect | Redistribution of records between workers, and its failures | | Arithmetic that overflows on one extreme record | Losing a worker mid-flight and recomputing its work | | A lookup that misses for a specific key | Memory limits of a worker process holding a whole working set | | An encoding the parser mishandles | Ordering and interleaving effects across parallel units | If the whole piece passes locally, that is a real result, not a failed attempt: the record-level logic survives those records, and the cause lies in what a single process cannot recreate. Say that, and redirect — to the piece's size, to the memory the step holds, or to the work that moves between machines.

  • The whole piece passes when you run the step's function over it locally. What does that rule out?
    It rules out the record-level logic on those records: the function survives all of them in isolation. What remains is what one process cannot recreate — redistribution between workers, the loss of a worker mid-run, the memory a worker holds for a whole piece, and interleaving across parallel units. It is a genuine result that redirects the search.
  • The engine fused several transformations into the unit, so the error names the chain. How do you find which one?
    Locally, apply the functions one at a time over the same range and see which throws; the fusion is an execution-time arrangement and does not exist in your loop. Engines differ in whether they report per-transformation or per-fused-chain, so on the cluster the only reliable lever is to split the chain with something the engine will not fuse across, and re-run.
  • Why is re-running the whole job with more workers a poor first move?
    It changes placement, not the data. If the cause is a record, the same unit fails again at higher cost and you have learned nothing; if it is capacity, you may mask the problem without understanding it. The cheap experiment is the retry the engine already performed for you.

saying these in an interview costs you the question

  • Re-running the whole job with more machines and hoping the error moves
  • Assuming the named step identifies the record that caused the failure
  • Concluding a bad record when a different unit fails on every attempt
  • Believing a passing single-process run proves the distributed run correct
  • Adding a per-record print across the entire input to find one record
  • Catching every exception in the local loop instead of stopping at the first