A single unit of a batch job fails identically on every attempt while its neighbours succeed: what does that pattern tell you?
answer
- environment or content
- does it follow the host or the piece
- same failure point every attempt
- retry budget spent on nothing
- progress needs a change, not attempts
basics
~20 sIdentical failure on every attempt, across machines, points at a deterministic cause: the records that unit holds, or the code path they take. Retrying cannot fix it — only changing the input, the code or the unit's tolerance can.
solid answer
~50 sTwo failure classes look the same in a status page and need opposite responses. A **transient** failure comes from the environment the attempt happened to run in — the machine died, a call timed out, a disk filled — so a fresh attempt somewhere else has a real chance, and that is what automatic retry is for. A **deterministic** failure comes from what the attempt carries: the content of the records in that unit, or a code path only those records take. It reproduces exactly, on every machine, so retrying cannot help — nothing that varies between attempts is what caused it. The tell is that the failure travels with the piece rather than with the host, and lands at the same point every time. Progress needs a change: fix the code, make the unit tolerate the value and count it, or change the input. Raising the attempt cap is not on that list.
go deeper
Recall the two causes: the environment the attempt ran in, or the records the attempt carried. The second reproduces on every machine, so retrying alone will never get past it.
Explain the evidence that separates them — whether the failure follows the host or the piece, whether the failure point is identical each time, and whether a re-run of that piece alone reproduces it.
Demonstrate the operational consequence: name what you would change to make progress tonight, what you would count if you made the unit tolerant, and what threshold on that count would page someone.
Argue the policy: who owns a record the pipeline cannot process, how much silent dropping the organisation is willing to buy for an on-time finish, and where that decision should be enforced.
## The two classes, and why the difference is operational - A **transient failure** is caused by the environment the unit happened to run in: the machine disappeared, a network call timed out, a local disk filled, a dependency was briefly unreachable. A fresh attempt elsewhere, or a minute later, has a genuine chance of succeeding. Automatic retry exists for exactly this, and it is why a low attempt cap is a bad idea on a long job. - A **deterministic failure** is caused by something the attempt carries with it: the content of the records in that unit, or a code path only those records take. It reproduces exactly on every attempt and on every machine. Retrying is not merely unlikely to help — it is structurally incapable of helping, because nothing that differs between attempts is what caused the failure. The sharp case of the second class deserves its name: **a record that fails every attempt** — an input whose own content, not the machine, causes the failure. A value the parser never anticipated, a denominator that is zero only in one day's data, a key whose group expands into something the next step cannot hold. ## The evidence that separates them - **Does the failure travel with the machine or with the piece?** Re-runs of the same piece failing on three different hosts, while other pieces on those hosts succeed, is the decisive observation. - **Is the failure point identical?** The same message, at the same place in the code, after the same number of records read, on every attempt. Transient failures are usually ragged; deterministic ones are photocopies. - **Do the unit's neighbours on the same machine survive?** If they do, the machine is not the story. - **Does it still fail hours later?** Time heals a full disk and a flaky dependency. It does not change a value already written into the input. - **Does it fail when that piece alone is re-run in isolation?** This is the cheapest decisive test and is worth naming in the interview, because it separates the two classes without waiting for the job. ## A third case that impersonates the second A piece that is simply too large for the room one worker has will also fail on every attempt and on every machine of the same size — yet the cause is the volume of data, not a value inside it, and the remedies differ (more room per worker, or a smaller piece). How a job handles a piece it cannot fit is its own subject in this tree. Classify by asking whether the failure names **a value** or names **a budget**. ## Why retrying a deterministic failure is worse than wasted 1. It spends the attempt budget that the genuine transient failures elsewhere in the run will need. 2. It costs the unit's full compute on every attempt, and delays everything waiting behind that unit. 3. In a long-running job it is far worse than a wasted unit. Where the recovery unit is a whole connected region rewound to its last recovery point — a durable copy of what the job would otherwise lose, written while it keeps running — the offending record is inside the replayed input, so it is reached again after every rewind. The job does not retry a unit; it restarts, replays and dies at the same place, in a loop whose period is the rewind rather than the unit. ## What actually makes progress | response | what it buys | what it costs | |---|---|---| | Fix the code path the value takes | the record is processed correctly from then on | a deploy, and a rewind or re-run to pick the record up again | | Make the unit tolerate the value: catch, count, continue | the job finishes tonight | records are now being dropped, so the count must be visible and bounded | | Change or repair the input before the job reads it | every downstream consumer benefits | somebody upstream has to own it, and the job waits | Note what is absent from that table: raising the attempt cap. The tolerate-and-count option deserves one caution — a drop rate that climbs run over run means the input has changed shape, not that the job got lucky, so the counter needs a threshold on it. Where rejected records are then sent, and the expectations that would have caught them before the job ever started, are a data-quality subject owned elsewhere. ## What varies between engines - Some engines report which record was in hand when the unit failed; others report only which unit. The first makes this a five-minute diagnosis, the second an afternoon of bisecting the piece. - Where the runtime processes records one at a time as they arrive and holds state between them, there is no per-unit retry to observe: the whole region goes down and comes back, so the same evidence has to be read from the restart history instead of an attempt counter. - In the two-phase disk-to-disk model, where each phase's output is materialised before the next reads it, a deterministic failure in the later phase is reproducible from those durable intermediate files without re-running the earlier phase — a real diagnostic advantage that engines holding intermediates in memory do not have.
- How do you tell a deterministic failure from a piece that is simply too large for one worker?Both fail on every attempt, so the attempt pattern cannot separate them. Ask what the failure names: a value, a field, a parse, a division — or a budget, as in out of room or out of time. The size case also succeeds when the piece is split smaller or the worker given more room, while a bad value survives both. Handling a piece too large to fit is a separate subject in this tree.
- What is the risk of making the unit skip the record that kills it?The job finishes, and the pipeline quietly loses data. It is a legitimate move only with a visible count and a bound on it: a rate that climbs run over run means the input changed shape and the exception has become the rule. Where the skipped records are sent, and the expectations that would have caught them upstream, belong to data quality rather than to recovery.
- Why can one bad record turn a long-running job into a restart loop rather than a retry loop?Because in a job that accumulates state there is no independent unit to re-run: recovery rewinds a whole connected region to its last recovery point and replays the input from the matching read position. The offending record is inside that replayed input, so it is reached again every cycle. Each iteration costs a full rewind and replay instead of one unit's work.
Re-dialling a dropped call is sensible: the line failed, not the number. Re-dialling a number that does not exist reaches the same dead end every time, however many attempts you are allowed — the only fix is to change the number.
saying these in an interview costs you the question
- Says every failure is transient and a bigger attempt cap will get the job through
- Blames the machine when the same unit fails at the same record on three different hosts
- Assumes a deterministic failure must be a code bug, never a data value nobody anticipated
- Thinks swallowing the failing record is free, with nothing to count or alert on
- Treats a transient failure that recurs every single night as needing no investigation