A job run finished successfully but retried four hundred times as many units of work as usual — why investigate?
answer
- green says nothing about the path
- retries are invisible unless counted
- compare against a normal run
- concentration: one host, one piece
- the allowance is finite
basics
~20 sA green run hides its retries: the engine re-ran failed units until they passed, so success proves only that the allowance was large enough. The count names a machine, an input piece or a memory limit that is failing today and will exhaust that allowance soon.
solid answer
~50 sSuccess is a verdict on the run, not on how it got there. A *unit of work* is the smallest thing the engine hands a worker process and the smallest thing it retries by itself, and those retries usually do not surface outside the run — the count is the only evidence they happened. A four-hundredfold jump says something is failing consistently and being papered over: one sick machine, a piece of input that kills whatever touches it, memory pressure killing workers, or capacity being reclaimed mid-flight. Three consequences follow. The run paid for all that repeated work in machine-time. The run got slower, which the duration figure will show without explaining. And the allowance for retries is finite, so the same rate on a slightly larger input produces a failed run. Investigate a green run with this signature exactly as you would a red one.
go deeper
Know that the engine re-runs a failed unit of work by itself, and that a run can be reported successful after many such retries. Success is not proof that nothing went wrong.
Explain the candidate causes and their tells: one sick machine, one poisonous input piece, memory pressure, a reclaimed machine. Say why the count is the only evidence retries happened.
Make the prediction: this run consumed most of its retry allowance, so the next slightly larger input fails. Then pair the count with memory headroom and spilled bytes to confirm the cheapest explanation.
Weigh whether a team should be paged on a green run at all. Decide what threshold on wasted work is worth someone's attention, and who carries the cost of the repeated work in a shared pool.
## Why a green run can be a bad run An engine's final verdict is binary and forgiving: it re-runs a failed **unit of work** — the smallest thing it hands to a **worker process** and the smallest thing it retries by itself — until the unit passes or an allowance is exhausted. If every unit eventually passes, the run is reported as successful, and everything that happened underneath is compressed out of that one word. Those retries usually do not surface outside the run at all. A consumer sees correct output. A schedule sees success. The only durable trace is a count, which is why it belongs on the short list of numbers you read on any running job. ## What a jump in the count actually means The plausible causes are few and each has a distinct tell: - **One sick machine.** Failures cluster on one host: bad disk, memory error, a degraded network link. The tell is the host distribution of the failures, if you kept it. - **A piece of input that kills whatever touches it.** The same piece fails on every worker that receives it. The tell is that retries concentrate on one piece rather than spreading. - **Memory pressure.** Units die as workers run out of room, come back, and sometimes pass on a less loaded machine. The tell is memory headroom and bytes spilled to local disk moving alongside the retries. - **Capacity reclaimed mid-flight.** On an interruptible worker — a cheaper machine the provider may reclaim at any moment — losing the machine loses everything it had in progress. The tell is retries correlating with reclaim events rather than with anything in the data. - **A flaky dependency the job calls out to.** Failures track the dependency's own availability. Note that retry-with-backoff against a dependency is a resilience pattern the *program* implements; do not confuse it with the engine re-running a unit. ## The three costs a success hides 1. **Money and machine-time.** Every retried unit re-did work the run had already paid for once. On a run billed for how long machines were held, four hundred extra units is real spend; the allocation of that spend across teams is a separate subject, but the waste is visible here. 2. **Duration.** Retries lengthen the run without changing anything about the program, so the wall-clock figure moves and gives no reason. A candidate who cannot explain a longer run often has not looked at this counter. 3. **The next run.** The allowance is finite. A run that used most of it is one slightly larger input, or one more reclaimed machine, away from a failure that will be reported as a mystery. ## What varies by execution model What gets retried differs across this family of engines, and it changes how much a single retry costs. The granularity itself is a recovery subject; what matters here is that the same count means different amounts of wasted work. | Execution model | What is re-run on failure | What a high count implies | |---|---|---| | a continuous job built from repeated small finite runs | the slice the failure occurred in | wasted work bounded by one slice's period | | a record-at-a-time runtime with key-bound state | the job resumes from a saved snapshot and replays from there | each occurrence can cost everything since the last snapshot | | the two-phase disk-to-disk model | the failed unit, with completed phase output already durable on shared storage | waste bounded by one unit, but a lost worker's local output may force recomputation | Because of that spread, 'four hundred retries' is alarming in a different way on each: on the first it is mostly noise plus cost, on the second it may mean the job has re-read hours of input, on the third it usually points squarely at one machine or one input piece. ## What to do with it Read the count against a normal run for the same job, not against zero — occasional retries are the mechanism working as designed. Then look for concentration: one host, one piece, one time window. Pair it with memory headroom and bytes spilled, because the memory explanation is the commonest and the cheapest to confirm. If the retries produced duplicate effects outside the job — a row written twice, a message sent twice — that is a question about what the outside world saw, which is its own subject and worth naming rather than guessing at. The habit the question is testing is simple and rare: treating the outcome field as one signal among several, rather than as the answer.
- How would you tell a poisonous input piece from one sick machine?Look at where the failures concentrate. A bad machine fails many different pieces, all on one host. A bad piece fails on every host that receives it, and the same piece identifier keeps reappearing in the failure records. Retaining the host and the piece identity alongside the count is what makes this a thirty-second answer instead of a day.
- Why is comparing the count against zero the wrong baseline?Occasional retries are the mechanism working: machines are lost, transient errors happen, and the engine absorbing them is the point. A baseline of zero produces an alarm on healthy runs and trains people to ignore it. The baseline is this job's own normal count, which is why the figure has to be retained per run.
- Do the retries mean the output is wrong?Not by themselves. A re-run unit recomputes the same result from the same input, so the answer is normally unaffected. The risk is external effects: if a unit wrote to a database or sent a message before failing, the retry may repeat it. Whether that is visible outside the job depends on how the writes were made, which is a separate question worth raising explicitly.
saying these in an interview costs you the question
- Treats a successful run as needing no investigation.
- Compares the retry count against zero rather than a normal run.
- Assumes retries are free because the output was correct.
- Confuses the engine re-running a unit with a program's own retry-with-backoff.
- Ignores that the retry allowance is finite and nearly spent.
- Assumes a retry costs the same amount of wasted work on every engine.