A nightly job with a very high per-unit attempt cap now finishes four hours late but green: what is the cap hiding?
answer
- a stopping rule, not a diagnosis
- broken job dressed as a slow job
- green status hides paid-for retries
- attempts per unit, not attempts total
- a rate catches what a count misses
basics
~20 sA high attempt cap converts a broken job into a merely slow one. The extra four hours are retry time: some failure recurs, is paid for on every attempt, and never surfaces because the final status is green.
solid answer
~50 sAn **attempt cap** is the number of times the engine will re-run a failing unit before giving up on the whole run. Set generously — which is reasonable, because transient failures are real and a cap of one would kill long jobs for nothing — it also absorbs failures that recur every time, and the job ends green with hours of wasted compute inside it. The signal is never the final status; it is the attempt counters. Look at whether attempts are scattered across many units and many machines, which is ordinary cluster noise, or concentrated on one unit, which means the failure travels with the data and the retries were never going to succeed. The durable fix is to treat the counters as a first-class output of the run, alarm on concentration rather than on totals, and prefer a cap expressed as a failure rate over a window to a bare count.
go deeper
Recall what the cap is: how many times the engine re-runs a failing piece before giving up on the run. Set high, it hides failures behind a job that merely takes longer.
Explain why the final status is the wrong signal and which counters replace it — attempts per unit, whether retries follow the piece or the host, and time spent in failed attempts.
Show the operating practice: publish attempt counts from every run, alarm on concentration rather than totals, and treat a green run that took twice as long as an incident.
Decide the policy across teams: what a stopping rule should express, who is paged when a run is green but expensive, and how the engine's rule and the scheduling layer's rule are kept from multiplying.
## What the cap is, and why a generous one is defensible An **attempt cap** is the number of times the engine will re-run a failing unit of work before it abandons the whole run. It exists because machines really do die in the middle of long jobs, and a job of eight hours that fails outright on the first lost worker is worse than useless. So the sensible default is generous, and every operator who has been woken by a job that failed on one flaky host raises it at least once. The cost of that generosity is that the cap does not distinguish failure classes. It counts attempts. A failure that will recur on every attempt — because the cause is the content of the records in that unit — consumes the budget just as happily as a genuine transient one, and it consumes it at full price: the unit's whole compute, every time, plus the delay to everything waiting behind it. ## Why "green" is not the signal The final status of a run answers one question: did every unit eventually produce output? It says nothing about how many times the cluster paid for the ones that did. A run that spent four hours retrying and finished is indistinguishable, at the status level, from a run that had a quiet night. That is the whole trap in this question: **the cap converted a job that is broken into a job that is merely very slow**, and slow does not page anyone. ## The numbers that expose it - **Attempts per unit, not attempts in total.** A total of 200 retries spread over 200 units across 90 machines is ordinary cluster noise. Twelve retries on one unit is a different condition entirely. - **Whether the retries follow the machine or the piece.** Concentration on one piece across several hosts means the failure is carried by the data. - **Time spent in failed attempts against time spent in first attempts.** This is the number that maps directly onto the four missing hours, and onto the bill. - **The trend, run over run.** A retry count that has doubled each week is describing an input that is changing shape, and it will cross the cap eventually — usually on the worst possible night. - **Whether any unit is near the cap.** A unit finishing on attempt eleven of twelve succeeded by luck, not by design. ## A count against a rate | stopping rule | catches | misses | |---|---|---| | a fixed number of attempts per unit | the unit that can never succeed, eventually | a job losing a steady trickle of units to a degraded cluster, since no single unit exhausts its count | | failures over a sliding window of time | a cluster going bad and a tight failure loop, quickly | a rare, slow, genuinely transient failure, which it correctly tolerates | | both together | most real conditions | nothing structural, at the price of two numbers to explain | Engines differ in which of these they offer, and several offer both, so the interview-safe formulation is that a stopping rule wants a *rate* as well as a *count* — the count bounds the work spent on one hopeless unit, and the rate detects a run that is failing steadily without any single unit standing out. ## The same knob in a long-running job Where the recovery unit is a whole connected region rewound to its last recovery point rather than a single unit, the cap governs **involuntary whole-job restarts**, and the arithmetic changes shape. Each attempt now costs a rewind and a replay of everything since that point, not one unit's work. A generous cap here does not produce a slow job; it produces a loop that restarts, replays, reaches the same record and dies, burning far more than it processes. This is the setting where a rate-based rule earns its keep, because a restart loop is visible in attempts-per-minute long before it is visible in a count. ## Two layers, two stopping rules The engine's cap is exhausted entirely inside one run. Only when it is does the run fail — and whatever schedules that run may then decide to run it again later, with rules of its own. Those are different graphs and different owners, and the practical warning is that they multiply: a run that retries a hopeless unit a dozen times, re-run three times by the layer above, has paid for that failure thirty-six times before anyone looks. ## The framing that reads as senior A cap is a **stopping rule, not a diagnosis**. It decides when to stop spending, and it deliberately hides why. Anything that must not be hidden has to be surfaced separately: publish attempt counts as an output of every run, alarm on concentration rather than on totals, and treat a green run that took twice as long as a failure that happened to produce output.
- Why is a stopping rule based on failure rate over a window better than a bare attempt count?A per-unit count says nothing about how often units fail across the run. A long job can legitimately absorb hundreds of scattered transient failures without any single unit approaching its count, while a cluster going bad produces exactly that pattern. A rate over a window catches the second without failing the first, which is why many engines offer both and the count alone is the weaker rule.
- What should be alarming about a run that finished green?Attempts concentrated on one unit rather than scattered; a growing ratio of failed-attempt time to first-attempt time; retries that follow the piece across hosts rather than following a host; and any unit that succeeded close to its cap. Each says the run survived by consuming budget that was meant for real transient failures.
- How does the engine's cap interact with something that re-runs the whole job later?They are separate layers with separate rules. The engine's cap is spent inside one run; only when it is exhausted does the run fail, and only then does the scheduling layer above consider running the job again. The costs multiply, so a hopeless unit can be paid for dozens of times before anyone investigates. The upper layer's rules are a different subject.
saying these in an interview costs you the question
- Raises the attempt cap whenever a job fails, without asking what failed
- Reads a green final status as proof that nothing went wrong during the run
- Thinks a high cap is free because failed attempts release their resources afterwards
- Uses one stopping rule for a dying machine and for a record that fails identically
- Argues that a cap of one is the safe setting, so any lost worker kills an eight-hour job