A suite's wait timeouts were all raised to 90 seconds to stop intermittent failures. What does that hide, and what should the team do instead?
answer
- A timeout is a claim, not a knob
- Three different failures were suppressed together
- Costs land on failing runs and diagnosis
- Measure the tail before arguing a number
- No defensible number means the wrong signal
basics
~20 sA blanket raise withdraws the claim that the system finishes in a known time. It hides performance regressions, wrong readiness signals and real races alike. Measure each wait's actual duration, set per-operation budgets from the tail, and fix the signal.
solid answer
~50 sA timeout is an assertion — "this completes within N" — so raising every one of them across a suite does not fix the failures, it withdraws the assertion. The suppressed failures usually belong to three different classes: a genuine slowdown the suite was accidentally monitoring, a wait watching the wrong readiness signal so an ordering assumption merely holds more often, and a real product race the test had caught. Note the cost is asymmetric: a big timeout costs nothing on a passing run, because the wait returns as soon as its condition holds. It costs on failing runs, where each one now burns the full budget, and it costs diagnosis permanently, because nothing distinguishes slow from broken any more. Instead: record wait durations on green runs, derive per-wait budgets from a measured tail plus stated headroom, check the signal before the number, and escalate a real slowdown as a regression.
go deeper
Know that a wait budget should reflect how long the operation really takes, and that making every budget huge to stop failures removes the suite's ability to tell you something got slower.
Explain the three classes of failure a blanket raise suppresses — a slowdown, a wrong readiness signal, a real race — and why one number cannot be right for operations of very different cost.
Show the investigation: split the failing waits by frequency, collect duration data from green runs, derive per-wait budgets from the tail, and demonstrate that you check what the wait watches before you touch its number.
Own timeouts as service-level assertions with named owners: defaults justified by data, overrides that must cite a measurement, and a review rule that treats a raised budget as a performance regression report rather than a maintenance edit.
### Why a global raise is tempting, and what it actually did Intermittent failures are expensive and demoralising, and raising timeouts makes them stop. That is the whole appeal: it is one change, it is reversible in principle, and the board goes green. The reason it is still the wrong move is that a timeout is not a tuning knob — it is an **assertion about the system's behaviour**. "This completes within fifteen seconds" is a claim someone could be held to. Raising every such claim to ninety seconds does not fix anything; it withdraws the claim across the whole suite. What was withdrawn matters, because the failures that were being suppressed fall into three quite different classes, and the raise treats them identically: - **A genuine slowdown.** Something got slower — a query lost an index, a batch grew, a dependency began throttling. The suite was, accidentally, the only performance monitor in the pipeline, and it has now been muted. The next signal will come from production. - **A wrong readiness signal.** The wait is watching something that does not mean "done" — typically an upstream stage's counter or status. A longer budget makes the ordering assumption hold more often; it does not make it true. The failure rate drops but does not reach zero, and now each occurrence costs ninety seconds. - **A real product race.** The system genuinely can produce the wrong result under interleaving, and the test caught it. Raising the timeout converts a reproducible bug report into a rare mystery. Only the first is even partly a timing problem, and even there the honest response is to record the regression rather than absorb it. ### The cost is asymmetric, and that is the nuance to raise A larger timeout costs **nothing** on a passing run: the wait returns the moment its condition holds. So "it slows the suite down" is not, by itself, the argument — and a candidate who leans on it will be corrected. The costs are these: - **Failing runs get much slower.** In a 27-minute suite, a run with six genuine failures that each now burn ninety seconds instead of fifteen adds about seven and a half minutes — and every one of those minutes is spent producing no new information. - **Diagnosis degrades.** A wait that can never plausibly trip stops discriminating between "slow" and "broken". Everything that goes wrong looks like the same long stall. - **The regression signal is lost permanently.** Nobody notices a stage that drifted from two seconds to forty, because forty is inside budget now. ### What to do instead 1. **Split the population before touching a number.** Identify which waits actually failed and how often. A single wait failing one run in forty-three is a different problem from thirty waits failing occasionally. 2. **Measure the real distribution.** Record every wait's duration on green runs as data. You cannot argue about a timeout without knowing the tail; the mean is not what trips it. 3. **Derive per-wait budgets from that data** — a high percentile plus stated headroom — instead of applying one number to a suite whose operations differ by orders of magnitude. A wait for an in-process state change and a wait for an overnight batch have no business sharing a budget. 4. **Check the signal before the budget.** In most of these investigations the wait turns out to be watching the wrong artefact, and the correct fix makes the timeout question disappear. 5. **Escalate the slowdown as a finding.** If the measured tail genuinely moved, the raise may be justified for now — but it is recorded as a performance regression with an owner, not applied quietly. ### A worked example On a smart-meter reading-feed suite, the raise hid an ordering assumption. Several tests waited for an ingest stage to report a batch of 8,412 readings accepted, then asserted the per-meter daily total produced by a later aggregation stage. Nothing guaranteed aggregation had run. At fifteen seconds the assumption held on all but about one run in forty-three; at ninety seconds it held almost always. Two weeks later the aggregation stage genuinely slowed after a schema change, and the suite said nothing, because the only test that would have complained had been given six times the budget it needed. The fix was to wait for the daily-total record's own window state and to restore a budget of roughly eighteen seconds derived from a 6.4-second tail — at which point the suite began failing immediately on the regression it had been trained to ignore. ### The one-line position to hold Timeouts should be **short enough to be informative and long enough to be fair** — set from measurement, per operation, with the failure message reporting the last observed state. When you cannot find a defensible number, that is evidence the wait is watching the wrong thing, and no number will rescue it.
- Someone objects that longer timeouts cannot be harmful because green runs finish just as fast. How do you answer?They are right about run time and wrong about value. The wait does return as soon as its condition holds, so a passing suite is unaffected. The damage is elsewhere: failing runs now spend the full budget each, producing no information; a timeout that can never plausibly trip no longer discriminates between slow and broken; and a stage that drifts from two seconds to forty raises no alarm because forty is inside budget. The argument is about signal, not speed.
- How would you collect the evidence needed to set defensible budgets?Have the shared wait helper record, for every wait on every run, what it waited for and how long it took, and publish that alongside the run. After a few hundred runs you have a per-wait distribution: a typical value, a tail percentile, and the spread across loaded and cold-start machines. Budgets then come from the tail plus stated headroom, and the same data doubles as an early performance-regression signal when a distribution shifts.
- When is raising a specific timeout the correct decision?When the measured distribution genuinely moved and the new latency is accepted — a larger batch, an added stage, a dependency whose service level changed. Even then it is a recorded decision: the new number cites the measurement, names who accepted the slower behaviour, and is applied to the affected waits only. What is never correct is applying one raised number across an entire suite because the board was red.
saying these in an interview costs you the question
- Treats a timeout as a knob rather than a claim
- Applies one budget to operations of wildly different cost
- Argues long timeouts are harmless because green runs are fast
- Never measures actual wait durations before choosing numbers
- Assumes every suppressed failure was a timing problem
- Raises the budget without recording a slowdown as a regression