skip to content

Component tests written with React Testing Library time out intermittently in CI, and a teammate proposes raising the global asyncUtilTimeout from one second to ten. How would you evaluate that proposal, and what waiting policy would you set for the suite?

level: principalimportance: should knowfreq 24%

answer

  1. the number reports the problem, it isn't the problem
  2. three failure shapes look identical from outside
  3. what does a bigger budget cost when tests fail
  4. fix the seam, not the deadline

basics

~20 s

Timeouts are a symptom, so classify the failures first: genuinely slow but deterministic, nondeterministic, or waiting on something that will never appear. A global bump only helps the first class and makes every real failure ten times slower to report, so keep the default tight and fix determinism instead.

solid answer

~50 s

I would ask what the failing tests have in common before changing any number. Three causes look identical from the outside: work that is genuinely slow but deterministic — big trees, a CI box several times slower than a laptop; work that is nondeterministic — an unmocked request, a real clock, state leaking between tests; and a test waiting for something that will never appear, where the timeout is the failure mechanism, not the problem. Raising the global default only addresses the first, and it taxes the other two: every genuine failure now costs ten seconds instead of one, multiplied across the suite and any retries. My policy is a tight global default, determinism enforced at the seams — mock at the network boundary, control the clock, reset state between tests — local timeout overrides with a comment where a slow path is real and measured, no arbitrary sleeps, and flake rate tracked as a number the team owns rather than absorbed by longer waits.

go deeper

for a junior

Know that a timeout is how an async utility reports a condition never happening, and that lengthening it does not make anything work — it only changes how long you wait to learn it did not.

for a middle

Be able to separate the three causes behind a timeout — slow but deterministic, nondeterministic, and impossible — and name the seams that make component tests deterministic: network mocks, clock control, state reset between tests.

for a senior

Diagnose from evidence rather than from the setting: repeat runs to see the distribution, isolation versus suite order to expose leaked state, and local timeout overrides with a written reason where a slow path is genuine.

for a principal

Own the suite as a product with a feedback-latency budget. Set the default, decide what evidence justifies changing it, make flake rate a tracked number with owners, and refuse retries as a standing substitute for fixing determinism.

## Why the number is the wrong place to start A timeout is how an async utility reports "the condition I was told to wait for did not happen". It is a reporting mechanism, not a cause. Changing it changes how long the suite takes to tell you something is wrong — and in one narrow case, whether it was wrong at all. Before touching it, get the failures into buckets. **Bucket one: deterministic but slow.** The condition always becomes true, just later than the budget. A large component tree in jsdom, an expensive provider stack, a CI runner with a quarter of the CPU share of a developer laptop. This is the only bucket a timeout increase legitimately serves — and even here, the honest question is why a component test needs a second of wall time at all. **Bucket two: nondeterministic.** Sometimes the condition happens in 80 ms, sometimes never. A request that escaped the mocks and hit a real endpoint. A real clock inside the component. State surviving between tests — a module-level cache, an unrestored fake timer, an unreset store — so the test passes in isolation and fails at position 40 in the file. A ten-second budget converts some of these from red to green, which is worse than leaving them red: the suite now lies, and it lies slowly. **Bucket three: waiting on something impossible.** A renamed label, an unrendered branch, a mock returning a shape the component cannot display. The test is correct to fail. Multiplying its cost by ten buys nothing. So the evaluation is empirical: sample the failing tests, bucket them, and see how much of the pain is bucket one. In my experience it is a minority, which is why the proposal is usually a suppression dressed as a configuration change. ## The cost side of a global increase Slowing failure reporting is a real cost, and it compounds: - Every failing test pays the new budget, and a broken shared component can fail hundreds of them. - If CI retries failing tests, the multiplier applies per attempt. - Longer budgets hide slow regressions. A test that quietly crept from 200 ms to 4 seconds stays green under a ten-second budget, so nobody notices the provider stack that got expensive. - Feedback latency changes behaviour. Once a suite takes long enough, people stop running it locally, and the quality signal degrades further. A tight default is not asceticism; it is a tripwire that makes slowness visible while it is still cheap to fix. ## The policy I would set **Determinism at the seams, first.** Every request mocked at one boundary so no test can reach the network. Timers controlled where they matter. State reset between tests — stores, caches, module registries — so ordering cannot change outcomes. Most bucket-two failures disappear here, and they disappear permanently. **A tight global default, changed only on evidence.** If measurement shows CI machines are uniformly, say, three times slower, a modest global adjustment is defensible and should be recorded with that measurement. Ten times the default without data is not a decision, it is a hope. **Local overrides, with a reason.** A genuinely slow path gets its own `{ timeout }` and a comment saying why. That keeps the cost attached to the test that incurs it and makes it reviewable, whereas a global number is invisible in every diff after the one that changed it. **No arbitrary sleeps, ever.** A fixed sleep is simultaneously too long on fast machines and too short on loaded ones. Ban it in lint and in review; awaiting a condition is always available. **Flake rate as an owned metric.** Track per-test failure rates across runs, quarantine the worst offenders with an owner and a deadline rather than muting them, and treat the number as a team commitment. Without a metric, every individual timeout bump looks locally reasonable and the suite decays by consensus. ## What I would say to the teammate Not no — not yet. Give me the sample: which tests, how often, and what they are waiting for. If most are bucket one and the CI box is measurably slower, we adjust the default by the measured factor and write down why. If they are bucket two, a bigger number will make the suite green and untrustworthy, and we fix the seam instead. The proposal is a reasonable instinct about a real pain; the disagreement is only about which layer the fix belongs in.

  • Is there any case where raising the global default is the right call?
    Yes — when measurement shows the failures are deterministic and the CI environment is uniformly slower than developer machines. Then adjust by the measured factor, not by an order of magnitude, and record the measurement next to the setting so the next person can re-evaluate it instead of guessing.
  • How do you tell a genuinely slow test from a nondeterministic one?
    Run it repeatedly and look at the distribution. A slow-but-deterministic test clusters tightly just above the budget; a nondeterministic one is bimodal — fast most of the time, never otherwise. Also run it in isolation versus in file order: a test that only fails in the suite is carrying leaked state, not a timing problem.
  • Why treat automatic CI retries as a warning sign rather than a fix?
    Retries convert a visible defect into an invisible cost. The flake keeps existing, the suite gets slower by the retry multiplier, and the team loses the signal that would have prompted a real fix. Retries are acceptable as a temporary shield while flaky tests are quarantined with owners and deadlines, not as the standing policy.

saying these in an interview costs you the question

  • Raising the global timeout as the first response to flakes
  • Adding fixed sleeps so tests "have enough time"
  • Enabling blanket retries and calling the suite stable
  • Assuming CI slowness explains failures without measuring
  • Treating a suite-wide green as proof that nothing leaks

context