A continuous-integration suite contains asynchronous tests that fail about one run in a hundred, with no obvious pattern. How would you handle that as a team, and what would you do to each individual test?
answer
- red must mean broken — protect the signal
- measure flake rate per test; owner + deadline
- quarantine visibly and time-boxed, never blanket retry
- classify: sleeps, shared state, ordering, scheduler, env, real race
- verify by many runs, not one
basics
~20 sTreat flakiness as a defect with an owner, not noise: detect and track it automatically, quarantine loudly with a deadline instead of blanket retries, then fix each test by removing the timing dependency — signals and virtual clocks instead of sleeps, isolated state, no ordering assumptions — and check whether the flake is actually a product bug.
solid answer
~50 sFirst, the systemic part. A one-percent failure rate across hundreds of tests means most runs are red, so the real damage is loss of trust: people re-run until green and stop reading failures. I would make flakiness measurable — record per-test pass/fail history across runs and flag tests whose result varies on identical code — and give each flaky test an owner and a deadline. Quarantine (excluded from the merge gate, still run and reported) is acceptable as a short-lived, visible state with an expiry; unlimited automatic retries are not, because they permanently hide both flaky tests and genuinely intermittent product bugs. Then, per test: classify the cause. Most fall into wall-clock waits, shared state between tests, dependence on ordering or on a real scheduler, unbounded external dependencies, or a real race in the product. Fix by construction — signal-based waits and quiescence instead of sleeps, injected clock and scheduler, fresh isolated fixtures, seeded randomness. Confirm the fix by running the test many times, not once.
go deeper
Focus on the per-test fixes you can perform: replace sleeps with waits on completion signals, isolate fixtures so tests do not share state, and do not assume ordering.
Add the classification framework (timing, shared state, ordering, environment, real race) and the point that a fix must be verified by repeated runs.
Cover detection and quarantine mechanics, why blanket retries are harmful, and how to make failures self-diagnosing with state and seed dumps.
Frame it as protecting the trustworthiness of the gate: measurable flake budgets, ownership and deadlines, intake control to stop new flakiness, explicit delete-versus-fix calls, and moving inherently probabilistic checks out of the blocking path.
## Why a 1% flake rate is a crisis, not a nuisance Flake rates compound. With 300 independent tests each failing one run in a thousand, roughly a quarter of runs are red. At one in a hundred per test the suite is essentially never green. The consequences are cultural before they are technical: engineers learn that red means "press retry", real regressions get retried away, and the suite stops being a gate. Any strategy has to restore the property that *red means broken*. ## The systemic layer **Measure it.** You cannot manage what you do not record. Store per-test outcomes keyed by commit; a test that both passes and fails on the same code is flaky by definition. Publish a ranked list by flake rate multiplied by how often the test runs — that ordering is where the pain actually is. **Make it owned.** Every flaky test gets an owner (the team owning the code under test) and a due date. Unowned flakiness is nobody's Tuesday. **Quarantine deliberately.** Move the test out of the blocking gate but keep running it and reporting it. Quarantine must be visible, time-boxed and small; an unbounded quarantine list is deletion with extra steps. If the deadline passes, the choice is fix or delete — a permanently muted test is worse than none, because it implies coverage that does not exist. **Ban blanket retries.** Automatic retry-on-failure across the whole suite converts every intermittent product bug — the expensive kind — into silence. A narrowly scoped, logged, alerted retry on a genuinely external dependency is a different thing from a global retry policy. **Fix the intake.** New sleeps and new uses of ambient global clocks/pools get caught at review or by lint, so the population stops growing while you drain it. ## The per-test layer: classify, then fix by construction 1. **Wall-clock dependence.** Sleeps or "finishes within X ms" assertions. Fix: wait on a completion handle, a latch, or a quiescence check with a generous timeout; move time-based behaviour onto an injected clock so the test asserts boundaries exactly instead of racing them. 2. **Shared mutable state between tests.** A static cache, a reused database row, a global counter, a leaked thread from a previous test. Symptom: failures depend on execution order or on parallel workers. Fix: fresh fixtures per test, unique keys/namespaces per test, explicit teardown, no ambient singletons. Deliberately randomising test order exposes this class. 3. **Ordering assumptions.** Asserting an order the system never guaranteed — completion order of concurrent tasks, iteration order of an unordered collection, log line sequence. Fix: assert on sets and invariants, not sequences, unless order is actually part of the contract. 4. **Real scheduler dependence.** Depending on how a thread pool happens to dispatch. Fix: inject a controllable scheduler so the interleaving is chosen rather than observed. 5. **Environmental dependence.** Real network, real filesystem timing, DNS, ports, container CPU quotas, unseeded randomness. Fix: stub the boundary, allocate resources dynamically, seed every random source and print the seed on failure. 6. **A genuine product race.** The test is fine; the code is broken and the test is the only thing noticing. This is the outcome that justifies the whole exercise, and the reason blanket retries are dangerous. Any triage must ask "could this be real?" before assuming test error. ## Verifying a fix A flaky test that passes once proves nothing. Re-run the fixed test many times — hundreds of iterations, ideally on a loaded machine and with randomised ordering — before declaring it stable, and keep watching its recorded history afterwards. Improve diagnosability at the same time: on failure, dump the relevant state, the seed, thread states and queue contents, so the next occurrence is one-shot diagnosable instead of another mystery. ## The judgment calls to voice - **Delete versus fix.** A high-flake, low-value test that duplicates coverage should be deleted outright; effort belongs where the test protects something real. - **Where determinism cannot reach.** Some properties are inherently probabilistic (timing under real load). Those belong in a separate non-blocking job with trend-based alerting, not in the merge gate. - **Budget explicitly.** Set a target — for example, no test above a defined flake rate in the blocking gate — and fund the work to hold it, because the alternative cost (ignored failures, escaped regressions, slow merges) is larger and quieter.
- Why not simply configure the CI system to retry failing tests up to three times?Because retries treat the symptom and destroy the signal. Any intermittent product bug — a real race that fails one run in fifty — is exactly what a retry policy erases, so the most valuable failures become invisible. Retries also let the flaky population grow unchecked, since nobody feels the pain, and they inflate build times. A narrow, logged retry around a genuinely flaky external dependency is defensible; a global retry policy is not.
- How do you decide whether an intermittent failure is a bad test or a real product race?Read the failure mode rather than the frequency: if the assertion failed because the system reached a state it should never reach — lost update, duplicated side effect, inconsistent invariant — that is a product bug regardless of how rare it is. If it failed because a wait expired while everything was still consistent, it is usually a timing-dependent test. When in doubt, reproduce the failing interleaving on a controllable scheduler; if the bad state is reachable in a deterministic replay, the product is wrong.
- What would you put in place so the next flaky test is diagnosable on its first failure?Make failures self-describing: dump the observed and expected state, the random seed, pending queue contents and thread states at the moment of failure, and attach them to the CI run. Record per-test outcome history so flakiness is detected automatically rather than by folklore, and randomise test ordering in at least one scheduled job to surface inter-test coupling early. These cost little and convert 'it failed again, no idea why' into a one-shot diagnosis.
A smoke alarm that chirps randomly: people take the battery out, and then it cannot warn about the real fire. Fixing the chirp is what preserves the alarm.
saying these in an interview costs you the question
- Adding automatic retries suite-wide and calling the problem solved.
- Assuming an intermittent failure is always a test defect and never a real race.
- Muting or deleting flaky tests without checking whether they protect something valuable.
- Declaring a fix verified after a single green run instead of many iterations.
- Treating flakiness as an individual's chore, with no measurement, ownership or deadline — so the population keeps growing.