What are the most common causes of a test that fails intermittently on unchanged code?
answer
- Something the test never controlled
- Six families, each with a signature
- Clock, time zone, locale, seed
- Which worker finished first
- Alone versus in the suite
basics
~20 sThe usual families are timing assumptions, concurrency races, uncontrolled inputs such as the real clock, time zone or random values, calls to real external systems, dependence on the order cases run in, and state leaked between cases.
solid answer
~50 sGroup them into families, because each has a recognisable signature. **Timing**: the test assumes work finishes within some elapsed period. **Concurrency**: the result depends on which thread or worker wins, inside the product or between the test and background work. **Uncontrolled inputs**: the real clock, time zone, locale, an unseeded random source, or generated identifiers leak into the expected value. **External dependencies**: a real service, the network, or a data store some other job also writes to. **Ordering and leaked residue**: a case that only passes after another has run, or that leaves rows, files or process-wide state behind. **Resource contention**: parallel workers competing for the same port, path, connection pool or CPU. The families matter operationally: if the case still flakes when run alone in a loop, the cause is inside it; if it is stable alone but flakes in the suite, the cause is between cases.
code
pseudocode · 7 linestest "billing run charges the mid-month tariff":
account = account_with_usage(units = 431)
invoice = run_billing(account, at = now()) # real system clock
assert invoice.total == 128.47 # hard-coded expectation
# passes most days; fails whenever the run crosses a tariff-period
# boundary, and on any machine set to a different time zonego deeper
Learn the list and be able to say it back: timing, concurrency, real clocks and randomness, external systems, ordering, and leaked state. Naming the usual suspects confidently is most of what this question checks at your level.
Do not stop at naming causes — give each family's signature. An interviewer here wants to hear how you would tell an ordering problem from a race before reading any code, using only how and where the failures appear.
Demonstrate a narrowing method: isolate the case, pin one uncontrolled input at a time, and reason about which family survives each experiment. Mention the causes that only surface once a suite runs in parallel, since those are the ones teams miss.
Talk about prevention at the level of design seams and harness defaults — what every case gets for free, such as controllable time, seeded data and per-case isolation — so entire families of flakiness cannot be authored in the first place.
## Why a taxonomy, and not a list of tricks An intermittent failure is only fixable once you know which input the test failed to control. A taxonomy is useful because each family produces a different *signature* under simple experiments, so naming the families is also a diagnostic procedure rather than trivia. ## The families **1. Timing assumptions.** The test assumes an operation completes within some period — a fixed pause, a generous timeout, an assumption that a background job has already run. It passes on a quiet machine and fails when the machine is busy, the data set is larger, or an unrelated job saturates the disk. The signature: failures cluster on loaded machines and disappear when the case runs alone. **2. Concurrency races.** Two or more threads, workers or processes touch the same state and the outcome depends on which one wins. The race can be inside the product (an unsynchronised accumulator, a cache written from two paths) or between the test and work the system started in the background. Signature: the failure mode varies — sometimes a missing value, sometimes a stale one, sometimes a corrupted one — and it gets more frequent on machines with more cores. **3. Uncontrolled inputs.** The expected value silently depends on the real clock, the time zone, the locale, an unseeded random source, or a generated identifier. These are the flakes with calendars: they appear at month boundaries, across a daylight-saving change, at the end of a quarter, or on whatever run the random source produces the awkward value. Signature: sharp, dateable onset, and perfect reproducibility once you pin the input. **4. External dependencies.** The test talks to a real service, the public network, or a shared data store another job also writes. Now someone else's availability, rate limit, latency or data is part of your verdict. Signature: failures correlate with someone else's schedule rather than with your change. **5. Ordering and leaked residue.** A case only passes because an earlier one left rows, files, cached values or process-wide state behind — or fails because an earlier one left the *wrong* state behind. Signature: stable when run alone, unstable in the suite, and sensitive to any change in execution order. **6. Resource contention.** Once cases run in parallel, they compete for ports, temporary paths, fixture rows, connection-pool slots and CPU. Signature: the suite was stable serially and became unstable the week parallelism was turned on. ## Where the six overlap Real flakes often combine families. A billing scenario that computes an expected total from the real clock (family 3) is also usually the one whose worker threads finish in a different order (family 2), and a currency-rounding drift appears only when both line up. Do not stop at the first plausible family; check whether pinning it actually removes the failure. ## Narrowing quickly The experiments follow the families directly: 1. **Run the single case in a loop, alone.** Still flaky? The cause is inside it: timing, concurrency, or an uncontrolled input. Stable alone? The cause is between cases: ordering, residue, or contention. 2. **Pin one input at a time** — fix the clock, fix the seed, fix the locale, fix the execution order — and see which pin makes the failure stop. The pin that stops it names the family. 3. **Vary the load** — run on a busy machine or with more parallel workers. If the rate rises with cores, suspect a race; if it rises with disk or network load, suspect a timing assumption. 4. **Read what the failing run recorded**: the actual value, not just the fact of mismatch. A total off by a hundredth of a unit points somewhere completely different from a total that is absent. ## Illustration A 340-case nightly regression pack for a utility billing run had two unrelated intermittent cases. The first computed the expected charge from the current instant, so it failed on every run that crossed a tariff-period boundary — an uncontrolled input, dateable and perfectly reproducible once the clock was pinned. The second failed roughly one run in sixty with a currency-rounding drift of a hundredth of a unit, and only on the machines with the most cores: the billing run summed line charges in worker-completion order and rounded per partial sum. Same symptom class in the report, two different families, and only one of them was a fault in the test. ## What interviewers listen for They want the families named and, more importantly, the *signature* of each, because that is what separates someone who has debugged flakiness from someone who has read about it. Saying "it is usually timing" and reaching for a longer pause is the answer that fails: it names one family, guesses, and treats the symptom.
- Which of these causes tend to appear only once a suite starts running cases in parallel?Contention causes above all: two workers writing the same fixture rows, competing for a port or temporary path, or exhausting a shared connection pool. Latent ordering coupling surfaces too, because parallel execution reshuffles what runs before what. None of them are visible on a serial run, so a suite that became unstable the week parallelism was enabled is usually sharing a resource rather than racing internally.
- How would you narrow an intermittent failure to one family quickly?Run the single case in a loop, alone. If it still flakes, the cause is inside it — timing, concurrency or an uncontrolled input. If it is stable alone but flakes in the suite, the cause is between cases: ordering, leaked residue or contention. Then pin one input at a time — clock, seed, locale, order — and the pin that makes the failure stop names the family.
- Can the same symptom come from two different families?Routinely, and assuming otherwise is how flake fixes fail. A wrong computed total can come from an uncontrolled clock, from workers accumulating in a nondeterministic order, or from residue left by an earlier case. The actual value matters: a result that is absent, stale, or off by a rounding step each point at different families. Verify the fix by reproducing the original failure and showing the pin removes it.
saying these in an interview costs you the question
- Blames the network without naming what actually varied
- Says a longer pause makes a test reliable
- Uses unseeded random data and calls failures noise
- Assumes every intermittent failure is a timing problem
- Dismisses ordering coupling because cases pass locally
- Never considers that the product is the nondeterministic part