skip to content

How do you tell a flaky test from a genuine intermittent defect in the product?

level: seniorimportance: should knowfreq 50%

answer

  1. Frequency classifies nothing
  2. Evidence from the failing run
  3. Replay with the recorded seed and order
  4. Ask what the specification actually promises
  5. A rare wrong answer is still wrong

basics

~20 s

They look identical in a run report, so separate them by cause, not frequency. Reproduce with evidence captured at the moment of failure, then decide whether the nondeterminism sits in the test's assumptions or in the behaviour it observed.

solid answer

~50 s

Start by refusing the shortcut: rarity is not evidence that a test is at fault. Work from the failing run's artefacts — the actual and expected values at full precision, the inputs and any generated data with their seed, the execution order, timestamps and time zone, and the system's own logs for that window — because a green re-run destroys all of it. Then try to reproduce under the recorded conditions and pin one variable at a time. Finally ask the oracle question: was the behaviour the test observed actually acceptable? If the assertion demanded more than the specification promises the test's oracle is wrong. If the system genuinely produced a wrong answer, it is a product defect, and the intermittency is a clue about ordering or concurrency inside it. Record the verdict and the evidence so nobody re-does the work.

go deeper

for a junior

Know the trap: a rare failure is not automatically a problem in the test. Before calling anything flaky you need a cause, and the fact that it passed on the next run is not one.

for a middle

Be able to say what a run must record so a rare failure is diagnosable — actual and expected values, inputs and seed, ordering, timestamps — and why a green re-run destroys exactly the evidence you needed.

for a senior

Show judgement about the test's oracle: decide whether the assertion demanded more than the specification promises, or whether the system genuinely produced a wrong result, and be willing to argue that a rare wrong answer is a real defect.

for a principal

Own the classification standard for the organisation: what evidence a run must produce by default, who decides the verdict, what closes a case, and how an unexplained intermittent failure gets escalated instead of quietly absorbed.

## The two things that look the same In a run report, a nondeterministic test and a rare product defect are indistinguishable: one case, red on this run, green on the next, no code change. The whole skill is refusing to classify on that appearance and getting to a cause. **Why the shortcut is so tempting.** Rare failures are expensive to investigate and cheap to dismiss, and the dismissal is self-reinforcing: re-run, green, move on, and the case never accumulates the evidence that would have argued otherwise. Frequency is not a classifier. A one-in-sixty failure is exactly what a narrow race in the product looks like, and it is also exactly what an unseeded random value looks like. ## Step 1 — preserve the evidence Diagnosis is decided before the investigation starts, by what the failing run kept. At minimum: * **actual and expected values at full precision**, not a truncated "values differed" line; * **the inputs**, including generated data and the seed that produced it; * **the execution order or worker identity** — what ran before this case, and on which worker; * **timestamps and the effective time zone and locale** of the run; * **the system's own logs and any relevant state** for the window of the failure, attached to that run. Diagnosing from a green re-run is guesswork wearing a lab coat. If your suite cannot hand you those artefacts for a failure that happens once in hundreds of runs, that gap is the first thing to fix, because everything downstream depends on it. ## Step 2 — reproduce under the recorded conditions Replay with the recorded seed, ordering and clock. Run the case in a loop, alone and then within the suite, and vary load and worker count. Two outcomes are informative: if pinning an input makes the failure vanish, the test was depending on something it did not control; if the failure persists under fully pinned inputs, the nondeterminism is downstream of the test, in the system. ## Step 3 — interrogate the test's oracle The **oracle** is the rule the test uses to decide whether the observed behaviour is wrong. Many "flakes" are oracle defects rather than control defects. Common shapes: * asserting exact equality on a value the specification only constrains approximately, or under a stated rounding rule; * asserting an order on a result the specification says is unordered; * asserting on a formatted string whose formatting depends on locale; * asserting a total that the system is permitted to compute by more than one legitimate path. If the assertion demands more than the specification promises, the test is wrong and the fix is in the oracle — assert the rule, not an accident. But be careful with the reverse move: loosening an assertion because it keeps failing is how a real defect gets legislated away. Change the oracle only when you can point at what the specification actually promises. ## Step 4 — decide, and weigh the severity If the system genuinely produced an unacceptable result, it is a product defect, and rarity does not reduce its severity — it only reduces the odds of catching it again. Consider a 340-case nightly regression pack for a utility billing run, where one case fails roughly one run in sixty with a currency-rounding drift of a hundredth of a unit. Pinning the clock and the seed did not stop it; running with more workers made it more frequent. The product summed line charges in worker-completion order and rounded per partial sum, so totals drifted whenever a worker landed late. That is not a flaky test. It is a wrong invoice, in production, at whatever rate real traffic produces — and the correct response is a defect report with the reproduction, not a quarantine ticket. The opposite verdict deserves the same rigour. When the cause really is in the test — an unpinned clock, an unseeded generator, residue from an earlier case — say so with the evidence that proves it, and fix the control rather than the symptom. ## Step 5 — write the verdict down Attach the cause, the reproduction and the decision to the case itself. Rare failures recur months later in front of a different engineer, and without a record the whole investigation is repeated from zero — or, worse, the second engineer takes the shortcut the first one refused. ## What interviewers listen for The strong answer states the trap first (frequency classifies nothing), then the method (evidence, reproduction, oracle interrogation), then the willingness to conclude *product defect* when the evidence points there. The weak answer is "I would re-run it a few times and see" — which is the behaviour that lets rare, real defects ship.

  • What must a run capture so that a one-in-hundreds failure is diagnosable the first time it happens?
    Actual and expected values at full precision, the inputs and generated data with their seed, the execution order and worker identity, timestamps with the effective time zone and locale, and the system's own logs for the failure window — all attached to the run that failed. Rare failures do not come back on demand, so anything not captured then is lost.
  • A test asserts an exact total and mismatches by a hundredth of a unit on rare runs. Test bug or product bug?
    It depends on what the specification promises. If the system is only required to produce a total under a defined rounding rule and the assertion demanded exact equality, the oracle is too strict and should assert the rule. If totals are specified to be exact and reproducible, the mismatch is a genuine defect, and its intermittency is a clue pointing at ordering or concurrency inside the calculation.
  • When is loosening an assertion the wrong response to an intermittent failure?
    Whenever you cannot point at the specification clause that permits the variation. Loosening on the grounds that the test keeps failing legislates the defect away and leaves a case that can no longer detect it. Tolerances are legitimate when they encode a stated rule and illegitimate when they encode fatigue.

An unreliable thermometer and a patient whose temperature genuinely spikes once a week produce the same chart; you can only tell them apart by instrumenting the moment of the reading, never by how often it happens.

saying these in an interview costs you the question

  • Classifies a failure as flaky because it is rare
  • Diagnoses from the green re-run rather than the red one
  • Loosens the assertion until the failure stops appearing
  • Reports could-not-reproduce after a single re-run
  • Assumes an unordered result always comes back ordered
  • Treats a rare wrong answer as low severity by default

context