skip to content

What is a flaky test, and why is it worse than a test that fails every time?

level: juniorimportance: must knowfreq 78%

answer

  1. Same code, different verdicts
  2. Not the single red run
  3. Something the test does not control
  4. Re-run reflex hides real regressions
  5. Consistent failure is honest; intermittent is not

basics

~20 s

A flaky test gives different verdicts on unchanged code, passing on one run and failing on the next. A consistent failure names a real problem; a flaky one teaches the team to re-run instead of investigate, destroying the suite's signal.

solid answer

~50 s

A test is flaky when it passes and fails across runs with no change to the code under test, to the test, or to its data. The verdict is being decided by something the test does not control: elapsed time, thread interleaving, a real clock or random source, an external service, the order cases run in, or state left behind by an earlier case. The damage is not the single red run — it is the erosion of what red means. A failure that might be nothing trains people to press re-run, and once that reflex exists genuine regressions ride out on the same shrug. A consistently failing test is honest: it points at one defect and blocks until someone deals with it. Flakiness turns a clear verdict into a probability nobody can act on, which is why teams treat a flaky test as a defect in its own right.

go deeper

for a junior

Be ready to define a flaky test in one sentence — different verdicts on unchanged code — and to say plainly that re-running until green is not a fix. This is used as a screening question, so a crisp definition matters more than a long list.

for a middle

Explain why the intermittency is the point: the verdict is being decided by an input the test does not control. Name that class of input before you name a fix, and distinguish a true flake from a case that fails deterministically in one environment.

for a senior

Show that you protect the suite's signal. Capture evidence on the failing run rather than the green re-run, refuse to classify a failure by how rare it is, and treat an unexplained intermittent failure as a possible product defect until you have a cause.

for a principal

Own the argument that flakiness is a cost centre rather than a nuisance: engineering hours, discounted red results, and defects that escaped because a real failure looked like noise. Be able to state the standard you would hold a blocking pipeline to.

## The definition, stated precisely A test is **flaky** when repeated executions of the *same* test against the *same* code produce **different verdicts**. Nothing in the change under test explains the difference: the code, the test and its data are identical between the green run and the red one. Something outside the test's control — timing, thread interleaving, the real clock, a random source, an external system, the order the cases ran in, or residue left by an earlier case — decided the outcome. Two boundaries matter, because candidates blur them constantly: * **Flaky is not the same as environment-dependent.** A case that fails *every* time on the build machine and passes *every* time on a developer machine is perfectly deterministic in each place. That is a valuable finding, not a flake: the difference between the two environments — locale, time zone, available cores, data volume, degree of parallelism — is the defect's address. * **Flaky is not the same as broken.** A suite that fails because the build does not compile, a dependency is missing, or a service is down for everyone is failing deterministically for a knowable reason. "Unstable" in casual speech covers both; in an interview, say which one you mean. Note also that *flaky* describes an observed **verdict**, not a location. The nondeterminism may live in the test, in the harness, in the environment, **or in the product**. That last possibility is why the label is dangerous when applied too early. ## Why it is worse than a hard failure A suite is a classifier: red is supposed to mean "there is a problem here". Its value depends entirely on how much information a red result carries. When a meaningful share of red runs turn out to mean nothing, the rational response of every engineer looking at one is to discount it — to re-run first and read second. That reflex is learned quickly and unlearned slowly, and it applies to **all** red results, not only the ones that were spurious. A single unfixed flake therefore taxes the credibility of every other test in the suite. The second cost is diagnostic. The re-run that turns the pipeline green also destroys the evidence: the actual values, the inputs, the ordering and the logs from the moment of failure are gone, and what remains is a green run that explains nothing. Teams that habitually re-run accumulate failures nobody can ever investigate, because the only artefact they keep is the one that succeeded. The third cost is the one that shows up in production. Intermittency in the *product* looks, in a run report, exactly like intermittency in the *test*. If the team has learned to treat rare red as noise, a genuine rare defect arrives wearing the same costume and is filed with the rest. By contrast, a test that fails every single time is a good citizen. It is unambiguous, it is reproducible on demand, it blocks the pipeline until someone resolves it, and it names one defect. Nobody has to make a judgement call about whether to believe it. ## A worked illustration Consider a nightly regression pack of 340 cases for a utility billing run, taking about 47 minutes. Over several weeks, 14 of the 340 cases had failed at least once and passed on re-run, so the team adopted an informal rule: re-run the pack before reading anything. One night, a case checking an account's invoice total failed on a currency-rounding drift — the total was off by one hundredth of a unit. It passed on the re-run and was mentally filed with the other 14. It was not a flake. The billing run summed line charges in the order its workers happened to finish and rounded each partial sum, so the total drifted by a hundredth whenever one worker landed late. The suite had caught a real, customer-visible defect on the first night it existed. The flakes around it made that signal unreadable, and the team's own coping mechanism completed the failure. ## What to say and what to do Define flakiness by the *verdict* (different results, unchanged code), then immediately say what it costs (the meaning of red, the evidence, and the ability to spot real intermittency). Say that re-running to green is a coping mechanism, never a fix: the fix is to find what the test failed to control and to control it, or to accept that the nondeterminism was real and was in the product. And treat the flaky test itself as a defect with an owner and a cause, rather than as weather the team has to live with.

  • Is a test that fails on every build-machine run and passes on every local run flaky?
    No — it is deterministic in each environment, which makes it far easier to work with than a true flake. The finding is the difference between the two environments: locale, time zone, core count, data volume, degree of parallelism, or available services. Chase that difference. It becomes flakiness only when the same environment produces different verdicts on unchanged code.
  • If a case fails once in several hundred runs, does that prove the product is fine?
    No. The nondeterminism may sit in the product rather than the test: a race, an unsynchronised cache, or a calculation whose result depends on the order work completes can produce rare, genuinely wrong, user-visible answers. Intermittency alone says nothing about which side the fault is on. Treat the failure as unexplained until you have a cause.
  • Why do teams call a flaky test a defect rather than a nuisance?
    Because it consumes what a product defect consumes — engineering hours, pipeline capacity and release confidence — and unlike a product defect it degrades the credibility of every other test. Once some share of red runs is known to be meaningless, every red run gets discounted. That is a systemic cost, so it deserves an owner, a cause and a fix like any other defect.

A smoke alarm that goes off whenever someone makes toast does not just annoy people — it teaches the household to ignore the alarm, which is exactly the day the kitchen catches fire.

saying these in an interview costs you the question

  • Calls any failure seen only on the build machine flaky
  • Says re-running until green counts as a fix
  • Assumes an intermittent failure can never be a product defect
  • Treats flakiness as unavoidable in any large suite
  • Confuses a flaky test with a merely slow test

context