skip to content

How do you measure a test suite's flake rate, and why is its trend more useful than its level?

level: middleimportance: must knowfreq 60%

answer

  1. Same revision, different verdict
  2. Choose the denominator before arguing
  3. Store every attempt, not the last one
  4. Multiply per-case rate by suite size
  5. Publish it beside the quarantine list

basics

~20 s

Flake rate is the share of runs or cases giving a different verdict on the same unchanged revision. Measure it by recording every attempt, not just the final one, and track the trend: an acceptable level depends on suite size.

solid answer

~50 s

Define it operationally: the same revision, re-executed, gave a different result for the same case. Measuring it needs results stored against a stable case identifier and the revision they ran on, and it needs the pipeline to record every attempt - a retry that overwrites the first result makes the metric unmeasurable. Pick your denominator deliberately: a per-run rate describes how often a run lied to the team, while a per-case rate over a rolling window ranks the offenders. The level alone is close to meaningless because it scales with suite size - a tiny per-case rate across thousands of cases still means most runs are unbelievable. So publish the trend over a two-to-four-week window, and publish it next to the size and age of the quarantine list, because the cheapest way to improve a flake rate is to stop counting cases.

code

pseudocode · 23 lines
pseudocode
# every execution recorded as: (case_id, revision, outcome, run_id, started_at)

function disagreeing_cases(results, revision):
    by_case = group_by(results.where(rev = revision), key = case_id)
    return by_case.keys().where(case -> distinct(by_case[case].outcome).size > 1)

function per_run_flake_rate(window):
    spoiled = 0
    for run in window.runs:
        bad = disagreeing_cases(window.results, run.revision)
        if run.outcome == FAIL and run.failing_cases.all_in(bad):
            spoiled = spoiled + 1
    return spoiled / window.runs.size

function per_case_flake_rate(window, case):
    execs = window.results.where(case_id = case)
    minority = execs.count_where(outcome != majority_outcome_for_same_revision)
    return minority / execs.size

function offender_ranking(window):
    return window.cases
        .map(c -> { case: c, spoiled_runs: per_case_flake_rate(window, c) * executions(window, c) })
        .sort_desc(by = spoiled_runs)

go deeper

for a junior

Be able to state the operational definition - same revision, re-run, different verdict - and to say why a suite that disagrees with itself stops being used. Knowing that both outcomes must be recorded already puts you ahead.

for a middle

Explain the two denominators and pick one for a stated purpose, and show the arithmetic that a small per-case rate across a large suite still produces mostly unbelievable runs. Naming retries-overwriting-results as the classic measurement bug is expected.

for a senior

Demonstrate that you have run this on a real suite: the results store you needed, how you kept case identifiers stable, how you excluded infrastructure incidents, and how you ranked offenders by runs spoiled rather than by rate.

for a principal

Own the incentive problem. Say out loud that the cheapest way to move this number is to stop counting cases, and describe the paired metric - exclusion-list size and age - that makes that move visible when it happens.

### What the number is measuring A **flake rate** quantifies how often the suite gives a different verdict on input that did not change. The operational definition that survives an interview is: *the same revision, run again, produced a different outcome for the same case*. Everything else — theories about why it happened — is separate work. To measure it you need two things the naive pipeline throws away: results keyed by a **stable case identifier**, and the **revision** each result was produced against. If the case identifier changes when someone renames a file or reorders a suite, you cannot build a history, and the metric quietly resets to zero every refactor. ### Two denominators, and why the choice matters Almost every disagreement about flake rate is really a disagreement about the denominator. - **Per-run flake rate** = runs that failed only because of a case that later passed on the same revision, divided by all runs. This is the number that describes the *experience* of the team: how often a run lied to somebody. - **Per-case flake rate** = for a given case, the share of its executions in a rolling window that disagreed with the majority outcome for the same revision. This is the number you rank an offender list by. They move independently and both are needed. A suite can hold its per-case rate flat while the per-run rate climbs, purely because the suite got bigger — which is the arithmetic below. ### The arithmetic that makes the level meaningless on its own Consider a dispatcher service with about **3,400** cases in its shared-branch tier, running as a **27-minute suite**. Suppose every case independently misbehaves on 0.08% of executions — a rate that sounds negligible per case. The chance a whole run is clean is roughly `0.9992 ^ 3400`, which is about **6.6%**. In other words, a per-case rate most people would call excellent produces a run that is believable about one time in fifteen. Reverse it: if you want three runs in four to be clean at that suite size, you need a per-case rate of roughly 0.008% — an order of magnitude better. This is why "what is an acceptable flake rate?" has no context-free answer, and why quoting a single percentage as a target is a trap. The pair that means something is *per-case rate together with suite size*, or simply the per-run rate. ### How you actually detect it Three mechanisms, usually combined: 1. **Re-execution on the identical revision.** When a case fails, run it again against the same revision. Disagreement is direct evidence. Record *both* results — the most common measurement bug is a pipeline that retries silently and stores only the final verdict, which makes the flake rate structurally unmeasurable. 2. **Repeat runs of an unchanged revision.** Schedule the shared-branch tier to run several times against a revision that is not moving. Any disagreement across those runs is unambiguous, and unlike per-failure re-execution it also surfaces cases that fail rarely enough that nobody has noticed. 3. **Cross-run comparison in a results store.** Keep every result with case id, revision, worker, shard, start time and duration. Then a case that produced both outcomes for one revision is detectable after the fact, without any special run mode, and the extra dimensions let you spot whether disagreements cluster on one worker or one shard. ### Why the trend beats the level Four reasons, and a good answer gives at least two: - **The acceptable level is context-dependent** — see the arithmetic above. A trend is comparable against yourself; a level is not comparable against anyone. - **A trend is actionable.** "1.9% and falling for six weeks" describes a team that is winning. "1.9%" describes nothing. - **A level invites a target, and a target invites gaming.** The cheapest way to lower a flake rate is to stop counting cases: quarantine them, mark them non-blocking, or delete them. All three improve the number while making the suite weaker. - **Levels are noisy at small numbers.** With a handful of unreliable events a week, a week-over-week percentage swings wildly; a rolling window of two to four weeks is what you plot. Because of the third point, the flake-rate trend is only honest when it is published next to the **size and age of the quarantine list** and the **count of blocking cases**. A falling flake rate with a growing exclusion list is not an improvement, it is a relabelling. ### What a useful report contains Rank offenders by cost, not by rate: a case that misbehaves on 4% of its executions but runs on every push blocks far more people than one that misbehaves on 30% of its executions in a weekly tier. A workable ordering is *disagreement probability multiplied by executions in the window*, which yields "runs this case spoiled" — a number you can put next to an owner's name. Add the age of the oldest unresolved offender, and the share of red runs attributable to unreliability rather than to real defects; that share is what tells you whether people are right to distrust the suite. ### Confounders to name unprompted Retries that overwrite results; suite growth changing the per-run rate with no change in behaviour; cases removed or quarantined mid-window; an infrastructure incident producing a day of mass disagreement that should be annotated rather than averaged in; and case identifiers that change under refactoring and silently truncate the history.

  • Your pipeline retries every failed case twice and stores only the final verdict. What has that cost you?
    The flake rate becomes unmeasurable. A case that failed and then passed is recorded as a pass, so the evidence of disagreement is destroyed at the moment it is produced, and the metric will report near zero no matter how unreliable the suite is. The fix is to keep every attempt with its outcome and mark the run as recovered-by-retry, so the run stays green for the developer while the disagreement is still counted.
  • Per-case flake rate is flat all quarter but the per-run rate has doubled. What happened?
    Almost certainly the suite grew. The chance of a clean run is roughly the per-case reliability raised to the number of cases, so doubling the case count degrades the per-run experience with no change in any individual case's behaviour. The response is different from a normal flake investigation: either improve per-case reliability to hold the product constant, or split the suite into tiers so fewer cases gate any one run.
  • How do you rank unreliable cases for repair when there are more of them than you can fix?
    Rank by blocked work, not by rate. Multiply each case's disagreement probability by how many times it executes in the window to get the number of runs it spoiled, and fix from the top. A moderately unreliable case on the per-push tier outranks a wildly unreliable one in a weekly tier. Cap the list, give each entry an owner and an age, and report the age of the oldest unresolved entry.

It is a measurement of the instrument, not of the patient: you are asking how often the thermometer disagrees with itself on the same forehead.

saying these in an interview costs you the question

  • Quotes a universal acceptable flake percentage with no reference to suite size
  • Measures from final verdicts only, so retries erase the evidence
  • Reports a falling flake rate while the exclusion list grows
  • Treats every red run as unreliability rather than separating real failures
  • Uses case names as identifiers so history resets on every rename
  • Averages an infrastructure outage into the trend instead of annotating it

context