How does Cypress Cloud turn a test's flake rate into a severity band?
answer
- It is a share of runs, not attempts
- The denominator is every run in the window
- Three bands above zero
- Ten and fifty are the boundaries
- Rate normalises across busy and quiet projects
basics
~20 sFlake rate is the share of recent runs where the spec had a flaky test, over all runs in the window. Cypress Cloud bands it Low above 0 to 10 percent, Medium above 10 to 50, and High above 50.
solid answer
~40 sCypress Cloud scores a flaky test by **how often** it flakes rather than how badly it failed. The **flake rate** is the number of runs in the current window where the spec had a flaky test, divided by the total number of runs in that window, so it is a share of runs and not a count of attempts. That rate maps to a **severity** band: no severity at 0%, **Low** above 0 to 10%, **Medium** above 10 to 50%, and **High** above 50%. The flaky-test analytics page groups tests by severity and orders its log by it, so a handful of high-severity specs float to the top of a long list. Because it is a rate, a spec that runs ten times a day is comparable with one that runs weekly.
go deeper
Know that Cypress Cloud scores flaky tests by how often they flake and shows that as Low, Medium or High, so you can point at the worst offenders in a list.
Be ready to state the formula: runs where the spec flaked over total runs in the window, and the boundaries at 10% and 50% that separate the three bands.
Explain why a rate beats a count when specs run on different schedules, and why a rate always lags a fix by the length of its window.
Own what the project-level flakiness number is used for: a goal a team can be held to, and its limits when the specs behind it differ wildly in how much they block.
## How the rate is computed Cypress Cloud does not score a flaky test by how loudly it fails; it scores it by **how often** it does it. The **flake rate** for a spec is: > runs in the current window where the spec had a flaky test **divided by** the total number of > runs in that window. Two things follow from that definition, and both catch people out: - The numerator counts **runs**, not attempts and not tests. A spec that flakes three times inside one run still contributes one run to the numerator. - The denominator is **every run in the window**, including the ones where the spec was perfectly well behaved. So the rate moves when the spec improves *and* when the project simply runs more often. ## The severity bands | Severity | Flake rate | | --- | --- | | — | 0% | | Low | above 0% up to 10% | | Medium | above 10% up to 50% | | High | above 50% | The bands are coarse on purpose. A spec that flakes in more than half its runs is qualitatively a different problem from one that flaked once last month, and a three-way split says that without inviting an argument about whether 12% is meaningfully worse than 14%. ## Why a rate rather than a count A raw count of flaky occurrences rewards the quietest schedule. In a storefront monorepo, the checkout package's specs may run on every push to every package, while the marketing package's specs run on a nightly schedule. Counting occurrences would make checkout look like the whole problem simply because it executes twenty times more often. Dividing by runs normalises that away: a nightly spec that flakes on two nights out of four scores 50%, and the same 50% means the same thing on the busiest spec in the repo. Organisation-wide flake reporting uses rate for exactly this reason — so a high-volume project does not look worse than a quiet one just for running. ## What the analytics page does with the number The flaky-test analytics page for a project turns the rate into four things you can act on: 1. A plot of the **number of flaky tests over time**, so reliability can be shown trending up or down across releases rather than asserted. 2. A single project-wide **flakiness level**, which is the number teams set goals against. 3. The count of flaky tests **grouped by severity**, so effort can be aimed at the top band. 4. A filterable **log of every flaky test, ordered by severity**. Selecting one test opens a details panel with a historical log of its latest flaky runs, the most common errors across those runs, the test case changelog, and a plot of its **failure rate and flake rate over time** side by side. ## Spec-level rate, test-level list One subtlety trips people up when they try to reconcile two screens. The rate is computed from runs in which **the spec** had a flaky test, while the analytics log lets you select an individual **test case** and open its own history. So a spec holding twelve tests contributes one affected run to the numerator no matter which of its tests flaked, or how many of them did. That is deliberate: the unit that costs you a rerun is the spec, not the individual assertion inside it. It does mean you should not expect the rates of a spec's tests to sum to the spec's own rate. ## Reading the rate honestly - **The window matters.** The rate is always "of recent runs". A spec fixed last week still carries the flake from the runs still inside the window, so the number lags the fix; watch the trend rather than one reading. - **A rate is not an impact score.** A Low-severity flake in the storefront's checkout spec blocks more merges than a High-severity flake in a spec nobody gates on. Severity ranks frequency, and frequency is only one input. - **Zero severity is not the same as no data.** A spec that has never run in the window has no rate at all, which is a different statement from a rate of 0%. - **Exported flake rates are whole numbers.** Organisation-level flake reports express the same measure as an integer percentage, so one flaky test over four runs is reported as `25`, matching the in-app rate rather than a fraction. - **A band is a starting point, not a verdict.** Two specs in the same band can have very different causes, and the rate says nothing at all about why either one is unstable. The practical value of the banding is that it converts an unbounded list into a short one. A monorepo with two hundred specs will have dozens with a non-zero rate; only a few of them will sit in the High band, and those few are the ones a reviewer can be asked to look at by name.
- Why does Cypress Cloud rank flaky tests by rate instead of by how many times they flaked?A raw count rewards specs that run rarely and punishes ones that run on every push. Dividing by the number of runs in the window makes a nightly spec and a per-push spec directly comparable, so the ranking reflects instability rather than schedule.
- A storefront spec was fixed yesterday but still shows Medium severity. Why?The rate is computed over a window of recent runs, and the runs where it flaked are still inside that window. The number lags the fix until those runs age out, so judge the fix by the flake-rate trend line rather than by today's band.
saying these in an interview costs you the question
- Thinks flake rate counts failed attempts
- Believes severity reflects how badly the test failed
- Ignores the window and treats the rate as all-time
- Compares raw flake counts across differently scheduled projects