Why can a Cypress Cloud test show flake while its failure rate is zero?
answer
- Attempt status is not test status
- Retries decide the final result
- Two lines on the same chart
- Green pipeline, non-zero instability
- Zero failures can still cost CI minutes
basics
~20 sThe two rates count different things. Failure rate counts runs where the test ended failed; flake rate counts runs where an attempt failed before it passed. With retries a test can fail attempts every run and still finish green.
solid answer
~40 sCypress Cloud tracks **failure rate** and **flake rate** separately for each test, and they answer different questions. The status of an individual attempt is not the status of the test: when retries are on, a test may fail its first two attempts, pass its third, and be recorded as **passing**. Repeat that on every run and the failure rate stays at zero while the flake rate climbs toward 100%. The test details panel plots both lines over time precisely so you can see them diverge. That divergence is the reason flake is easy to miss without dedicated tracking: the pipeline is green, the merge is unblocked, nobody is paged, and the instability is still there, still costing CI minutes and still capable of hiding a real regression behind a retry.
code
bash · 2 linescy-cloud test list --projectId abc123 --runNumber 4021 --status passed \
| jq '.tests[] | select((.attempts | length) > 1) | {testName, attempts}'go deeper
Know that a test can be recorded as passing even though one of its attempts failed, and that Cypress Cloud keeps a separate number for how often that happens.
Explain the mechanism: the final status comes from the attempt that passed, so runs with a failing first attempt raise the flake rate without touching the failure rate.
Show you read the two lines as a pair and can say what each combination implies, and that you know to open the failing attempt rather than the one that passed.
Own the argument that a green pipeline is a lossy summary, and be able to say what your organisation does with the instability the red/green signal discards.
## Two rates, two questions Cypress Cloud tracks two different numbers for a test over a window of recorded runs, and a lot of confusion comes from reading one as if it were the other. | Metric | What it counts | Goes up when | | --- | --- | --- | | Failure rate | Runs where the test's **final** status was failed | Every attempt failed | | Flake rate | Runs where the test flaked at least once | An attempt failed and a later one passed | The failure rate answers *did this test block anyone?* The flake rate answers *did this test behave differently from itself on identical code?* Those are independent facts, and the test details panel plots them as two lines on the same chart because the interesting information is in the gap between them. ## Where the divergence comes from When retries are enabled, a test that fails is attempted again on the same code. The **status of an individual attempt is separate from the final status of the test**. A project that retries a failing test up to three times may fail attempt one, fail attempt two, pass attempt three, and record a final status of **passing**. Nothing about that is a bug or a rounding artefact. Consider the storefront monorepo's `cypress/e2e/checkout.cy.ts` spec over twenty CI runs: - In every one of the twenty runs, `applies a discount code` fails its first attempt on the cart total and passes its second. - Final status in all twenty runs: **passed**. - Failure rate: **0%**. Flake rate: **100%**. A test can therefore show a zero final failure rate while exhibiting flake in every run it appears in. The build goes green each time, and the underlying instability — plus the extra wall-clock time of the wasted attempts — is entirely invisible in the pass/fail signal. ## Why the pair is worth reading together Reading the two lines side by side tells you things neither one says alone: - **Flake high, failures zero.** The retry is absorbing the problem every time. The suite is reporting green results it did not really earn, and the cost is showing up as run duration rather than as red builds. - **Flake high, failures rising.** The retry is no longer enough. Whatever was intermittent is becoming reliable, which usually means a real defect rather than timing noise. - **Flake falling, failures rising.** Often a genuine regression landing on top of a test that used to be merely unstable. - **Both zero.** Either the test is healthy, or it stopped running in the window at all — worth checking before you celebrate. ## Finding the attempt that actually failed Once you decide to look at a flaky test, the evidence you want is in the attempt that **failed**, not the one that passed. Two details matter here: 1. There is no `flaky` status to filter on when listing a run's tests programmatically. A flaky test's status is `passed`; you spot it by having **more than one attempt**. 2. Replay queries default to the **last** attempt, which for a flaky test is by definition the attempt that succeeded. You have to ask for attempt 1 explicitly to see the failure. Both points are easiest to see from the Cypress Cloud CLI (`cy-cloud`), whose replay commands address one attempt at a time and take an explicit attempt number. The example below pipes its JSON output through `jq`, the command-line JSON processor, to keep only the passed tests that needed more than one attempt. Pulling the failing attempt and the passing attempt for the same test and comparing them — a command that ran slower, a request that came back late, a value that differed — is usually what turns "this flakes" into a cause. ## What the zero does not mean - It does not mean the test is fine. A zero failure rate on a test with a high flake rate means the retry policy is doing work every run. - It does not mean the run was cheap. Every extra attempt is real time on a CI machine, multiplied by however many runs the window covers. - It does not mean a regression cannot hide there. A test that already fails half its attempts has no headroom left to signal a genuine break. That last point is the reason the two rates are shown together at all: the pipeline's red/green result is a lossy summary of what the suite actually observed, and the flake rate is the part it threw away.
- In a Cypress Cloud replay, which attempt of a flaky test do you open first, and why?The attempt that failed, normally attempt 1. Replay queries default to the most recent attempt, which for a flaky test is the one that passed and therefore shows nothing wrong. Ask for the earlier attempt explicitly, then compare it against the passing one to find what differed.
- How would you list a run's flaky tests when there is no flaky status to filter on?A flaky test's recorded status is passed, so filter the run's tests to passed and keep the ones with more than one attempt. Each result carries the failing attempt's error beside the passing attempt, which is enough to triage without opening the Cloud UI.
Failure rate is the log of buildings that burned down; flake rate is the log of alarms that went off. A test can trip the alarm every week and never once appear in the fire log.
saying these in an interview costs you the question
- Treats a zero failure rate as proof the test is healthy
- Thinks flake rate and failure rate must move together
- Says retries make the failing attempts disappear from the record
- Opens the passing attempt and concludes nothing went wrong
- Assumes a green run means the suite cost nothing extra