skip to content

Flake Signals

What the service hands back after a run: which tests only passed on a retry, how often each one does it, and a replay of the run itself rather than a video of the screen.

on this pageshow

explore

questions

6

In Cypress Cloud, what does it mean when a recorded run is flagged flaky?

level: juniorimportance: must knowfreq 65%

answer

  1. Same commit, two disagreeing outcomes
  2. Count the attempts, not the runs
  3. Needs retries to produce a second attempt
  4. A badge on the Latest Runs row
  5. Tests for Review orders notable results

basics

~20 s

At least one test in that run failed an attempt and then passed on a later attempt of the same code. Cypress Cloud counts those tests and badges the run, so a run can finish green and still be flaky.

solid answer

~40 s

A **flaky** test in Cypress Cloud is one that both failed and passed inside a single recorded run, with no code change between the attempts. Cloud can only see that pattern because test retries produced more than one attempt, giving it a failing attempt and a passing one to compare. On the **Latest Runs** page the run then carries a flaky-test count, and a **Flaky** filter shows or hides runs that contain them. Inside the run, the **Tests for Review** panel lists the notable results ordered failed, then flaky, then modified, and badges each flaky one so you can open its attempts directly. The run's own status is unaffected: a test that eventually passed leaves the run green, which is exactly why the separate count exists.

go deeper

for a junior

Be ready to say plainly what flaky means: failed once, passed again, same code. Know that Cypress Cloud shows it as a count on the run rather than as a failure.

for a middle

Explain the mechanism, not just the word. Attempts are what Cloud compares, retries are what produce them, and the final status of the test is decided by the attempt that passed.

for a senior

Show that you read the flaky count as a second, independent signal beside the run's red or green, and that you know a green pipeline is not evidence the suite was stable.

for a principal

Own the argument that flake has to be measured rather than remembered, and be clear about what a per-run flag can and cannot support before anyone builds a policy on top of it.

## What Cypress Cloud means by flaky A **flaky** test is one that both fails and passes on the *same code*. Cypress Cloud does not infer that from a hunch or from a long history of red builds. It sees it directly, inside one recorded run, because the test produced more than one **attempt** and those attempts disagreed: the first attempt failed, a later attempt passed, and nothing about the application or the spec changed in between. The only honest label for that is unstable. That has a hard prerequisite. Attempts only exist when test retries are enabled for the run, and the run has to be recorded to Cypress Cloud for the service to see them at all. A recorded run with retries switched off reports passes and failures and nothing else — there is no second data point to compare, so no test in it can ever be flagged flaky. This is the most common reason a team says "Cypress Cloud never shows us any flake": the detector has nothing to work with. ## The three outcomes, side by side | First attempt | Later attempts | Final test status | Flagged flaky | | --- | --- | --- | --- | | Passed | none needed | passed | no | | Failed | one of them passed | **passed** | **yes** | | Failed | all of them failed | failed | no | The middle row is the whole point of the feature. Its final status is *passed*, so it contributes nothing to the run's failure count; if you only read the red/green result of the pipeline you will never learn that it happened. The top and bottom rows are what everybody already sees. ## Where the flag actually appears - On the **Latest Runs** page, a run containing flaky tests carries a **flaky-test count** beside its status, and a **Flaky** filter shows or hides those runs. That is the day-to-day view of how much flake reached the pipeline. - Inside a run, the **Tests for Review** panel consolidates the results that need a human and orders them **Failed**, then **Flaky**, then **Modified**, badging each with the reason it was considered notable. Hovering a row exposes that test's artifacts. - The **Test Results** tab lets you sort and filter the full result list by status, flaky included. - Opening the test gives you its individual attempts, so you can look at the attempt that failed rather than the one that finally passed. ## A storefront monorepo example Suppose the checkout package's spec `cypress/e2e/checkout.cy.ts` holds a test `Checkout > applies a discount code`. On CI it fails its first attempt with an assertion error on the cart total, then passes its second attempt against the very same commit. Cypress Cloud records three separate things: 1. the **run status** as passed, because nothing ultimately failed; 2. a **flaky count of 1** on the run's row in Latest Runs; 3. the test in **Tests for Review**, badged flaky, carrying both attempts and their artifacts. A teammate reading the pipeline sees green. A teammate reading Cypress Cloud sees green **and** a 1. Those are two genuinely different pieces of information about one run, and only the second one survives into next week. ## What the flag does and does not say - It says the test disagreed with itself on one commit. It does **not** say the application is fine — a real race in the storefront's cart code produces exactly this signal. - It is **not** a failure. A test that fails every attempt is *failed*, not flaky, and Tests for Review lists it above the flaky ones. - It is not the same as *modified* or *pending*, which the same panel also surfaces. Those describe what changed about the test, not how it behaved. - One flagged run is a single data point, not a rate. How often a given test does this across many runs is tracked separately, and that is the number worth comparing between tests. The reason the flag exists at all is that the alternative is anecdote. Without it, "checkout is flaky sometimes" lives in one engineer's memory and leaves with them; with it, the same claim is a count on a run anyone can open.

  • Can a Cypress Cloud run be flagged flaky and still show a failed status?
    Yes. The flaky count and the run status are independent. One run can hold a test that failed and then passed, making it flaky, and another test that failed every attempt, making the run failed. Tests for Review lists both, failures first, so you can tell which is which without opening each spec.
  • Why does a Cypress Cloud run with retries disabled never report flaky tests?
    Flake detection works by comparing attempts of the same test on the same code. With retries off, every test produces exactly one attempt, so a failure stays a failure and there is nothing to disagree with. The run still records fully; it simply carries no flake signal.

saying these in an interview costs you the question

  • Thinks a flaky flag means the run failed
  • Believes Cypress Cloud detects flake from one attempt
  • Assumes a green run cannot contain flaky tests
  • Confuses a flaky test with a skipped or pending one
  • Calls a test that fails every attempt flaky
open as a page

Why can a Cypress Cloud test show flake while its failure rate is zero?

level: seniorimportance: must knowfreq 62%

basics

~20 s

The two rates count different things. Failure rate counts runs where the test ended failed; flake rate counts runs where an attempt failed before it passed. With retries a test can fail attempts every run and still finish green.

open as a page

How does Cypress Cloud turn a test's flake rate into a severity band?

level: middleimportance: should knowfreq 44%

basics

~20 s

Flake rate is the share of recent runs where the spec had a flaky test, over all runs in the window. Cypress Cloud bands it Low above 0 to 10 percent, Medium above 10 to 50, and High above 50.

open as a page

In Cypress Cloud Branch Review, what do new, existing and resolved mean for flake?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Branch Review compares a changed run against a base run. On its Flaky tab, new means the flake was not captured on the base, existing means it was already there, and resolved means it was there before and no longer appears.

open as a page

In Cypress Cloud, how do you decide what Test Replay captures and who can see it?

level: principalimportance: should knowfreq 28%

basics

~20 s

Treat it as an exposure decision, not just a debugging one. Test Replay uploads DOM, styles, the Command Log, network traffic and console logs to everyone with project access. Keep redaction on, tighten project visibility, and opt out last.

open as a page

Why is Test Replay unavailable for a Cypress run recorded in Firefox?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

Test Replay captures through the Chrome DevTools Protocol, so it only records on Chromium-based browsers: Chrome, Edge and the deprecated Electron. Firefox and WebKit runs still record to Cypress Cloud, but their Test Replay button stays disabled.

open as a page