skip to content

In Cypress Cloud, what does it mean when a recorded run is flagged flaky?

level: juniorimportance: must knowfreq 65%

answer

  1. Same commit, two disagreeing outcomes
  2. Count the attempts, not the runs
  3. Needs retries to produce a second attempt
  4. A badge on the Latest Runs row
  5. Tests for Review orders notable results

basics

~20 s

At least one test in that run failed an attempt and then passed on a later attempt of the same code. Cypress Cloud counts those tests and badges the run, so a run can finish green and still be flaky.

solid answer

~40 s

A **flaky** test in Cypress Cloud is one that both failed and passed inside a single recorded run, with no code change between the attempts. Cloud can only see that pattern because test retries produced more than one attempt, giving it a failing attempt and a passing one to compare. On the **Latest Runs** page the run then carries a flaky-test count, and a **Flaky** filter shows or hides runs that contain them. Inside the run, the **Tests for Review** panel lists the notable results ordered failed, then flaky, then modified, and badges each flaky one so you can open its attempts directly. The run's own status is unaffected: a test that eventually passed leaves the run green, which is exactly why the separate count exists.

go deeper

for a junior

Be ready to say plainly what flaky means: failed once, passed again, same code. Know that Cypress Cloud shows it as a count on the run rather than as a failure.

for a middle

Explain the mechanism, not just the word. Attempts are what Cloud compares, retries are what produce them, and the final status of the test is decided by the attempt that passed.

for a senior

Show that you read the flaky count as a second, independent signal beside the run's red or green, and that you know a green pipeline is not evidence the suite was stable.

for a principal

Own the argument that flake has to be measured rather than remembered, and be clear about what a per-run flag can and cannot support before anyone builds a policy on top of it.

## What Cypress Cloud means by flaky A **flaky** test is one that both fails and passes on the *same code*. Cypress Cloud does not infer that from a hunch or from a long history of red builds. It sees it directly, inside one recorded run, because the test produced more than one **attempt** and those attempts disagreed: the first attempt failed, a later attempt passed, and nothing about the application or the spec changed in between. The only honest label for that is unstable. That has a hard prerequisite. Attempts only exist when test retries are enabled for the run, and the run has to be recorded to Cypress Cloud for the service to see them at all. A recorded run with retries switched off reports passes and failures and nothing else — there is no second data point to compare, so no test in it can ever be flagged flaky. This is the most common reason a team says "Cypress Cloud never shows us any flake": the detector has nothing to work with. ## The three outcomes, side by side | First attempt | Later attempts | Final test status | Flagged flaky | | --- | --- | --- | --- | | Passed | none needed | passed | no | | Failed | one of them passed | **passed** | **yes** | | Failed | all of them failed | failed | no | The middle row is the whole point of the feature. Its final status is *passed*, so it contributes nothing to the run's failure count; if you only read the red/green result of the pipeline you will never learn that it happened. The top and bottom rows are what everybody already sees. ## Where the flag actually appears - On the **Latest Runs** page, a run containing flaky tests carries a **flaky-test count** beside its status, and a **Flaky** filter shows or hides those runs. That is the day-to-day view of how much flake reached the pipeline. - Inside a run, the **Tests for Review** panel consolidates the results that need a human and orders them **Failed**, then **Flaky**, then **Modified**, badging each with the reason it was considered notable. Hovering a row exposes that test's artifacts. - The **Test Results** tab lets you sort and filter the full result list by status, flaky included. - Opening the test gives you its individual attempts, so you can look at the attempt that failed rather than the one that finally passed. ## A storefront monorepo example Suppose the checkout package's spec `cypress/e2e/checkout.cy.ts` holds a test `Checkout > applies a discount code`. On CI it fails its first attempt with an assertion error on the cart total, then passes its second attempt against the very same commit. Cypress Cloud records three separate things: 1. the **run status** as passed, because nothing ultimately failed; 2. a **flaky count of 1** on the run's row in Latest Runs; 3. the test in **Tests for Review**, badged flaky, carrying both attempts and their artifacts. A teammate reading the pipeline sees green. A teammate reading Cypress Cloud sees green **and** a 1. Those are two genuinely different pieces of information about one run, and only the second one survives into next week. ## What the flag does and does not say - It says the test disagreed with itself on one commit. It does **not** say the application is fine — a real race in the storefront's cart code produces exactly this signal. - It is **not** a failure. A test that fails every attempt is *failed*, not flaky, and Tests for Review lists it above the flaky ones. - It is not the same as *modified* or *pending*, which the same panel also surfaces. Those describe what changed about the test, not how it behaved. - One flagged run is a single data point, not a rate. How often a given test does this across many runs is tracked separately, and that is the number worth comparing between tests. The reason the flag exists at all is that the alternative is anecdote. Without it, "checkout is flaky sometimes" lives in one engineer's memory and leaves with them; with it, the same claim is a count on a run anyone can open.

  • Can a Cypress Cloud run be flagged flaky and still show a failed status?
    Yes. The flaky count and the run status are independent. One run can hold a test that failed and then passed, making it flaky, and another test that failed every attempt, making the run failed. Tests for Review lists both, failures first, so you can tell which is which without opening each spec.
  • Why does a Cypress Cloud run with retries disabled never report flaky tests?
    Flake detection works by comparing attempts of the same test on the same code. With retries off, every test produces exactly one attempt, so a failure stays a failure and there is nothing to disagree with. The run still records fully; it simply carries no flake signal.

saying these in an interview costs you the question

  • Thinks a flaky flag means the run failed
  • Believes Cypress Cloud detects flake from one attempt
  • Assumes a green run cannot contain flaky tests
  • Confuses a flaky test with a skipped or pending one
  • Calls a test that fails every attempt flaky