skip to content

Scaling a Run

When one machine is too slow, every answer runs through a recorded run: specs pulled off a shared queue, flake reported back, and a bill for both. Know what you buy.

on this pageshow

explore

questions

15

In Cypress Cloud, what does it mean when a recorded run is flagged flaky?

level: juniorimportance: must knowfreq 65%

answer

  1. Same commit, two disagreeing outcomes
  2. Count the attempts, not the runs
  3. Needs retries to produce a second attempt
  4. A badge on the Latest Runs row
  5. Tests for Review orders notable results

basics

~20 s

At least one test in that run failed an attempt and then passed on a later attempt of the same code. Cypress Cloud counts those tests and badges the run, so a run can finish green and still be flaky.

solid answer

~40 s

A **flaky** test in Cypress Cloud is one that both failed and passed inside a single recorded run, with no code change between the attempts. Cloud can only see that pattern because test retries produced more than one attempt, giving it a failing attempt and a passing one to compare. On the **Latest Runs** page the run then carries a flaky-test count, and a **Flaky** filter shows or hides runs that contain them. Inside the run, the **Tests for Review** panel lists the notable results ordered failed, then flaky, then modified, and badges each flaky one so you can open its attempts directly. The run's own status is unaffected: a test that eventually passed leaves the run green, which is exactly why the separate count exists.

go deeper

for a junior

Be ready to say plainly what flaky means: failed once, passed again, same code. Know that Cypress Cloud shows it as a count on the run rather than as a failure.

for a middle

Explain the mechanism, not just the word. Attempts are what Cloud compares, retries are what produce them, and the final status of the test is decided by the attempt that passed.

for a senior

Show that you read the flaky count as a second, independent signal beside the run's red or green, and that you know a green pipeline is not evidence the suite was stable.

for a principal

Own the argument that flake has to be measured rather than remembered, and be clear about what a per-run flag can and cannot support before anyone builds a policy on top of it.

## What Cypress Cloud means by flaky A **flaky** test is one that both fails and passes on the *same code*. Cypress Cloud does not infer that from a hunch or from a long history of red builds. It sees it directly, inside one recorded run, because the test produced more than one **attempt** and those attempts disagreed: the first attempt failed, a later attempt passed, and nothing about the application or the spec changed in between. The only honest label for that is unstable. That has a hard prerequisite. Attempts only exist when test retries are enabled for the run, and the run has to be recorded to Cypress Cloud for the service to see them at all. A recorded run with retries switched off reports passes and failures and nothing else — there is no second data point to compare, so no test in it can ever be flagged flaky. This is the most common reason a team says "Cypress Cloud never shows us any flake": the detector has nothing to work with. ## The three outcomes, side by side | First attempt | Later attempts | Final test status | Flagged flaky | | --- | --- | --- | --- | | Passed | none needed | passed | no | | Failed | one of them passed | **passed** | **yes** | | Failed | all of them failed | failed | no | The middle row is the whole point of the feature. Its final status is *passed*, so it contributes nothing to the run's failure count; if you only read the red/green result of the pipeline you will never learn that it happened. The top and bottom rows are what everybody already sees. ## Where the flag actually appears - On the **Latest Runs** page, a run containing flaky tests carries a **flaky-test count** beside its status, and a **Flaky** filter shows or hides those runs. That is the day-to-day view of how much flake reached the pipeline. - Inside a run, the **Tests for Review** panel consolidates the results that need a human and orders them **Failed**, then **Flaky**, then **Modified**, badging each with the reason it was considered notable. Hovering a row exposes that test's artifacts. - The **Test Results** tab lets you sort and filter the full result list by status, flaky included. - Opening the test gives you its individual attempts, so you can look at the attempt that failed rather than the one that finally passed. ## A storefront monorepo example Suppose the checkout package's spec `cypress/e2e/checkout.cy.ts` holds a test `Checkout > applies a discount code`. On CI it fails its first attempt with an assertion error on the cart total, then passes its second attempt against the very same commit. Cypress Cloud records three separate things: 1. the **run status** as passed, because nothing ultimately failed; 2. a **flaky count of 1** on the run's row in Latest Runs; 3. the test in **Tests for Review**, badged flaky, carrying both attempts and their artifacts. A teammate reading the pipeline sees green. A teammate reading Cypress Cloud sees green **and** a 1. Those are two genuinely different pieces of information about one run, and only the second one survives into next week. ## What the flag does and does not say - It says the test disagreed with itself on one commit. It does **not** say the application is fine — a real race in the storefront's cart code produces exactly this signal. - It is **not** a failure. A test that fails every attempt is *failed*, not flaky, and Tests for Review lists it above the flaky ones. - It is not the same as *modified* or *pending*, which the same panel also surfaces. Those describe what changed about the test, not how it behaved. - One flagged run is a single data point, not a rate. How often a given test does this across many runs is tracked separately, and that is the number worth comparing between tests. The reason the flag exists at all is that the alternative is anecdote. Without it, "checkout is flaky sometimes" lives in one engineer's memory and leaves with them; with it, the same claim is a count on a run anyone can open.

  • Can a Cypress Cloud run be flagged flaky and still show a failed status?
    Yes. The flaky count and the run status are independent. One run can hold a test that failed and then passed, making it flaky, and another test that failed every attempt, making the run failed. Tests for Review lists both, failures first, so you can tell which is which without opening each spec.
  • Why does a Cypress Cloud run with retries disabled never report flaky tests?
    Flake detection works by comparing attempts of the same test on the same code. With retries off, every test produces exactly one attempt, so a failure stays a failure and there is nothing to disagree with. The run still records fully; it simply carries no flake signal.

saying these in an interview costs you the question

  • Thinks a flaky flag means the run failed
  • Believes Cypress Cloud detects flake from one attempt
  • Assumes a green run cannot contain flaky tests
  • Confuses a flaky test with a skipped or pending one
  • Calls a test that fails every attempt flaky
open as a page

What does a project need before `cypress run --record` uploads to Cypress Cloud?

level: juniorimportance: must knowfreq 62%

basics

~10 s

Two things: a projectId in the Cypress config (or CYPRESS_PROJECT_ID) naming the Cloud project, and a record key passed as --key or exported as CYPRESS_RECORD_KEY. The projectId says where; the key authorises writing there.

open as a page

Why does `cypress run --parallel` fail unless you also pass Cypress's `--record`?

level: middleimportance: must knowfreq 70%

basics

~10 s

Because parallelisation is a Cypress Cloud feature, not a runner feature. The spec-splitting coordinator lives on the server, so Cypress rejects --parallel, --group, --tag, --ci-build-id and --auto-cancel-after-failures up front when --record is absent.

open as a page

Under `cypress run --parallel`, how does Cypress Cloud decide what each machine runs?

level: middleimportance: must knowfreq 68%

basics

~20 s

Cypress Cloud keeps one queue of whole spec files ordered by predicted duration, longest first, and hands the next spec to whichever machine is free. Nothing is assigned in advance, so spec order is not guaranteed.

open as a page

Why can a Cypress Cloud test show flake while its failure rate is zero?

level: seniorimportance: must knowfreq 62%

basics

~20 s

The two rates count different things. Failure rate counts runs where the test ended failed; flake rate counts runs where an attempt failed before it passed. With retries a test can fail attempts every run and still finish green.

open as a page

How does Cypress Cloud turn a test's flake rate into a severity band?

level: middleimportance: should knowfreq 44%

basics

~20 s

Flake rate is the share of recent runs where the spec had a flaky test, over all runs in the window. Cypress Cloud bands it Low above 0 to 10 percent, Medium above 10 to 50, and High above 50.

open as a page

Your CI splits a Cypress suite across four jobs with hand-written `--spec` globs. What does that cost?

level: seniorimportance: should knowfreq 48%

basics

~10 s

It costs maintenance and silent coverage loss. The partition never rebalances as spec durations drift, and a new spec matched by no job's glob is simply never run while every job still exits green.

open as a page

In Cypress Cloud Branch Review, what do new, existing and resolved mean for flake?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Branch Review compares a changed run against a base run. On its Flaky tab, new means the flake was not captured on the base, existing means it was already there, and resolved means it was there before and no longer appears.

open as a page

Your parallel `cypress run --record` jobs land in Cypress Cloud as separate runs. Why?

level: seniorimportance: should knowfreq 45%

basics

~20 s

They are not sharing a CI build id. Cypress ties machines into one run by a build id it auto-detects from the CI environment, so when that value differs per job each job opens its own run. Pass --ci-build-id explicitly.

open as a page

How do you decide whether a Cypress suite should depend on Cypress Cloud to parallelise?

level: principalimportance: should knowfreq 38%

basics

~20 s

Measure the serial wall clock against your feedback budget, try a derived --spec split first, and price metered recording against the suite you are growing into. Then decide what a service outage may do to a deploy.

open as a page

In Cypress Cloud, how do you decide what Test Replay captures and who can see it?

level: principalimportance: should knowfreq 28%

basics

~20 s

Treat it as an exposure decision, not just a debugging one. Test Replay uploads DOM, styles, the Command Log, network traffic and console logs to everyone with project access. Keep redaction on, tighten project visibility, and opt out last.

open as a page

Which Cypress runs do you let Auto Cancellation stop early, and which must run whole?

level: principalimportance: should knowfreq 34%

basics

~20 s

Cancel early where the run exists to give one author a fast verdict: pull-request and branch runs. Let it run whole where the deliverable is a complete picture: release branches, nightly coverage runs, and flake investigations.

open as a page

In a recorded `cypress run`, what does setting `video: true` cost you?

level: juniorimportance: nice to knowfreq 32%

basics

~20 s

One video per spec file, plus per-spec processing time. As of Cypress 16 video defaults to false; turning it on adds capture, optional encoding, and under --record an upload after every spec file, whether it passed or failed.

open as a page

Why is Test Replay unavailable for a Cypress run recorded in Firefox?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

Test Replay captures through the Chrome DevTools Protocol, so it only records on Chromium-based browsers: Chrome, Edge and the deprecated Electron. Firefox and WebKit runs still record to Cypress Cloud, but their Test Replay button stays disabled.

open as a page

In a parallel Cypress run, what does `--auto-cancel-after-failures 3` actually stop?

level: middleimportance: nice to knowfreq 28%

basics

~10 s

Once three tests have failed anywhere in the run, Cypress Cloud stops handing out specs and marks the run canceled. Specs already running finish and report; every spec not yet started is reported skipped.

open as a page