skip to content

In Cypress Cloud Branch Review, what do new, existing and resolved mean for flake?

level: seniorimportance: should knowfreq 50%

answer

  1. Two runs: a base and a changed one
  2. Three words describe every result
  3. Not previously captured, but now is
  4. Proves a stability fix took effect
  5. One recorded run per commit

basics

~20 s

Branch Review compares a changed run against a base run. On its Flaky tab, new means the flake was not captured on the base, existing means it was already there, and resolved means it was there before and no longer appears.

solid answer

~40 s

Branch Review sets one recorded run as the **base** and another as the **changed** run, then reports how each status shifted between them across its Failures, Flaky, Pending, Added and Modified tabs. On the **Flaky** tab every spec that flaked is marked **new** (not previously captured, so possibly introduced by this branch), **existing** (already present on the base branch) or **resolved** (present before, gone now). That three-way split is what lets a reviewer hold an author accountable for the flake their branch introduced without blocking them on flake they inherited, and it is how you confirm a stability fix actually took effect rather than merely coinciding with a quiet run. Opening a marked test gives a side-by-side comparison of both branches' attempts, artifacts and code diff.

go deeper

for a junior

Know that Branch Review compares two recorded runs and that each flaky result is labelled new, existing or resolved relative to the base.

for a middle

Explain what each label is derived from, and that all the counts describe the changed run while the base run only supplies the comparison.

for a senior

Demonstrate the review judgement: hold an author to the new flake, leave the existing flake to its owner, and use resolved as the evidence a fix landed.

for a principal

Own the conditions that make the comparison trustworthy across a monorepo, especially one recorded run per commit, and what you do when the base commit has no run.

## The comparison Branch Review makes Branch Review is built on exactly one idea: a pair of recorded runs, a **base** run and a **changed** run. The changed run is the subject — usually the latest run on a pull request's branch — and the base run is the point of comparison, normally the branch it will merge into. Every number on the screen refers to the changed run; the base run only supplies the "was it like this before?" half of each answer. The review surface splits results into tabs: **Failures**, **Flaky**, **Pending**, **Added** and **Modified**. The Flaky tab lists every spec that flaked in the changed run, and marks each result with how it compares against the base. ## The three states | State | Meaning | What a reviewer does with it | | --- | --- | --- | | **new** | The state was not previously captured, but now is | Look here first — this branch may have introduced it | | **existing** | The state was captured before and still is | Not this branch's doing; do not block the author on it | | **resolved** | The state was captured before and no longer is | Evidence a stability fix worked | Alongside the states, count indicators show the direction of travel for the whole tab: how many were introduced in this branch, how many decreased or resolved, and the plain total. So a Flaky tab reading "3 new and 1 existing" is a different conversation from one reading "4 existing". ## Why the split matters at review time Without the comparison, a reviewer sees a flaky count on a pull request run and has two bad options: treat all of it as the author's problem, or ignore all of it. Both are wrong, and both erode the signal. The three states let you take a position that survives argument: - **new** flake is the branch's responsibility while it is still cheap to fix, before it merges and becomes everyone's background noise. - **existing** flake is pre-existing debt. Blocking a storefront checkout change on flake that has been in the marketing package for a month teaches people to bypass the check. - **resolved** flake is the only honest way to demonstrate that a stability fix landed. A spec that simply did not flake this time looks identical to a spec that was fixed, until you compare it against a base run where it did. Opening a marked test drops into a side-by-side view: the base branch's result on one side, the feature branch's on the other, the attempts for each in descending order, the artifacts and Test Replay per attempt, and the diff of the test's own code. Comparing a failing attempt on your branch against a passing one on the base is often the shortest path to what changed. ## When the comparison is not what you think Branch Review can only compare what has been recorded, and the fallbacks are quiet: | Base branch | Changed branch | What you actually see | | --- | --- | --- | | has a run | has a run | A real comparison using both | | has a run | no run | Comparison against the last run on the feature branch | | no run | has a run | The feature run's data, with no comparison | | no run | no run | The last run on the feature branch, no comparison | If the required run is missing on the merge base or the feature commit, a banner names the commit that is missing it. You can also pick the two runs by hand, which is worth doing when the automatic base is not the run you meant. ## Keeping the comparison honest - **Record one run per commit.** Branch Review relies on the *latest* run for a commit. If a commit produces several separately recorded runs — one per package in a storefront monorepo, say — the comparison silently uses only the last of them and undercounts everything else. Several invocations for one commit must be tied together into a single Cloud run. - **Do not read new as proven.** A spec marked new may simply never have run on the base branch, or may have been renamed. New means "not previously captured", which is weaker than "introduced here", and renaming a spec file or a test title makes Cypress treat it as a brand-new test. - **The comparison works without a pull request.** Any two runs can be compared, which is how you check a nightly run against the last release build rather than only reviewing branches.

  • A spec is marked new on the Flaky tab. Why is that not proof the branch caused it?
    New means the state was not previously captured on the base run, which also covers a spec that never ran there, a spec added by this branch, and a test whose title or file was renamed and is therefore treated as brand new. Check whether the base run actually exercised it before assigning blame.
  • Why is recording one run per commit a prerequisite for a trustworthy Branch Review?
    The comparison uses the latest run on each side. If a commit records several independent runs, everything outside the last one is invisible, so failures and flake from the other runs vanish from the review. Several invocations for one commit have to be recorded as a single Cloud run instead.

saying these in an interview costs you the question

  • Treats every flaky test on a branch as newly introduced
  • Thinks resolved means the test was deleted
  • Ignores that the base run may not exist
  • Reviews the flaky count without comparing against a base
  • Records many runs per commit and trusts the comparison