skip to content

Your browser end-to-end suite fails on about one run in ten, a different test each time, and every one of them passes when re-run. How do you triage that?

level: seniorimportance: must knowfreq 62%

answer

  1. one reliability problem, not ten bugs
  2. measure before you fix
  3. rank by merges blocked
  4. artifacts from the failing attempt
  5. some flakes are real product races

basics

~20 s

Treat it as one reliability problem, not ten separate bugs. Record per-test outcomes so you can rank flakes by how often they block a merge, reproduce from the failing run's trace, video and logs, classify the cause, quarantine the worst offenders, then fix them.

solid answer

~50 s

First measure, because the failure you are looking at today is rarely the one costing the most: record every test's pass, fail and passed-on-retry outcome so you can rank flakes by frequency and by merges blocked, and work the top three. Then reproduce from the artifacts of the failed attempt — the trace, video, console and network log — rather than from a fresh local run that will pass. Classify by cause family: waiting on the wrong signal, state or ordering shared between tests, an under-resourced CI runner turning a tight timeout into a failure, app nondeterminism such as animations, real clocks and third-party scripts, and genuine product race conditions. That last family matters most: some flakes are real intermittent bugs, and retrying them ships the bug. Quarantine the worst tests off the merge gate with a named owner and a deadline so the gate is trustworthy while the fixes land.

go deeper

for a junior

Know that re-running until green is not a fix, and that the first useful step is looking at what the failing run actually captured rather than guessing.

for a middle

Explain the common cause families — bad waits, shared state, slow environment, animations and clocks — and how you would reproduce each one deliberately, for example by looping a single test or forcing full parallelism.

for a senior

Demonstrate the process: measure per-test flake rate, rank by merges blocked, work from artifacts, classify structurally, and quarantine with an owner. Be explicit that some flakes are real product races and that retrying those ships the bug.

for a principal

Own the reliability target itself: what flake rate the organisation tolerates, who is accountable when it drifts, how quarantine is prevented from becoming permanent, and how you fund deflaking work against feature pressure.

## Stop debugging one failure at a time A suite that is red one run in ten with a rotating cast of tests does not have ten unrelated bugs; it has a reliability problem with several contributing causes. The trap is to fix whichever test failed this morning, feel productive, and be red again tomorrow. Triage means deciding *which* flakes to spend effort on, and that requires data before analysis. ## Instrument first Record, per test and per run: passed, failed, and passed-only-on-retry. That third outcome is the important one — it is the flake signal that a green pipeline otherwise hides. With a few weeks of history you can rank tests by flake rate and, better, by merges blocked, which weights a rare flake in a test everyone runs above a frequent flake in a nightly-only spec. Almost always a small number of tests produce most of the pain, and fixing three of them changes the suite's character. ## Reproduce from the failed run, not a fresh one Re-running locally is how a flake escapes. The evidence you want belongs to the attempt that failed: the trace or step timeline showing which action timed out and what the DOM looked like at that instant, the video showing whether a modal was open or a spinner was still spinning, the console log for an exception thrown at the moment of failure, and the network log for a request that hung or 500'd. Make sure the pipeline captures those artifacts on failure and retains them long enough to be useful — a flake you cannot see is a flake you cannot fix. ## Cause families and their structural fixes - **Waiting on the wrong signal.** Fixed sleeps, waits on spinners that may never render, assertions on transient chrome, absence assertions with no anchor. Fix by waiting on the durable end state. - **Shared state and ordering.** The test passes alone and fails in the suite, or only when workers run in parallel. The failure is that two tests contend for the same record or account; the fix is that each test owns the data it touches. - **Environment.** Failures cluster on the CI runner and not locally, and the failing step is always the slowest one. A machine running many workers on few cores is genuinely slower; either give the job room or set timeouts that reflect the real environment. Beware of concluding "just CI" — contention frequently exposes real races. - **Application nondeterminism.** Animations and transitions, real clocks and dates near midnight, randomised content, retry banners, and third-party scripts (ads, chat widgets, session recorders) that inject elements or slow the page. Fix by controlling what you can — disabling animations in the test environment, freezing time, blocking third-party requests — and accepting that the rest needs a stronger signal to wait on. - **Genuine product races.** The test is right and the app is wrong: a double-submit, a stale-while-revalidate flash, a request whose ordering is not guaranteed. This is the family that makes retries dangerous, because retrying converts a customer-facing intermittent bug into a green build. ## Amplify to reproduce Once you have a hypothesis, force it. Run the single test many times in a loop; run it with the suite's full parallelism to reproduce contention; throttle CPU or network in the browser to stretch the window; run the specs in a different or randomised order to expose ordering dependence. A flake you can reproduce on demand is a normal bug; the goal of triage is to get every candidate into that state. ## Quarantine as a process, not a graveyard While the fix is being written, a chronically flaky test should stop blocking merges — but it must keep running and keep reporting, with a named owner and an expiry date. Without the owner and the date, quarantine becomes a place tests go to be forgotten, and coverage quietly rots. A good policy makes an expired quarantine entry itself a build failure, which forces the decision: fix it, or delete it and admit the coverage is gone. ## The line to hold The end state of triage is not zero flakes forever; it is that a red suite means something again. Everything above serves that: measurement so you fix the right thing, artifacts so you can see what happened, classification so the fix is structural, and quarantine so the gate stays credible while the work is done.

  • A test passes in isolation but fails when the full suite runs in parallel. What does that tell you?
    That the test does not own everything it depends on. Two workers are contending for the same record, account or shared backend state, or an earlier spec leaves residue the later one trips over. The fix is isolation — each test creating and owning its own data — not a longer timeout, which only changes who wins the race.
  • How do you tell a flaky test apart from an intermittent product bug?
    Look at what actually failed in the artifacts. If the app reached a state a user could reach — a duplicate submission, a stale value, an unhandled rejection in the console — the product raced, and the test caught it. If the app was correct and the test merely observed too early or clicked the wrong node, the test is at fault. When you cannot tell, treat it as a product bug until proven otherwise.
  • What do you do about failures caused by third-party scripts on the page?
    Decide deliberately whether they are under test. Usually they are not, so block those requests in the test environment or stub them, which removes both the timing noise and the network dependency. If a third party is genuinely part of the flow being verified, keep it in one narrowly scoped test and accept that this test needs longer waits and its own reliability budget.
  • Why are artifacts from the failed attempt more valuable than a local re-run?
    Because the local re-run is a different execution — different machine speed, different data, different ordering — and it usually passes, which tells you nothing. The trace, video, console and network log capture the actual moment: which step timed out, what the DOM held, whether a request hung. Triage speed depends almost entirely on whether the pipeline retains that evidence.

saying these in an interview costs you the question

  • Adds retries and calls the problem solved
  • Fixes whichever test failed today with no data
  • Dismisses runner-only failures as just CI noise
  • Re-runs locally instead of reading the failing run's artifacts
  • Skips flaky tests permanently with no owner or expiry

context