skip to content

Your Playwright suite runs with retries: 2 and trace 'on-first-retry', yet flaky tests only leave traces of passing runs — why?

level: seniorimportance: should knowfreq 40%

answer

  1. Recording starts later than you think
  2. The traced attempt is the one that passed
  3. Record the first run, keep on failure
  4. Video and screenshots follow the same rule

basics

~10 s

Because on-first-retry starts recording on the second attempt, not the first: the run that actually failed was never traced. Switching to retain-on-first-failure records the original run and keeps it only when it failed.

solid answer

~40 s

Each `trace` value picks which **attempt** is recorded. `'on-first-retry'` records the first retry — attempt two — so when the retry passes, the only artifact you have describes a clean run, and the transient condition that broke attempt one is gone. In Playwright 1.63 the fix is `trace: 'retain-on-first-failure'`, which records the first run of every test and keeps the recording only when that run failed; `'retain-on-failure'` also works and additionally covers the retries, at the cost of recording every attempt before discarding the passing ones. `'on-all-retries'` does not help, because it still skips the original run. The same attempt-selection logic governs `video: 'retain-on-failure'` and `screenshot: 'only-on-failure'`, and because each attempt has its own output directory, the failing attempt's artifacts are never overwritten by the passing one.

code

typescript · 10 lines
typescript
import { defineConfig } from '@playwright/test';

export default defineConfig({
  retries: 2,
  use: {
    trace: 'retain-on-first-failure',
    video: 'retain-on-failure',
    screenshot: 'only-on-failure',
  },
});

go deeper

for a junior

Recall that the trace option chooses which attempt is recorded, and that on-first-retry means the retry rather than the run that failed.

for a middle

Explain the modes side by side: which attempts each records, and which of them discard the recordings of attempts that passed.

for a senior

Diagnose it end to end. Walk the attempt timeline, name retain-on-first-failure or retain-on-failure as the fix, and say why the retry's clean trace misled the reader.

for a principal

Own the trade between evidence and cost: recording the first run of every test slows the suite and grows the artifact store, and the alternative is a retry policy that produces flaky counts nobody can act on.

## Which attempt each capture mode records The `trace` option in the `use` block does not simply mean "record traces". Each value picks a different **attempt** to record, and picking the wrong one is why a flaky test can leave behind a trace of a run that passed. In Playwright 1.63 the retry-relevant values are: | `trace` value | Attempts recorded | What is kept | |---|---|---| | `'off'` | none | nothing | | `'on'` | every attempt of every test | everything, pass or fail | | `'retain-on-failure'` | every attempt of every test | only attempts that failed | | `'retain-on-first-failure'` | the first run of each test only | that run, only if it failed | | `'on-first-retry'` | the first retry only (attempt two) | that attempt, pass or fail | | `'on-all-retries'` | every retry (attempts two onward) | those attempts | `'on-first-retry'` is popular because it is cheap: on a suite where almost everything passes first time, almost nothing is recorded. Its cost is written into its own name — it starts recording **after** a failure has already happened. ## Why the traced attempt is the wrong one Follow one payroll test through a run configured with `retries: 2` and `trace: 'on-first-retry'`: 1. Attempt one runs untraced, fails on a `expect(locator).toBeVisible()` timeout, and is recorded as a failed attempt with an error message and stack. 2. Attempt two starts, and because it is the first retry, tracing is switched on. 3. Attempt two passes. The test is reported **flaky**. 4. The report offers exactly one trace — belonging to attempt two, the run that worked. Nothing is broken; the mode did what it says. But the artifact answers the wrong question. The transient condition — the slow payroll aggregation, the late-arriving XHR, the animation that moved the button — existed during attempt one and is gone by attempt two. The trace shows a clean run, so a reader concludes "cannot reproduce" and the flake stays in the suite. ## Fixing it Switch the mode so the **first** run is the one recorded: - `trace: 'retain-on-first-failure'` records the first run of every test and keeps that recording only when the run failed. Passing tests leave nothing behind, so the artifact store stays small, and every failed first attempt has a trace. - `trace: 'retain-on-failure'` also keeps the failing first attempt, and additionally records the retries; it costs more because every attempt of every test is recorded before the passing ones are discarded. - `trace: 'on'` records everything unconditionally and is only sensible on a small suite or a targeted debugging run. - `trace: 'on-all-retries'` does **not** fix this. It still skips the first run and records attempts two and three, which helps when the failure reproduces and not at all when the original failure is the only one. ## What the fix costs Recording the first run of every test is not free, and on a large regression suite the difference is visible: - tracing adds overhead to every test, including the vast majority that pass and whose recording is then thrown away; - disk and upload time grow with the number of failures retained, not the number of tests, so the steady-state cost is modest but the bad-day cost is not; - `'retain-on-first-failure'` deliberately does not record the retries, so if you also want to compare the failing and passing attempts side by side, `'retain-on-failure'` is the mode that gives you both. ## Video and screenshots follow the same shape The same attempt-selection logic governs the other artifacts, so a config fixed only for traces still loses the other evidence: - `video: 'retain-on-failure'` records every test and deletes the recordings for attempts that passed, so the failing first attempt keeps its video; - `video: 'on-first-retry'` has exactly the problem described above, one artifact type over; - `screenshot: 'only-on-failure'` captures at the end of any attempt that failed, including the first. Because each attempt writes into its own output directory, none of these overwrite each other: the failing attempt's video and the passing attempt's video coexist, and the report shows them under separate tabs. The rule of thumb is simply that the capture mode must name a condition — failure — rather than an attempt number, unless you are certain the failure repeats.

  • Would switching the trace mode to 'on-all-retries' solve this?
    No. It records attempts two and three but still skips the original run, so a failure that does not reproduce is still untraced. It helps only when the failure repeats across retries, which is precisely the case that was never hard to diagnose.
  • Why not simply set trace: 'on' for the whole payroll suite?
    It records every attempt of every test and keeps them all, so the overhead is paid by the thousands of tests that pass and the artifacts pile up. `'retain-on-first-failure'` gives the same diagnostic reach for failures while discarding the recordings nobody will open.

saying these in an interview costs you the question

  • Thinks on-first-retry records the first attempt
  • Assumes the retry reproduces the original failure
  • Believes later attempts overwrite earlier artifacts
  • Says retain-on-failure keeps traces for passing tests
  • Expects a flaky test's trace to show the failure