skip to content

A Playwright CI job set trace: 'on-first-retry' but a failing test produced no trace zip; how do you diagnose it?

level: seniorimportance: should knowfreq 48%

answer

  1. Walk the chain, do not guess
  2. First ask whether an attempt was traced
  3. Config resolves in layers, last one wins
  4. A dead worker finalises nothing

basics

~20 s

Check whether a retry actually ran. That mode records only during a retry, so zero retries, a run aborted before the retry, or a command-line trace override all leave a failure with no trace at all.

solid answer

~50 s

Work down the chain that has to hold for a zip to exist. First, did the retry attempt run at all — is `retries` above zero for the project that failed, and did a fail-fast limit or a timeout kill the run before the retry started? Second, is the mode you think is in force actually in force: a per-project `use`, a `test.use` in the file, or `--trace off` on the command line all override the top-level config. Third, did the worker crash or the process get killed mid-attempt, which prevents the zip from being finalised. Fourth, is the artefact simply not where you are looking — the zip lands in the attempt's folder under `test-results/` and is attached to the result, so a job that publishes only the report directory can lose it. Reproduce locally with `--trace on` to confirm capture works at all.

code

bash · 2 lines
bash
npx playwright test driver-map.spec.ts --project=order-tracker --trace on --retries=1
ls -R test-results | grep trace.zip

go deeper

for a junior

Start with the simplest link: was a retry allowed to run at all. A failure with no retries configured produces no trace under this mode, and that is normal behaviour.

for a middle

Explain how the trace option resolves across config, project, file and command line, and why a retried attempt writes its artefacts into its own folder.

for a senior

Diagnose in order rather than by guesswork, separate a broken pipeline step from a mode that never asked for capture, and say when the incident means the mode itself is wrong.

for a principal

Own the evidence chain end to end: capture settings, retry budget and artefact collection are one system, and a gap in any of them costs an investigation nobody can repeat.

"No trace" is almost never a bug in tracing. It is a chain of preconditions, and one link is broken. Diagnose it by walking the chain in order rather than guessing. ## Step 1 — Did a traced attempt ever run? `'on-first-retry'` records during the retry, so no retry means no recording. 1. Confirm `retries` is above zero **for the project that failed**, not just at the top level — a project's own value overrides the global one. 2. Confirm the run was not stopped before the retry: a `maxFailures` limit or a global timeout can end the run while the failing test is still on its first attempt. 3. Confirm the failure was a test failure at all. A worker process that is killed outright never reaches the retry logic, and a failure raised in a global setup does not belong to any test attempt. ## Step 2 — Is the mode you think is set actually set? The `trace` option resolves through several layers, each beating the last: - top-level `use` in `playwright.config.ts` - a project's own `use` - `test.use({ ... })` in the file or describe block - the command line: `npx playwright test --trace off` A CI script that appends `--trace off` for speed, or a project that was copied without its `use` block, will silently win over the config you are reading. Print the resolved configuration or add a temporary run with `--trace on` to see whether capture works at all. ## Step 3 — Was the attempt able to finish? A trace is finalised when the attempt ends. If the worker crashed, ran out of memory, or the job was cancelled, the recording never gets written out. Signals to look for: - the reporter shows the test as interrupted rather than failed; - other artefacts for that attempt are missing too, not just the trace; - the job log ends abruptly rather than with a summary. When several artefact kinds vanish together, the cause is the process, not the trace option. ## Step 4 — Are you looking in the right place? The zip is written into the failing attempt's own folder under `test-results/` and registered on the result as an attachment named `trace`. Two things follow: a retried attempt has its own folder, so the trace is not beside the first attempt's files; and a job step that uploads only a report folder, or that runs before the run finishes, can drop it. | symptom | likely link in the chain | |---|---| | no zip, and `retries: 0` in that project | no traced attempt ever ran | | zip appears locally but not on CI | a command-line or project override, or the upload step | | every artefact for the attempt is missing | the worker died before the attempt finished | | zip exists but only for some projects | the mode is set per project, not globally | ## Step 5 — Decide whether the mode is right at all Once the chain is intact, ask the design question the incident actually raised. If the failure is one that does not survive a re-run, `'on-first-retry'` will keep producing traces of healthy attempts. A retain-style mode records the original attempt and keeps it when it fails — that is the setting for evidence of the failure itself. ## Working habit Make the check cheap to repeat: - Keep a local command that reproduces with `--trace on` so you can separate "tracing is broken" from "tracing was never asked to run". - Assert the mode in review: a project without an explicit `use` block inherits, and inheritance is what most of these incidents come down to. - Treat a missing trace as a signal about the pipeline's evidence chain, not as a mystery about the tool.

  • The trace exists on CI but the job's artefacts do not contain it — where do you look?
    At the upload step's paths and its ordering. The zip is written under the attempt's folder in test-results, so a step that collects only the report directory misses it, and a step that runs before the suite finishes collects nothing. Check also that a cleanup step is not deleting the output directory between the run and the upload.
  • How would you change the configuration so the original failing attempt is always captured?
    Move that project to a retain-style mode. Recording then happens on the real attempt and the zip is kept only when it failed, so you get evidence of the failure itself rather than a re-run. Pay for it selectively — scope it to the flaky project or spec rather than turning it on for the whole suite.

saying these in an interview costs you the question

  • Assumes tracing is broken rather than never started
  • Forgets that project settings override the top level
  • Ignores a command-line trace flag in the CI script
  • Looks for the zip beside the first attempt's files
  • Expects a trace after a worker process was killed