A Playwright CI job set trace: 'on-first-retry' but a failing test produced no trace zip; how do you diagnose it?
answer
- Walk the chain, do not guess
- First ask whether an attempt was traced
- Config resolves in layers, last one wins
- A dead worker finalises nothing
basics
~20 sCheck whether a retry actually ran. That mode records only during a retry, so zero retries, a run aborted before the retry, or a command-line trace override all leave a failure with no trace at all.
solid answer
~50 sWork down the chain that has to hold for a zip to exist. First, did the retry attempt run at all — is `retries` above zero for the project that failed, and did a fail-fast limit or a timeout kill the run before the retry started? Second, is the mode you think is in force actually in force: a per-project `use`, a `test.use` in the file, or `--trace off` on the command line all override the top-level config. Third, did the worker crash or the process get killed mid-attempt, which prevents the zip from being finalised. Fourth, is the artefact simply not where you are looking — the zip lands in the attempt's folder under `test-results/` and is attached to the result, so a job that publishes only the report directory can lose it. Reproduce locally with `--trace on` to confirm capture works at all.
code
bash · 2 linesnpx playwright test driver-map.spec.ts --project=order-tracker --trace on --retries=1
ls -R test-results | grep trace.zipgo deeper
Start with the simplest link: was a retry allowed to run at all. A failure with no retries configured produces no trace under this mode, and that is normal behaviour.
Explain how the trace option resolves across config, project, file and command line, and why a retried attempt writes its artefacts into its own folder.
Diagnose in order rather than by guesswork, separate a broken pipeline step from a mode that never asked for capture, and say when the incident means the mode itself is wrong.
Own the evidence chain end to end: capture settings, retry budget and artefact collection are one system, and a gap in any of them costs an investigation nobody can repeat.
"No trace" is almost never a bug in tracing. It is a chain of preconditions, and one link is broken. Diagnose it by walking the chain in order rather than guessing. ## Step 1 — Did a traced attempt ever run? `'on-first-retry'` records during the retry, so no retry means no recording. 1. Confirm `retries` is above zero **for the project that failed**, not just at the top level — a project's own value overrides the global one. 2. Confirm the run was not stopped before the retry: a `maxFailures` limit or a global timeout can end the run while the failing test is still on its first attempt. 3. Confirm the failure was a test failure at all. A worker process that is killed outright never reaches the retry logic, and a failure raised in a global setup does not belong to any test attempt. ## Step 2 — Is the mode you think is set actually set? The `trace` option resolves through several layers, each beating the last: - top-level `use` in `playwright.config.ts` - a project's own `use` - `test.use({ ... })` in the file or describe block - the command line: `npx playwright test --trace off` A CI script that appends `--trace off` for speed, or a project that was copied without its `use` block, will silently win over the config you are reading. Print the resolved configuration or add a temporary run with `--trace on` to see whether capture works at all. ## Step 3 — Was the attempt able to finish? A trace is finalised when the attempt ends. If the worker crashed, ran out of memory, or the job was cancelled, the recording never gets written out. Signals to look for: - the reporter shows the test as interrupted rather than failed; - other artefacts for that attempt are missing too, not just the trace; - the job log ends abruptly rather than with a summary. When several artefact kinds vanish together, the cause is the process, not the trace option. ## Step 4 — Are you looking in the right place? The zip is written into the failing attempt's own folder under `test-results/` and registered on the result as an attachment named `trace`. Two things follow: a retried attempt has its own folder, so the trace is not beside the first attempt's files; and a job step that uploads only a report folder, or that runs before the run finishes, can drop it. | symptom | likely link in the chain | |---|---| | no zip, and `retries: 0` in that project | no traced attempt ever ran | | zip appears locally but not on CI | a command-line or project override, or the upload step | | every artefact for the attempt is missing | the worker died before the attempt finished | | zip exists but only for some projects | the mode is set per project, not globally | ## Step 5 — Decide whether the mode is right at all Once the chain is intact, ask the design question the incident actually raised. If the failure is one that does not survive a re-run, `'on-first-retry'` will keep producing traces of healthy attempts. A retain-style mode records the original attempt and keeps it when it fails — that is the setting for evidence of the failure itself. ## Working habit Make the check cheap to repeat: - Keep a local command that reproduces with `--trace on` so you can separate "tracing is broken" from "tracing was never asked to run". - Assert the mode in review: a project without an explicit `use` block inherits, and inheritance is what most of these incidents come down to. - Treat a missing trace as a signal about the pipeline's evidence chain, not as a mystery about the tool.
- The trace exists on CI but the job's artefacts do not contain it — where do you look?At the upload step's paths and its ordering. The zip is written under the attempt's folder in test-results, so a step that collects only the report directory misses it, and a step that runs before the suite finishes collects nothing. Check also that a cleanup step is not deleting the output directory between the run and the upload.
- How would you change the configuration so the original failing attempt is always captured?Move that project to a retain-style mode. Recording then happens on the real attempt and the zip is kept only when it failed, so you get evidence of the failure itself rather than a re-run. Pay for it selectively — scope it to the flaky project or spec rather than turning it on for the whole suite.
saying these in an interview costs you the question
- Assumes tracing is broken rather than never started
- Forgets that project settings override the top level
- Ignores a command-line trace flag in the CI script
- Looks for the zip beside the first attempt's files
- Expects a trace after a worker process was killed