skip to content

How does Playwright's --last-failed flag know which tests failed on the previous run?

level: middleimportance: should knowfreq 38%

answer

  1. The flag reads state left on disk
  2. Lives in the test output directory
  3. Fresh container means no selection
  4. Sibling flags select by git or title
  5. An empty selection fails by default

basics

~20 s

Playwright writes a .last-run.json file into the test output directory recording the run status and the ids of failed tests. The --last-failed flag reads that file, so it selects nothing useful if the directory is not preserved.

solid answer

~40 s

At the end of every run Playwright writes `.last-run.json` into the output directory (`test-results/` by default), holding the overall status and the ids of the tests that failed. `--last-failed` reads that file and runs only those tests. The consequence matters more than the mechanism: the selection is **state on disk**, so it works naturally on a developer's machine and works in CI only if the output directory survives between runs — a fresh container has no file and therefore no selection. It sits beside the other selection flags: `--only-changed` picks test files by comparing against a git ref, `--grep` and `--grep-invert` filter by title, and `--repeat-each=N` multiplies rather than narrows. All of them can select nothing, which fails the run unless you pass `--pass-with-no-tests`.

code

bash · 10 lines
bash
# Full payroll run leaves state behind in the output directory
npx playwright test
cat test-results/.last-run.json   # { "status": "failed", "failedTests": [ ... ] }

# Re-run just what was red
npx playwright test --last-failed

# Or narrow by diff, or by title tag
npx playwright test --only-changed=main
npx playwright test --grep @payroll --pass-with-no-tests

go deeper

for a junior

Know the loop it enables: run the suite, fix what broke, then re-run only the failures instead of the whole file. Remember that it depends on the previous run's output directory still being there.

for a middle

Explain the mechanism: the runner writes .last-run.json into the output directory with the failed test ids, and the flag reads it. Contrast that with --only-changed, which asks git, and --grep, which matches titles.

for a senior

Show where each flag belongs. Disk state suits a laptop loop; a diff against the base branch suits CI; a tag pattern suits a curated fast subset. Note that an empty selection fails by default, which is a feature in a pipeline.

for a principal

Own the selection strategy: what a pre-merge run is allowed to skip, what evidence a full run must still produce, and how narrowing interacts with the confidence the team claims from a green check.

Narrowing a run is the cheapest speed-up available: the fastest test is the one you had a reason not to run. Playwright ships several flags for that, and `--last-failed` is the one whose mechanism surprises people, because it depends on a file rather than on anything in your code. ## The file behind the flag Every run ends by writing `.last-run.json` into the configured output directory — `test-results/` unless you have changed `outputDir`. It records the run's overall status and the ids of the tests that failed. `--last-failed` reads exactly that file and restricts the run to those ids. Two practical consequences follow: - The selection is **state on disk**, not something derived from your source. Delete or ignore the output directory and the flag has nothing to work from. - It is per workspace, so it reflects *your last run*, including one you narrowed yourself. Chaining it after an already-filtered run narrows twice. ## Why it often selects nothing in CI On a laptop this is the natural debugging loop: run everything, fix a payroll rounding bug, re-run only what was red. In CI it usually needs help, because each job typically starts from a clean container with no `test-results/` from the previous build. Unless the directory is restored from a cache or an artefact, there is no previous run to consult and the flag selects nothing. That is a workflow decision, not a bug in the flag. ## The neighbouring selection flags | Flag | Selects by | Depends on | |---|---|---| | `--last-failed` | The previous run's failures | `.last-run.json` in the output directory | | `--only-changed[=ref]` | Test files changed against a git ref | A git repository; defaults to uncommitted changes | | `--grep` / `--grep-invert` | Test title, including tags such as `@payroll` | Nothing external — pure pattern matching | | `--repeat-each=N` | Nothing; runs each selected test N times | Nothing; it multiplies rather than narrows | `--grep` matches against the full test title, so a describe block's name and any `@tag` in the title are both matchable — which is why tag-based selection is written as a grep pattern rather than a dedicated flag. `--grep-invert` is the complement, useful for excluding a slow group from a fast pre-merge job. `--only-changed` reads git, so it does the sensible thing on a feature branch: `--only-changed=main` runs the specs that differ from the base branch. Because the runner resolves imports, a changed page object or payroll helper pulls in the specs that import it, not just files whose own text changed. ## An empty selection is a failure All of these can narrow to zero tests. By default Playwright treats "no tests found" as an error and exits non-zero, which is the right default for a pipeline: a stage that silently passes because it ran nothing is exactly the failure mode a suite exists to prevent. When an empty selection is legitimate — a shard-and-filter matrix where some combinations are genuinely empty — pass `--pass-with-no-tests`. ## Using them together 1. Reproduce locally with `--last-failed` after a red full run, so the loop is seconds rather than minutes. 2. On a branch, use `--only-changed=main` for a fast signal before the full suite is worth paying for. 3. Reach for `--grep` when the subset is a concept rather than a diff, for example a `@smoke` tag on the payroll checks that must never break. 4. Add `--repeat-each` when the question is stability rather than correctness: run the narrowed selection several times over. ## Common misreadings - Believing `--last-failed` inspects the HTML report or a junit file. It reads `.last-run.json` in the output directory. - Expecting it to work in a fresh CI container without preserving that directory. - Assuming `--only-changed` works without git, or that `--repeat-each` narrows the selection.

  • What would you have to do to make --last-failed useful in a CI job?
    Preserve the output directory between builds — cache or upload `test-results/.last-run.json` and restore it before the next run. Without that state a clean container has nothing to select from. Most teams find `--only-changed` against the base branch a better fit for CI, since git is already there.
  • Why does Playwright fail a run when a filter matches no tests?
    Because a stage that passes without executing anything looks identical to a healthy stage. Failing on an empty selection turns a typo in a grep pattern into a visible error rather than a false green. Where empty is legitimate, `--pass-with-no-tests` opts out deliberately.

saying these in an interview costs you the question

  • Says the flag parses the HTML or junit report
  • Expects it to work in a clean CI container
  • Thinks --only-changed works without git
  • Believes --repeat-each narrows the selection
  • Assumes an empty grep result passes the run