In Playwright, does a CI run end green when three tests failed once and passed on retry?
answer
- Four verdicts, not two
- Eventually green counts as green
- The report keeps every attempt
- One CLI flag flips the exit code
basics
~20 sYes. Playwright records a test that failed then passed within its retries as flaky, counts it separately in the report, and still exits with code zero, so the job goes green unless the run is started with --fail-on-flaky-tests.
solid answer
~40 sPlaywright reports four verdicts — passed, flaky, failed and skipped — and only `failed` makes the process exit non-zero. A test that failed at least once but passed within its allowed retries is **flaky**, so a run of 12 passed, 3 flaky, 0 failed exits `0` and CI is green. The flakiness is still visible: the `list` and `line` reporters print a flaky count, and the HTML report gives flaky its own filter plus one tab per attempt, each carrying that attempt's trace, video and attachments. In Playwright 1.63 you can change the contract with the `--fail-on-flaky-tests` CLI flag, which makes any flaky test produce a non-zero exit code while still running the retries and keeping their evidence.
code
bash · 2 linesnpx playwright test --retries=2 --reporter=list
npx playwright test --retries=2 --fail-on-flaky-testsgo deeper
Remember that a pass on a retry is labelled flaky, not passed, and that by default it does not turn the run red.
Explain the four verdicts and which one drives the exit code, and know that the reporters count flaky separately from passed and failed.
Show how you would stop a suite degrading silently: read the flaky count, and use --fail-on-flaky-tests when the pipeline must react rather than merely record.
Own the signal design. The exit code is the only thing most pipelines read, so decide deliberately whether the flaky bucket belongs in it and what the team is expected to do with the count.
## The verdicts a run records Playwright's reporters do not use a two-value pass/fail vocabulary. A completed run of a large payroll regression suite is summarised with four counts, and only one of them makes the process exit non-zero. | Verdict | What produced it | Exit code effect | |---|---|---| | passed | The test passed, on its first attempt | none | | flaky | The test failed at least once, then passed within its allowed retries | none by default | | failed | The test failed on every allowed attempt | non-zero | | skipped | The test was never run | none | `flaky` exists only because `retries` is greater than zero. With `retries: 0` the same unstable test is simply `failed`, because there is no second attempt in which it could pass. ## Why the run is still green Playwright's exit code answers "did every test eventually pass?", not "did every attempt pass?". So a run reporting 12 passed, 3 flaky, 0 failed exits `0`, the CI job goes green, and a pipeline that only inspects the exit code merges the change. In Playwright 1.63 the CLI flag `--fail-on-flaky-tests` changes that contract: with it, a run containing any flaky test exits non-zero, and the same three tests turn the job red. Nothing else about the run changes when that flag is set — the tests are still retried, the flaky count is still reported, and the artifacts are unchanged. The flag only decides whether the flaky bucket contributes to the exit code. ## What the reporters show - The **list** and **line** reporters print a trailing summary that names the flaky count separately from passed and failed, and mark the re-run lines with their attempt number. - The **html** reporter gives flaky its own filter chip; opening the test shows one tab per attempt, each with the trace, video, screenshots and attachments that attempt produced. - The **json** and **junit** reporters carry the per-attempt results too, so a downstream dashboard can distinguish "passed" from "passed on the third go" without re-parsing logs. - The **github** reporter annotates the pull request; a flaky test is reported as such rather than as a failure. ## Attempts are kept, not overwritten Each attempt runs with its own output directory, so the evidence of the failing attempt survives the passing one. That is what makes a flaky verdict actionable at all: opening the flaky test in the report gives you the first attempt's failure and the second attempt's success next to each other, rather than only the successful one. The order matters when reading a report: 1. Find the flaky count in the summary — it tells you how many tests needed more than one attempt. 2. Open the test and select the **first** attempt's tab; that is the one that failed. 3. Read its error and artifacts. The later attempt only tells you that the failure did not reproduce. ## Making a flaky run fail There are two ways to stop a flaky test slipping through a payroll pipeline, and they are not equivalent: - Set `retries: 0`, so the unstable test is reported `failed` outright and the run is red. This also removes the second data point that would have told you whether the failure reproduces. - Keep the retries and add `--fail-on-flaky-tests`, so the retry still runs and its evidence is still captured, but the run's exit code reflects the flaky bucket. The second keeps the diagnostic value of the retry while removing the silent-green behaviour, which is why it is the usual choice when a team wants the report but not the amnesty. ## The failure mode to watch for The reason interviewers ask this is the failure mode it produces on a large suite: with retries on and nothing inspecting the flaky count, the suite degrades quietly. Tests that fail one run in five stay green, the flaky count climbs from two to twenty over a quarter, and nobody notices because the only signal anyone watches — the exit code — never changed. The verdict vocabulary exists precisely so that signal is available; it is the pipeline's job to read it.
- What verdict does a test get when it fails on every allowed attempt?It is reported failed, not flaky, and the run exits non-zero. Flaky requires at least one failing attempt followed by a passing one; exhausting the retries without a pass is an ordinary failure with several attempts attached to it.
- Where do the first attempt's screenshots go once the test has been retried?Each attempt writes to its own output directory, so the first attempt's artifacts are not overwritten. In the HTML report they appear under that attempt's own tab, which is the one to open because it holds the failure.
saying these in an interview costs you the question
- Thinks flaky tests make the run exit non-zero
- Believes only the last attempt's artifacts survive
- Says flaky is a kind of skipped test
- Confuses flaky with failing on every attempt
- Assumes the report hides the earlier attempts