Your Playwright payroll suite sometimes burns the whole CI hour with no verdict — how do globalTimeout and maxFailures each bound it?
answer
- Two axes: elapsed time versus failure count
- A hang produces no failures to count
- Whole-run budget, not a per-test one
- Let the runner end the run, not CI
- Leave headroom under the job limit
basics
~20 sIn Playwright, globalTimeout caps wall-clock time for the whole run and ends it whatever the results look like; maxFailures caps how many failed tests the run tolerates. A suite that hangs without failing is stopped only by globalTimeout.
solid answer
~50 sThe two controls bound different axes. `globalTimeout` is milliseconds of wall clock for the **entire run**: when it expires Playwright stops the run and fails it, regardless of how many tests passed. `maxFailures` (or `--max-failures`) bounds the **number of failed tests** before the runner gives up scheduling. A suite that is slow, stuck on a retrying navigation, or waiting on a dead payroll API produces no failures for a long time, so only the wall-clock cap ends it; a suite that fails fast and often is truncated by the failure budget long before the clock matters. Per-test timeouts do not solve the run-level problem either, since a thousand tests each just under their own limit can still exceed any job budget. In practice CI sets both, with `globalTimeout` a little under the job's own limit so Playwright, not the CI system, produces the report.
code
typescript · 8 linesimport { defineConfig } from '@playwright/test';
export default defineConfig({
// Whole-run wall clock: end the run ourselves, under the 60-minute CI job limit.
globalTimeout: 50 * 60 * 1000,
// Failure budget: off for the nightly full regression run.
maxFailures: 0,
});go deeper
Learn the two names and what each measures: one caps elapsed time for the whole run, the other caps how many tests may fail before the runner gives up. Both are configured once, for the run, not per test.
Explain why a hang is invisible to a failure budget and why per-test limits cannot bound total duration. Know that an expired run is interrupted, reported and exited non-zero, with unreached tests marked as not run.
Demonstrate the operating judgment: pick a value under the CI job limit with headroom for artefacts, base it on observed healthy duration, and use the resulting report to tell a stuck dependency apart from ordinary suite growth.
Own the budget across the estate: what a run is allowed to cost, which jobs may truncate on failures and which must stay complete, and how shard counts multiply the wall clock you thought you had capped.
A run that dies at the CI system's own limit gives you the worst artefact possible: a killed process, half-written reports, and no statement about what passed. Playwright's two run-level bounds exist so the runner ends the run itself, on its own terms, with output intact. ## Two different axes | Control | Bounds | Fires when | Unit / default | |---|---|---|---| | `globalTimeout` | Wall-clock time for the whole run | The clock expires, whatever the results | Milliseconds; `0` means no limit | | `maxFailures` | Number of failed tests | The Nth test finishes red | Count; `0` means no limit | They are not substitutes. Ask which axis the pathology lives on: - A payroll suite whose staging API is refusing connections fails fast and often — the failure budget truncates it in a minute. - A payroll suite whose API answers slowly, or whose login page never settles, produces *no* failures for a long time. Nothing about a failure count can end it. Only the wall-clock cap does. ## Why per-test timeouts are not enough Every test already has its own limit, and a hung test does eventually end. The run-level problem is arithmetic, not correctness: a suite of eight hundred tests that each take just under their per-test allowance still blows any job budget, and no single test has misbehaved. Run-level bounds are the only place that arithmetic can be capped, which is why they are configured separately from the per-test values. ## What happens when globalTimeout expires 1. Playwright stops the run at the point it has reached. 2. Tests already in flight are interrupted, and tests never reached are reported as not run. 3. The run is failed and the process exits non-zero, so no downstream step mistakes it for success. 4. Reporters still finalise, so you keep the report and the artefacts produced up to that moment. That last point is the entire argument for setting it: compare it with the CI system killing the container, where reporters never run and you are left reading raw logs. ## Choosing values for a large suite - Set `globalTimeout` **below** the CI job's own limit, with enough headroom for reporters and artefact upload. If the job allows an hour, the run should give up before then. - Base the number on the observed duration of a healthy full run plus real headroom, not on a round figure. A cap that trips on a merely busy runner trains the team to re-run rather than investigate. - Keep the failure budget separate and looser. `maxFailures` on a full regression job trades completeness for speed, so many teams leave it off there and enable it on the fast pre-merge job instead. - Remember that a sharded run is several processes: each shard carries the wall-clock cap and the failure budget on its own, so the matrix as a whole can consume a multiple of either. - Mark the handful of genuinely long tests with `test.slow()` rather than raising limits for everyone; it widens the allowance for that test alone and keeps the run-level budget meaningful. ## Diagnosing the reported case For a suite that intermittently consumes the full hour with no verdict, the sequence is: 1. Add a `globalTimeout` under the job limit so the runner, not the CI system, ends the run and emits a report. 2. Read that report: not-run tests tell you where the run stalled, and the last completed test names the neighbourhood. 3. Decide whether the pathology is one stuck test, a slow dependency shared by many tests, or genuine growth in suite size — each has a different fix, and only the report distinguishes them. ## Where teams go wrong - Treating `globalTimeout` as a per-test value, which produces a cap so small that healthy runs abort. - Expecting `maxFailures` to rescue a hang. Hangs produce no failures to count. - Setting the cap equal to the CI job limit, so the container is killed first and the report is lost anyway.
- Why set globalTimeout below the CI job's own timeout rather than equal to it?Because whoever fires first owns the outcome. If Playwright expires first it stops the run, finalises reporters and leaves you a report naming the tests that never ran. If the CI system kills the container first you get a dead process, no report and no artefacts, which is the case you were trying to diagnose.
- What does globalTimeout leave you that a killed CI container does not?A finished report. Playwright interrupts in-flight tests, marks unreached tests as not run, finalises the configured reporters and exits non-zero. That report tells you where the run stalled and what had already passed, which is exactly the evidence a hard kill destroys.
- Does a sharded run share one globalTimeout across the matrix?No. Each shard is a separate Playwright process with its own clock, so every shard gets the full allowance and the matrix can consume a multiple of it. Size the value for one shard's expected work, and reason about total spend at the pipeline level.
A wall-clock cap is the parking meter — it expires whether or not anything went wrong. A failure budget is the rule that you leave after the third ticket. Neither one covers the other's case.
saying these in an interview costs you the question
- Thinks globalTimeout is a per-test timeout
- Expects a failure budget to stop a hanging run
- Sets the cap equal to the CI job limit
- Believes per-test timeouts bound total run time
- Assumes the shards share one run-level clock
- Thinks an expired run still exits zero