Playwright discards a worker process after any test failure - what does that guarantee cost a large suite?
answer
- Isolation bought with throughput
- Cost scales with failures, not workers
- Cascading failures are the thing avoided
- Restart price equals worker start-up cost
- No switch to reuse a failed process
basics
~20 sIt buys certainty that no test inherits a process a failure left dirty. It costs a cold browser launch plus per-worker set-up for every failure, so a red payroll run is measurably slower than a green one.
solid answer
~50 sThe guarantee is that a failure can never propagate: the next test starts in a process the failing test never touched, so a single broken test cannot quietly poison the ten that follow. That is worth a great deal in a large payroll suite, where a cascading failure costs hours of misdirected investigation. The cost is throughput. Each failure buys another browser launch and another run of whatever per-worker set-up you have, so runs get slower exactly when they are already red and people are waiting. The lever is not a switch to disable the restart — the worker options size the pool, not the reuse policy. It is to make per-worker start-up cheap and idempotent, to move expensive seeding to a one-off step before the run, and to treat a high failure count as the thing to fix.
code
typescript · 9 linesimport { defineConfig } from '@playwright/test';
export default defineConfig({
testDir: './tests/payroll',
workers: 6,
// Every failure throws its worker away, so a red run pays worker start-up
// again. Seed the payroll tenants once here, not inside each worker.
globalSetup: './seed-payroll-tenants.ts',
});go deeper
Understand the shape of the deal: a failed test costs the whole process, and the run restarts one to keep later tests clean. Speed is being traded for trustworthy results.
Explain what a restart re-does — process start, browser launch, per-worker set-up — and that the total scales with the number of failures rather than with the number of workers.
Show that you have measured it: time from worker start to first test, multiplied by the restarts a bad day produces, and the work you moved out of the worker to shrink it.
Own the trade explicitly. Cap what per-worker start-up may cost, budget the pipeline for a red day rather than a green one, and refuse both non-existent workarounds and worker inflation as answers.
## What the guarantee actually buys Discarding a worker after a failure removes an entire class of investigation from your life. When a test fails, you know that the following tests ran in a process it never touched — not a cleaned-up process, a **different** one. That rules out, by construction: - A leaked event listener from the failing test firing during a later one. - A half-mutated module-level cache serving stale payroll data to everything after it. - A browser left in a state no one designed, quietly changing later behaviour. - The worst failure mode of all: **a cascade**, where one real defect produces thirty red tests and the report no longer tells you how many problems you have. In a large regression suite the value of this is asymmetric. Throughput costs are measured in minutes. A cascade costs a person a day, and it costs the team's trust in the suite, which is far more expensive to rebuild than it is to lose. ## What it costs Each restart puts real work back on the critical path: 1. **Process start-up** — a new Node process importing your test files again. 2. **Browser launch** — a cold browser, which is the largest fixed cost in most Playwright runs. 3. **Per-worker set-up** — everything a worker does before its first test: signing in, resolving configuration, claiming a payroll tenant, warming a client. 4. **Lost warmth** — caches and connections the discarded process had built up, all rebuilt from nothing. That cost is multiplied by the failure count, not the worker count. A run with two failures never notices. A run with eighty failures pays it eighty times, which is why a badly broken pipeline often appears to have got slower as well as redder — and why "the suite is slow" complaints frequently trace back to a suite that is simply failing a lot. ## Where the levers actually are | Lever | Effect | Watch out for | |---|---|---| | Cheap per-worker set-up | Cuts the price of every restart | Set-up must be idempotent; it will rerun | | Seed data once before the run | Removes seeding from the restart path | Shared data must stay isolated per worker | | Drive the failure count down | Fewer restarts to pay for | The real fix, and the slowest to deliver | | Cap a hopeless run early | Stops paying for a run already lost | A separate setting from the worker count | | More workers | Hides latency, does not remove cost | Oversubscription creates new failures | Note what is missing from that table: a switch that keeps a failed worker alive. The worker options size the pool; they do not make the runner reuse a process that has just failed a test. Designing around the guarantee is the only available strategy. ## The judgment call to make explicit The question a lead actually has to answer is **how much per-worker set-up the suite can afford**, because that number is the multiplier on every failure. A payroll suite whose worker start-up signs in through the UI and seeds a company can pay a minute per restart; the same suite whose workers claim a pre-seeded tenant and reuse a saved session pays a second. The two behave identically when everything is green and diverge catastrophically on the day something is broken — which is precisely the day the feedback loop matters most. - **Measure the restart cost** rather than estimating it: time from process start to first test, and multiply by a realistic failure count for a bad day. - **Budget for the bad day**, not the good one. A pipeline sized only for green runs becomes unusable exactly when it is needed. - **Push expensive work out of the worker** and into a one-off step before the run, so a restart re-does the cheap part only. - **Treat a climbing worker index as a metric**, not trivia: it is a direct count of how much of a run went into re-launching browsers. ## What not to argue Do not propose reusing the process to save time; you would be trading a bounded, visible cost for an unbounded, invisible one. And do not compensate by raising the worker count — that shortens the run only until oversubscription starts producing new failures, each of which buys another restart. The correct escalation, once set-up is already cheap, is to spend fewer failures, not more processes.
- How would you quantify what worker restarts cost a given payroll run?Measure the interval between a worker process starting and its first test beginning, then multiply by the number of restarts, which the highest worker index tells you. Compare that total with the run's wall-clock time; if it is a meaningful share, the payback is in cheaper worker start-up rather than in more workers.
- A team asks to keep failed workers alive to speed up red runs. What do you tell them?That the runner does not offer it, and that the trade is bad anyway: you would swap a visible, bounded cost for cascading failures that are unbounded and hard to attribute. Point them at the levers that do exist — cheaper worker start-up, seeding before the run, and fixing what fails.
- When is the right answer to accept slow red runs rather than optimise them?When failures are genuinely rare. Restart cost only bites at volume, so a suite that fails twice a week gains nothing from optimisation and should spend the effort elsewhere. Revisit the moment the failure rate rises, because the same suite then pays that cost dozens of times a day.
saying these in an interview costs you the question
- Looks for a setting to reuse failed workers
- Claims restarts are free because tests are short
- Adds workers to hide restart cost
- Ignores that per-worker set-up reruns on failure
- Assumes cascading failures are easy to spot
- Optimises restart cost before measuring it