skip to content

Playwright's default worker count makes your payroll suite time out in a CPU-limited CI container - how do you diagnose and fix it?

level: seniorimportance: should knowfreq 36%

answer

  1. CI-only timeouts, scattered, not reproducible
  2. More workers made it slower
  3. Reported cores are not the container's allowance
  4. Confirm the count from the run header
  5. Pin workers rather than raising timeouts

basics

~20 s

The default is half the cores the operating system reports, which is usually the host's, not the container's CPU allowance. Read the worker count from the run header, compare it with the container limit, then pin workers explicitly.

solid answer

~50 s

Start by confirming the count rather than guessing it: the `list` reporter's header line names how many workers the run chose. Compare that with what the container is actually allowed — a two-CPU container on a thirty-two-core host still reports the host's cores, so Playwright's default of half the logical cores can start sixteen workers and sixteen browsers on two CPUs. The symptoms match: unrelated tests time out, failures are spread thinly across the suite rather than clustered on one feature, nothing reproduces locally, and the run is slower than a smaller worker count would be. Confirm by re-running with `--workers=2`; if the timeouts vanish, it was oversubscription, not the payroll application. Fix it by pinning `workers` in the config for CI, sizing it to the container's CPU and memory allowance rather than to the reported core count.

code

bash · 8 lines
bash
# What the run actually chose - the header names tests and workers
npx playwright test --reporter=list 2>&1 | grep -m1 -i 'workers'

# What the container is really allowed (cgroup v2 CPU quota, in microseconds)
cat /sys/fs/cgroup/cpu.max

# Re-run the same payroll suite at a count the container can afford
npx playwright test -j 2

go deeper

for a junior

Know that the number of workers is a setting you can see and change, and that a run header tells you how many were used. If CI times out and your machine does not, ask about the worker count.

for a middle

Explain why a container is the trap: the default comes from the reported logical cores, and a CPU quota does not lower that number, so the runner can start far more browsers than the container can drive.

for a senior

Walk the diagnosis in order — read the header, compare with the container limits, halve and re-run, check memory and the backend — and pin the number rather than reaching for a longer timeout.

for a principal

Own the policy across pipelines: where the count is declared, how it is sized against CPU, memory and shared environment capacity, and what review the number gets when runners change shape.

## The symptom pattern Oversubscription has a recognisable signature, and it is worth learning because it looks like application flakiness: - Failures are **timeouts**, not assertion mismatches — actions and navigations that expire rather than see the wrong value. - They are **scattered** across unrelated payroll features instead of clustering on one screen or endpoint. - They **do not reproduce locally**, and often not even in the same container when the suite is re-run alone. - Adding more workers makes things **worse**; the run gets slower as well as redder. - Workers occasionally vanish outright, which is the kernel reclaiming memory rather than any test failing. ## The cause Playwright's default worker count is derived from the machine's **logical CPU cores as reported by the operating system**. Inside a container that reporting is normally the host's core count: a CPU quota restricts how much CPU time the container may consume, but it does not change the number of cores the process sees. On a thirty-two-core build host, a container limited to two CPUs still reports thirty-two cores, so the default starts sixteen workers — and sixteen browsers — into an allowance of two. Every worker is a full browser as well as a Node process, so oversubscription is doubly expensive: CPU time is sliced sixteen ways, and memory is multiplied sixteen times. Playwright's waits are generous but not infinite, so a click that normally resolves in fifty milliseconds eventually exceeds the timeout, and the test fails with no defect anywhere in the payroll application. ## Diagnosing it in order 1. **Read the header.** Run with the `list` reporter and look at the opening line, which names the test count and the worker count actually in force. Do not infer it from the config; a stray `-j` in a CI script overrides the file. 2. **Read the container's limits.** Compare that worker count with the CPU allowance and the memory limit of the job, not with the core count the host advertises. 3. **Halve it and re-run.** Re-run the same suite with `--workers=2`. If the timeouts disappear, the finding is oversubscription; if the same tests fail identically, look at the application instead. 4. **Check memory.** If workers disappear without reporting failures, you are hitting the memory limit rather than the CPU quota, and the fix is the same lever pushed further. 5. **Look at the far side.** Confirm the payroll backend is not itself the bottleneck; a shared test environment can be the thing that slows under sixteen concurrent sessions. ## Fixing it | Fix | When it is right | |---|---| | Pin `workers` for CI in the config | The default answer: state the number the container can afford | | Pass `--workers` in the pipeline | The count differs per job and belongs with the job definition | | Express it as a share of cores | Only where the reported cores match the real allowance | | Raise the timeout | Almost never — it hides oversubscription and slows every failure | Pinning the number is the fix that matches the cause. Raising timeouts is the fix that looks appealing at eleven at night and turns a fifteen-minute red run into a fifty-minute red run, because every genuinely broken test now waits longer before giving up. ## Sizing the pinned number Start from the container's CPU allowance rather than the host's cores, and expect roughly one worker per available CPU as an upper bound, fewer if the tests are heavy on rendering or if the payroll backend is a shared environment. Then validate empirically: run the suite at two or three candidate counts and compare wall-clock time and failure count together. The fastest configuration is not the highest count — it is the highest count that has not yet started producing timeouts. ## Keeping it fixed Record the number in the repository with a comment explaining what it is sized to, so the next person who moves the job onto a larger runner knows to revisit it. The failure mode is quiet: nothing warns you that a config tuned for a two-CPU container is now running on eight, or the reverse. And whenever a CI-only wave of timeouts appears after an infrastructure change, check the worker count before opening a bug against the payroll application.

  • Why not simply raise the test timeout until the CI failures stop?
    Because it treats the symptom. Oversubscription means every action is slow, so a higher timeout keeps the run red-adjacent and much slower: each genuinely failing test now burns the longer budget before reporting. It also destroys the timeout's value as a signal, since nothing distinguishes a wedged page from a merely starved one.
  • How would you prove the payroll backend, rather than CPU, is the constraint?
    Hold the worker count fixed and watch the backend: response times and error rates rising in step with concurrency point at the service, while flat backend latency alongside stretched local action times points at CPU. Running the same worker count against a private environment, where the timeouts vanish, settles it.
  • What is the first thing to check when CI timeouts appear right after a runner upgrade?
    The worker count in the run header. A larger host advertises more cores, so a config that relies on the default silently starts more workers than before while the container's CPU allowance may not have moved at all. That is oversubscription introduced by infrastructure, with no change in the suite.

saying these in an interview costs you the question

  • Blames flaky tests without checking the worker count
  • Raises timeouts to make CI green again
  • Assumes a container reports its own CPU quota
  • Thinks more workers always shortens the run
  • Ignores memory limits when sizing workers
  • Tunes worker count without measuring wall-clock time