Enabling Playwright's fullyParallel made one payroll spec fail intermittently in CI — how do you contain it?
answer
- Reproduce under both schedules first
- Attribute before you configure
- Smallest scope that holds
- Weakest mode that fixes it
- Label the quarantine as debt
basics
~20 sConfirm the schedule is the trigger by rerunning that spec with and without test-level parallelism, then apply the narrowest override that holds: default mode for the file, or a serial describe around only the dependent tests.
solid answer
~40 sFirst establish that scheduling is the cause rather than a coincidence: run the suspect spec alone with `--fully-parallel` and again without it, and repeat each run a few times, since a genuine ordering problem reproduces under one schedule and not the other. Then contain it at the smallest scope that works. Adding `test.describe.configure({ mode: 'default' })` at the top of the file restores ordered execution on one worker for that file only; wrapping just the two or three genuinely dependent tests in a `test.describe` with `mode: 'serial'` is tighter still and leaves the other tests in the file parallel. Prefer the weaker override, record it as debt with the reason in a comment, and keep the rest of the suite on the fast schedule while the underlying collision is fixed.
code
typescript · 23 linesimport { test, expect } from '@playwright/test';
// Rest of the file still runs fully parallel.
test('payslip list loads', async ({ page }) => {
await page.goto('/payroll/payslips');
await expect(page.getByRole('heading', { name: 'Payslips' })).toBeVisible();
});
// Quarantine: these two share the export queue (PAY-4821).
test.describe('bank export', () => {
test.describe.configure({ mode: 'serial' });
test('queues the export', async ({ page }) => {
await page.goto('/payroll/exports');
await page.getByRole('button', { name: 'Queue export' }).click();
await expect(page.getByText('Queued')).toBeVisible();
});
test('downloads the export file', async ({ page }) => {
await page.goto('/payroll/exports');
await expect(page.getByRole('link', { name: 'Download' })).toBeVisible();
});
});go deeper
Know that the fix belongs in the file or block, using a describe mode, rather than in a global setting that slows the whole run.
Be able to compare the containment widths and say what each one costs, and why default mode is often enough where serial is reached for.
Show the method: reproduce under both schedules, attribute, contain narrowly, and record the override as tracked debt with a route back to parallel.
Own the rollout policy, including who is allowed to add a quarantine line, how the list is reviewed, and what the suite's speed target is worth defending.
## Confirm the schedule is the cause An intermittent failure that appeared alongside a `fullyParallel` rollout is *probably* an ordering or collision problem, but probably is not a diagnosis. Cheap steps that settle it: 1. Run the one spec file on its own with `--fully-parallel`, several times. If it fails some of those runs, the schedule reproduces the problem in isolation. 2. Run the same file without the flag, the same number of times. Consistent green here is the contrast that implicates scheduling. 3. Note **which** tests fail and which pass. A failure that always lands on the same test, and mentions state a sibling test creates, points at a specific pair rather than the whole file. The artefacts your reporters already collect will tell you what the failing test saw; the point of the two-schedule comparison is to attribute it, not to explore it. ## Choose the narrowest override that works Playwright gives you three widths of containment, and the instinct to reach for the widest is what makes suites slow again: | containment | scope affected | what it costs | |---|---|---| | Turn `fullyParallel` off in the config | every file in the run | the entire speed-up | | `test.describe.configure({ mode: 'default' })` at file top | one spec file | that file's internal concurrency | | `test.describe` with `mode: 'serial'` | one block of tests | that block only, plus dependency semantics | Start at the bottom row and move up only if it does not hold. A payroll suite that reverts the global flag because one export spec is unhappy has traded a forty-file win for a one-file problem. ## Default or serial for that scope Both pin the scope to one ordered worker, so both fix a pure race. Pick between them on what you are claiming: - If the tests merely must not overlap, use `mode: 'default'`. Every test still runs on every attempt and still reports on its own merits, so you keep the diagnostic value of the file. - If a later test genuinely cannot mean anything when an earlier one failed, use `mode: 'serial'`. Accept that the tail is skipped on failure and that a retry replays the whole block. Over-claiming with `serial` when `default` would do costs you information: the skipped tail hides whatever else was broken that run. ## Make the containment visible - Put the reason and a ticket reference in a comment on the configure line, so the next reader knows it is a quarantine and not a design choice. - Keep the serial span to the smallest set of tests that actually depend on each other; every extra test in the block is time paid again on each retry. - Track the list of quarantined scopes somewhere the team looks. Lines like this are invisible in a green build and survive for years unless someone counts them. - Re-check periodically by deleting the line and running the file a few times under the parallel schedule; the fix that landed last month may have already made it unnecessary. ## What not to do - **Do not raise retries to paper over it.** A retry hides the symptom and leaves the collision live for whichever test loses the race next time. - **Do not squeeze the run down to a single worker** to make the symptom vanish. That is the global revert wearing a different hat, and it slows every unrelated file. - **Do not mark the whole file serial** when two tests are the problem; the rest of the file loses concurrency and gains skip-on-failure it never needed. ## How to present the answer Interviewers are listening for a method rather than a setting: reproduce under one schedule and not the other, contain at the smallest scope, choose the weakest mode that holds, label it as temporary, and keep the global speed-up intact. Naming the exact API — file-level `mode: 'default'` versus a block-level `mode: 'serial'` — is what turns that method into a concrete plan.
- How would you prove the failure is scheduling rather than an unrelated flaky assertion?Run the file repeatedly under both schedules. A scheduling problem shows a clear asymmetry: failures appear with `--fully-parallel` and vanish without it. An assertion that is flaky on its own fails at a similar rate under both, which points somewhere else entirely.
- Why not just set fullyParallel to false until the spec is fixed?Because it charges the whole suite for one file. The override mechanisms are scoped precisely so one unhappy spec can be quarantined while the other files keep test-level scheduling. Reverting globally also removes the pressure that would surface the next collision.
- When is marking the group serial the right permanent answer rather than a quarantine?When the block models a genuinely sequential business flow whose steps are not meaningful alone, and the team accepts skip-on-failure and whole-group retries as honest reporting for it. Even then, keep the block short so the retry cost stays small.
saying these in an interview costs you the question
- Reverts the global flag for one bad file
- Adds retries instead of finding the collision
- Marks the whole file serial without checking scope
- Runs the suite on one worker to hide the symptom
- Leaves the quarantine line unlabelled and permanent