skip to content

Your Playwright suite spends six minutes in setup projects before any test starts -- how do you decide what stays a suite-level precondition?

level: principalimportance: should knowfreq 27%

answer

  1. Work sitting on the critical path
  2. Who actually needs each step
  3. Cost multiplied by every shard
  4. One graph, several narrow projects
  5. Speed traded against reproducibility

basics

~20 s

Judge each step by who needs it, how often it must run, and whether it needs a browser. Split one broad setup project into narrow ones so a project waits only for what it lists.

solid answer

~50 s

Measure per-step duration and failure rate first, because a monolithic setup project usually hides one slow or flaky step behind several cheap ones. Then place each step by three questions: **who needs it** -- if only some projects do, give it its own setup project so only they list it in `dependencies`; **how often must it run** -- anything genuinely per-test belongs in a fixture, not at the front of the graph; and **does it need a browser** -- if not, `globalSetup` or the `webServer` option is a cheaper home. Mind the multipliers: the phase is on the critical path for every dependent project, it repeats inside every shard, and one failure there skips everything downstream. Attach a `teardown` project to anything with a lifetime, and state the target as a property -- no project waits for a precondition it does not use -- rather than a stopwatch number.

code

typescript · 11 lines
typescript
import { defineConfig } from '@playwright/test';

export default defineConfig({
  projects: [
    { name: 'flags', testMatch: /flags\.setup\.ts/, teardown: 'flags-reset' },
    { name: 'flags-reset', testMatch: /flags\.teardown\.ts/ },
    { name: 'search-index', testMatch: /search-index\.setup\.ts/ },
    { name: 'reports', dependencies: ['flags'] },
    { name: 'staff-directory', dependencies: ['flags', 'search-index'] },
  ],
});

go deeper

for a junior

Understand that anything at the front of the graph delays every test that follows, and ask which project actually needs a preparation step before adding one to it.

for a middle

Be able to explain the placement options and their frequencies: a setup project runs once per invocation, a worker fixture once per worker, a test fixture once per test.

for a senior

Diagnose with data. Measure per-step duration and failure rate, split the phase so blast radius and waiting shrink, and keep cleanup attached to whatever creates state.

for a principal

Own the line between reproducibility and feedback speed. Write down what the suite may assume about its environment, and revisit it as the suite and the pipeline change.

## Why the phase is expensive A precondition expressed as a setup project is not simply "some slow tests". Because dependent projects cannot start until it finishes, its duration is added to the front of the run for everybody, and the worker pool sits mostly idle while a small number of preparation tests execute. Three multipliers make this worse than the stopwatch suggests: - **It is on the critical path.** Six minutes of preparation is six minutes added to the wall-clock time of every dependent project, however parallel the rest of the suite is. - **It repeats per invocation.** Sharding a run across eight machines runs the graph eight times, so a six-minute phase costs forty-eight machine-minutes, not six. - **It concentrates risk.** Every project downstream is skipped when one preparation test fails, so the flakiest step in the phase sets the reliability ceiling for the whole suite. ## The questions worth asking about each step 1. **Who actually needs it?** If only the reporting projects need a warmed search index, only they should list it in `dependencies`. One monolithic setup project makes every project pay for the union of all preconditions. 2. **How often must it really run?** Once per invocation, once per worker, or once per test? Anything genuinely per-test belongs in a fixture, not in the front of the graph. 3. **Does it need a browser?** Non-browser work -- starting a stub, writing a config file, checking environment variables -- is cheaper and clearer in `globalSetup`, or is really a service the `webServer` option should own. 4. **Is it idempotent?** A step that can be skipped safely when its result already exists lets you use `--no-deps` locally and keeps reruns cheap. 5. **Who cleans up?** Anything with a lifetime needs a `teardown` project attached to the setup project, or the phase quietly becomes someone else's problem. ## Shapes to move toward | Shape | What it buys | What it costs | |---|---|---| | One setup project for everything | Simplest config; one place to look | Every project waits for every precondition; one failure skips the whole suite | | Several narrow setup projects | Each project waits only for what it lists; failures have small blast radius | More project names; the graph needs a comment explaining it | | Precondition moved into a fixture | Runs only for tests that name it, and only as often as its scope demands | Repeated per worker or per test, so it must be genuinely cheap | | Precondition moved out of the suite | Nothing on the critical path at all | The suite now depends on state it does not own, and reruns are less reproducible | ## How to argue it with a team Measure before you restructure. A per-step duration and a per-step failure rate over the last few weeks of runs usually shows that a small part of the phase accounts for most of the time or nearly all of the flakiness, and that is the part to move. Then state the target as a property, not a number -- for example "no project waits for a precondition it does not use", which the dependency graph can be checked against, rather than "setup must take under two minutes", which invites the shortcut of pushing work somewhere unmeasured. ## The tradeoff you own The honest tension is between **reproducibility** and **feedback speed**. Every precondition removed from the graph makes the run faster and makes it depend on something the run does not control; every precondition kept makes the run trustworthy and slower. A lead's job is to place that line deliberately, write down where it sits, and revisit it when the suite or the environment changes -- not to let it drift one `--no-deps` at a time. ## Signals that the phase has drifted too far - A project waits for a precondition none of its tests use. - Nobody can say what the phase leaves behind, or what removes it. - Local work is done almost exclusively with `--no-deps`, so the graph is rarely exercised outside CI. - The phase grows whenever a new suite is added, because adding to it is easier than adding a project.

  • How does sharding change the arithmetic of a slow precondition phase?
    Each shard is a separate invocation of the runner, so it executes the dependency graph itself. A six-minute phase across eight shards costs forty-eight machine-minutes and adds six minutes to every shard's wall clock, so splitting the phase pays back eight times over -- and a flaky step there now has eight chances per run to fail.
  • What target would you give the team instead of a duration for the setup phase?
    A structural property they can check: no project waits for a precondition it does not use, every precondition is idempotent, and anything with a lifetime has a teardown project. A duration target invites moving work somewhere unmeasured, while a property target survives the suite growing and can be reviewed in the config itself.

saying these in an interview costs you the question

  • Puts every precondition in one setup project by default
  • Ignores that each shard re-runs the whole graph
  • Treats --no-deps in CI as a performance fix
  • Forgets cleanup for state the phase creates
  • Optimises the phase without measuring per-step cost