An end-to-end suite runs eight workers in parallel against one shared staging environment. Giving each test its own freshly created records fixed most of the interference, but a handful of tests still fail only when the whole suite runs together. What kinds of state cannot be made unique per test, and what do you do about them?
answer
- some things exist exactly once
- widen the unit before serialising
- wrong value, not missing element
- one inbox, eight readers
- a small serial lane is fine
basics
~20 sSingleton state resists per-test uniqueness: application-wide feature flags and settings, limited inventory, rate limits, a shared mailbox, background jobs, caches and the clock. Partition what you can into a per-worker tenant or account, and run the few tests that must mutate genuinely global state serially.
solid answer
~50 sPer-test records solve interference only for state that can be duplicated. What remains is state the system has exactly one of: a feature flag or admin setting that changes behaviour for everyone, a shared inventory or seat count, a per-IP rate limit, one email inbox all tests read, a scheduled job that runs over everything, a cache or search index everybody shares, and the environment's clock. The first move is to widen the unit of isolation — give each worker its own tenant, organisation or account so most "global" settings become per-tenant and can genuinely be duplicated. Whatever is still global after that gets a serial lane: a small group of tests that runs alone, with the parallel lane blocked or those tests scheduled apart. Accept that lane being slower; it is a handful of tests. For shared inboxes, per-test plus-addressing gives each test its own address inside one mailbox.
go deeper
Understand that some state exists only once for the whole application — a global setting, one shared inbox — so creating fresh records per test cannot isolate it, and those tests need special handling.
Name the usual singletons — application-wide flags, limited inventory, rate limits, a shared mailbox, background jobs, caches, the clock — and explain why they surface as failures that appear only when the full suite runs.
Show the ordering of fixes: widen the isolation unit to a per-worker tenant or account first, use plus-addressed inboxes and per-identity rate allowances, and only then put the genuinely global cases into a tagged serial lane.
Decide whether the product should grow the isolation seam the suite needs — per-tenant settings, a scoped test identity, provisionable environments — and weigh that platform investment against permanently carrying a serial lane and its runtime.
## The category the record-per-test rule misses Generating fresh records fixes interference for anything the system can hold many of: users, orders, documents, projects. It does nothing for state the system has **exactly one of**, because there is nothing to make unique. Recognising which of your failures is in this category is most of the work — otherwise the symptom (fails only under load, passes when re-run alone) looks exactly like a timing problem and gets a wait added to it. ## The usual singletons - **Application-wide flags and settings.** A feature flag, a maintenance banner, a global tax rate, an admin toggle. One test flips it and every concurrent test is now running a different application. - **Constrained resources.** Limited stock, a seat count on a plan, a licence pool, a single "active" record where only one may exist at a time. Two tests both take the last one. - **Rate limits and quotas.** Eight workers hammering login or search from one IP hit a per-IP or per-account limit that never triggers when a developer runs one test. - **A shared mailbox or SMS endpoint.** Every test that verifies a signup email reads the same inbox and grabs someone else's message. - **Background jobs and schedulers.** A nightly digest or a queue worker processes everything in the environment, including records another test is mid-way through asserting on. - **Caches, search indexes and projections.** Shared, eventually consistent, and warmed by whatever ran last. - **The clock.** Anything that freezes or advances the environment's time affects every worker at once — this is why time manipulation belongs at a lower test level. - **Third-party sandboxes.** One sandbox account with one set of test cards and one webhook stream. ## Fix one: widen the isolation unit Before serialising anything, ask whether the "global" thing can be scoped. In a multi-tenant product, most settings are per-tenant, so giving each worker its own organisation converts a singleton into a duplicable record. Even in a single-tenant product, many settings turn out to be per-account or per-project. This is the highest-leverage move, because it keeps everything in the parallel lane. Concretely: allocate one organisation (and the data inside it) per worker at the start of the run, and let every test in that worker create inside it. For the mailbox, plus-addressing usually works: `[email protected]` routes to one mailbox but gives each test an address only it will match on, so tests never compete for the same message. ## Fix two: a serial lane for what is genuinely global Whatever remains — a truly application-wide flag, an environment-wide setting — gets an explicit serial group. The important part is that this is a **deliberate, named** arrangement, not an accident: those tests are tagged, they run in a lane that does not overlap the parallel lane, and the tag says why. A handful of serial tests costs a minute; the alternative is a suite nobody trusts. A lighter variant is a lock: a test acquires an exclusive lease on the flag, mutates, asserts and releases. That preserves parallelism for everything else, but it introduces its own failure mode — a crashed test holding a lease — so the lease needs a timeout. ## Fix three: move the scenario down a level Some of these scenarios do not belong in an end-to-end suite at all. Verifying that the UI renders differently when a feature flag is on is a component-level concern where you control the input directly and no shared world exists. Reserve the end-to-end version for the case where the flag's real propagation path is the thing under test — and then accept it is serial. ## Diagnosing which one you have The signature of singleton interference is: fails only when the full suite runs, passes in isolation, and the failure is a *wrong value* rather than a missing element — the page rendered fine, it just rendered someone else's world. When you see that, resist adding a wait; instead ask what the two tests both touched that only one of exists. ## Rate limits deserve a special note Parallel workers concentrate traffic from few addresses, so a limit sized for humans trips immediately. The right response is usually an environment-level allowance for the test identity rather than slowing the suite down — and if you disable the limit entirely in staging, remember you have also stopped testing it, so it needs coverage somewhere else.
- How do you tell singleton interference apart from a timing problem?Look at the failure shape. Timing failures are missing or not-yet-rendered elements; singleton interference renders fine but shows the wrong value, and it appears only when the full suite runs. If a lone re-run passes and the page looked complete, ask what the tests share exactly one of.
- Every test reads the same inbox to grab a signup email. What is the fix?Give each test its own address that still lands in one mailbox — plus-addressing on the recipient, with a run and test tag in the suffix — then match on that address rather than on the newest message. Tests stop competing for the same email without needing eight mailboxes.
- Is disabling the rate limit in staging an acceptable answer?It is often the pragmatic one, since parallel workers concentrate traffic from few addresses and trip a limit sized for humans. Prefer raising the allowance for the test identity over switching the feature off, and remember that anything disabled in staging now needs coverage at another level.
- When should a scenario like "the banner appears when the flag is on" not be an end-to-end test at all?When the flag's value is just an input to rendering. That is a component-level test where you supply the input directly and no shared world exists. Keep the end-to-end version only for when the propagation path — how the flag actually reaches the running app — is the subject.
Eight cooks can each bring their own ingredients, but the kitchen has one oven. Duplicating the ingredients is easy; the oven temperature has to be taken in turns.
saying these in an interview costs you the question
- Calls every parallel-only failure a timing issue and adds a wait
- Tries to make a global feature flag unique per test
- Retries the interfering tests until an ordering works
- Assumes parallel workers never hit shared rate limits
- Has every test read the newest message from one shared inbox