skip to content

questions

4

A harness keeps the current case's test user in one process-wide variable; why does that break once two test workers run at once?

level: middleimportance: must knowfreq 60%

answer

  1. One slot cannot describe two cases
  2. Serial order was the hidden guarantee
  3. Last write wins; the reader gets it wrong
  4. Scope it to the case, not the process

basics

~20 s

Both workers write the same variable, so a case reads whatever the other stored last. Serial execution was the only thing making that safe. The result is wrong-data assertions and cross-talk between cases, not a crash.

solid answer

~40 s

A single process-wide holder has exactly one slot, so it can only ever describe one case. While the suite runs one case at a time that is fine: the value is written in setup, read by helpers, and replaced by the next case. Add a second test worker and the two cases interleave — one worker writes its user, the other overwrites it, and the first worker's helper reads the wrong one. The failure is quiet. Nothing throws; assertions simply compare the wrong record, so the red case moves between runs and passes when re-run alone. The fix is scope, not locking: give the value a scope no wider than the thing it describes — one holder per test worker, or the value carried on the case's own context object.

code

pseudocode · 19 lines
pseudocode
# broken: one slot for the whole process
CURRENT_USER = null

setup(case):
    CURRENT_USER = create_user(case.role)

place_order(item):
    return submit_order(user = CURRENT_USER, item = item)

# worker A: setup -> CURRENT_USER = alice-01
# worker B: setup -> CURRENT_USER = bob-02      (overwrites)
# worker A: place_order(...) -> submits as bob-02

# scoped: one slot per worker
setup(case):
    store_for_this_worker("user", create_user(case.role))

place_order(item):
    return submit_order(user = read_for_this_worker("user"), item = item)

go deeper

for a junior

Be ready to say plainly what a process-wide holder is: one named slot shared by everything in the run. Know that a second case writing it replaces what the first case stored, and that nothing warns you.

for a middle

Explain the interleaving out loud — write, overwrite, read — and name the alternatives: one slot per test worker, or the value carried on the case's own context object. Say why locking is not the fix.

for a senior

Show the diagnosis as well as the mechanism: the red case is the reader, so hunt the write. Expect to be asked what you would change first in a suite that already has a dozen of these holders.

for a principal

Own the rule rather than the instance. Decide what the harness is allowed to hold at all, make the scope of each value visible in the code, and say how you stop the next convenience from reintroducing the same slot.

A test harness accumulates conveniences. Early on, a setup step needs to tell a helper which user the current case is acting as, and the shortest path is a value that lives for the whole process: write it in setup, read it anywhere. It works. It keeps working — for exactly as long as one case runs at a time. ## What serial execution was quietly providing A process-wide holder has one slot, and one slot can describe one case. Running cases one at a time guarantees that only one case sits between its setup and its teardown at any moment, so the slot is always describing the case that is asking. That guarantee is a property of the **schedule**, not of the code, and it is written down nowhere. Nobody chose to depend on it; the dependency arrived for free. Enabling a second test worker changes the schedule. Nothing about the holder changed — the assumption underneath it did. ## The interleaving, step by step Two cases, two test workers, one slot: 1. Worker A's setup creates and stores the user `alice-01`. 2. Worker B's setup creates and stores `bob-02`, overwriting the slot. 3. Worker A's helper reads the slot and receives `bob-02`. 4. Worker A places an order on the wrong account and asserts against `alice-01`'s history. Notice what did not happen. Nothing threw. Nothing timed out. No lock was contended and no operation was slow. The write succeeded and the read succeeded; the suite is simply describing the wrong case. That is why the symptom is an assertion mismatch — an unexpected total, an empty list, a record belonging to someone else — rather than an error naming the shared value. ## Why the attribution is wrong by default The case that reports red is the one that **read** the mixed-up value. The mistake was made by whichever case **wrote** it, and that case very often passes, because it read its own value before anyone replaced it. Three consequences follow, and each of them costs a day: - The failing case name moves between runs, because which worker loses the race depends on timing. - Re-running the failing case alone passes, which reads as evidence that the case is fine. - The bug is assigned to whoever last touched the failing case, and there is nothing wrong with it. ## Scope is the whole decision The repair is not to make the shared value safe to touch from two threads. Wrapping reads and writes in a lock removes a data race but leaves the mix-up untouched: two cases still take turns writing one slot, so a case still reads a value another case put there. The question is not *is access to this value serialized* but *how many cases can this value describe*. | Where the value lives | Correct while | What it costs | | --- | --- | --- | | One slot for the whole process | one case runs at a time | breaks silently at the second worker | | One slot per test worker | each worker runs one case at a time | helpers still read rather than receive | | Carried on the case's own context object | always | one more value threaded along the path | | Passed as a parameter to each helper | always | every signature on the path grows | The middle two are the usual landing places, and which one you pick depends on how much of the harness you are willing to change at once. ## Where these holders hide They are rarely called *the shared holder*. Look for: - the acting user, or the credentials the run authenticated with - the address of the deployed target, when different cases point at different ones - a correlation tag attached to outbound requests so the run can find its own traffic later - accumulated results a reporter reads at the end of the run - a lazily built client kept around because constructing it is slow The last is the most seductive, because caching a costly object is a genuinely good instinct. It is safe only when the cached object is either immutable or built once per test worker; a cached object that carries the current case's identity is the same bug wearing a performance justification. ## Making the change without a rewrite Change scope one value at a time, and make the tooling do the finding. Rename the holder as you scope it, so every existing read breaks loudly at build time rather than silently picking up a default. Where renaming will not help, make an unset read raise instead of returning empty — a loud failure in a run you are already debugging is far cheaper than a quiet one months later. Finally, prove the fix instead of assuming it. A suite that passes once at four test workers has proved very little. Re-run the reduced pair of cases several times, and keep a single-worker run in the pipeline for a while, so a regression appears as a difference between two runs rather than as a mystery.

  • Why is the case that reports red usually not the case with the defect?
    The red case is the reader. It received a value that some other case wrote, and that other case very often passes, because it read its own value before the replacement happened. So the failure names the victim. The useful search is for the write that landed between the red case's setup and its assertion, not for anything inside the red case itself.
  • If a value is written once before any case runs and never changed, is a process-wide holder still a problem?
    No. An immutable value shared by every worker is safe, because there is no write to lose a race with. The hazard is specifically a slot rewritten per case. Watch for the drift, though: a value that starts run-wide and later gains a per-case override reintroduces the bug, and nothing about the holder's shape shows that it happened.
  • Does putting a lock around the shared holder fix it?
    No. A lock makes each read and write atomic, which removes a torn or half-written value, but the slot still holds one value for the whole process. Two cases take turns writing it, so a case still reads what another case put there. The problem is how many cases the value can describe, and a lock does not change that number.

One coat hook shared by a whole building works perfectly while there is one guest. The second coat does not fail loudly; it simply replaces the first.

saying these in an interview costs you the question

  • Says a shared holder is fine because reads and writes are fast
  • Adds a lock around the holder and calls the problem solved
  • Blames the machine or the runner for intermittent failures
  • Assumes the defect lives in the case that reported red
  • Treats a passing solo re-run as proof the case is correct
open as a page

What do you gain by passing a case's test data into a helper as parameters instead of letting the helper read a shared holder?

level: juniorimportance: should knowfreq 40%

basics

~20 s

The helper's inputs become visible in its signature: a reader knows what it needs, a caller can supply different data, and nothing outside the call can change it midway. The cost is longer signatures on every step of the chain.

open as a page

A case fails only when the suite runs with several test workers, yet passes alone — how do you find the shared harness state?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Treat worker count as a variable you control: reproduce at two workers, shrink to the smallest failing pair, then run the case beside a copy of itself. Then hunt the writer, not the reader that reported red.

open as a page

Giving each parallel test worker its own current-case context fixes cross-talk; what does a helper that reads that context implicitly still cost you?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Scoping removes the mix-up, not the hidden dependency. A helper that reaches for an ambient current-case value has an input its signature does not declare, cannot be reasoned about locally, and misbehaves wherever no case is in scope.

open as a page