A case fails only when the suite runs with several test workers, yet passes alone — how do you find the shared harness state?
answer
- Worker count is a knob, not luck
- Shrink to the smallest failing pair
- Run the case beside a copy of itself
- The red case is the reader
basics
~20 sTreat worker count as a variable you control: reproduce at two workers, shrink to the smallest failing pair, then run the case beside a copy of itself. Then hunt the writer, not the reader that reported red.
solid answer
~50 sThe useful property is that this is not luck: worker count is a knob, so the failure is reproducible on demand. Start by shrinking. Drop to two workers, bisect down to the smallest failing set of cases, then run the failing case beside a copy of itself — if it still fails, the conflicting write happens in every case; if it passes, one specific partner case is involved. Then invert the search. The case that reports red is the reader, and the defect belongs to whatever wrote the value it read, which usually passes. So log every write to a harness-held value with the value name, the worker identity and the case name, and line that log up against the failing assertion. Confirm by repeating the reduced run several times, not by one green result.
code
pseudocode · 16 lines# 1. make every write to harness-held data name its author
set_case_value(name, value):
log("write", name = name, value = value,
worker = current_worker_id(), case = current_case_name())
store[name] = value
# 2. or make the collision itself the failure
set_case_value(name, value):
if store.has(name) and store.owner(name) != current_case_name():
fail("value " + name + " already owned by " + store.owner(name))
store.put(name, value, owner = current_case_name())
# 3. reduce along the axis you control
for workers in [1, 2]:
for candidate in smallest_failing_sets(suite):
run(candidate, workers = workers)go deeper
Know the first fact: a case that passes alone but fails alongside others is telling you something is shared. Report it that way, with the worker count you used, rather than re-running until it goes green.
Be ready to describe the reduction — fewer workers, fewer cases, then the case run beside a copy of itself — and say what each of those outcomes rules in or out.
Demonstrate the inversion: the red case is the reader, so instrument the writes with the value name, worker identity and case name. Say how you confirm a fix rather than trusting a single green run.
Own the default. Decide what the harness records about writes to shared values on every run, so the next failure of this shape costs an hour of reading rather than a week of guessing, and say who pays for that.
A case that passes alone and fails when the suite runs on several test workers looks like bad luck. It is not: the number of workers is a knob you control, which makes this one of the most tractable failures in a suite. ## Why this is reproducible rather than random An intermittent failure with no controllable input is expensive because you cannot make it happen on demand. Here you can. If the failure appears at four workers and never at one, then the parallel schedule is a necessary ingredient, and any run you construct with enough concurrency has a real chance of reproducing it. That changes the job from waiting to searching. ## Reduce along the axis you control Shrink the run in this order, and treat each result as an answer rather than a step: 1. **Fewer workers.** Drop to two. Reproducing at two workers means only two participants are needed, and every later run is cheap. 2. **Fewer cases.** Bisect the suite by halves at fixed worker count until the smallest failing set remains. Two cases is the usual floor. 3. **The case beside a copy of itself.** Run the failing case twice, concurrently, and nothing else. If it still fails, the conflicting write is something *every* case does, so any case would have served as the partner. If it passes, one specific other case is involved, and you have halved the search again. 4. **Repeat the reduced run.** A shrunken run that fails three times out of five is a reproduction; one that fails once in twenty has not been reduced yet. | What you observe | What it usually means | | --- | --- | | Fails at two workers, passes at one | a value the harness shares across concurrently running cases | | Fails beside a copy of itself | the conflicting write happens in every case, not one special case | | A different case reports red each run | the reader and the writer are racing; the loser moves | | Same case always red, others green | that case is the slowest reader, not the culprit | | Also fails serially in that case order | an ordering dependency on data, which is a different hunt | ## Invert the search: look for the writer This is the step people skip. The case that reports red **read** a value that was wrong; the defect belongs to whatever **wrote** it, and the writer usually passes. Debugging the red case therefore inspects the victim. Instrument the writes instead. Give every write to a harness-held value a log line carrying the value name, the worker identity and the case name, then compare that log against the failing assertion's timestamp. A write from worker B landing between worker A's setup and worker A's assertion is the whole diagnosis, and it takes one run to see once the logging exists. Two cheaper variants of the same idea, when adding logging is awkward: - Make an unset or replaced read **raise** with the identity of the previous writer, so the failure reports the collision instead of a mismatched value. - Tag every value with the case that created it and assert on read that the tag matches the running case. That converts a silent wrong answer into a loud, precise one. ## Confirming, not hoping A single green run proves almost nothing here, because the failure was always probabilistic given a fixed schedule. Confirm a fix three ways: re-run the reduced set enough times to beat its original failure rate, run the full suite at the worker count that failed, and keep the reduced run around as a case in its own right if the shape is likely to recur. ## Where this method stops It finds state the harness itself holds. It will not find two cases fighting over the same record in the deployed target, or over one shared account, or over a setting one of them toggles — those reproduce serially in the right order, and the fifth row of the table above is how you tell the difference early rather than after two days of reading harness code. If reducing the pair and running it serially in that order reproduces the failure, stop looking at concurrency: the schedule was revealing an ordering dependency, not creating one.
- The failing case is different on almost every run. Does that change your approach?No, it supports the diagnosis. A moving failure means the loser of the race moves, which is what a shared value produces. Keep reducing: fewer workers, then fewer cases. The pair that reproduces at two workers is the pair to instrument, whichever of the two happens to report red on the run you are looking at.
- The reduced pair passes at two workers but the full suite still fails. What now?The interaction needs a third participant, or timing that two cases alone do not create. Keep the worker count that failed and bisect by halves of the suite instead of by pairs: run half at full parallelism, then the failing half again, until the smallest failing set remains. If nothing reduces, instrument writes across the whole run instead and let the log do the search.
- How do you know the fix worked, given the failure was never certain to appear?Beat the original failure rate with repetition rather than trusting one green run. Re-run the reduced set enough times that its old rate would almost certainly have shown, run the full suite at the worker count that failed, and keep the reduced run as a case of its own if the shape is likely to recur elsewhere in the harness.
Two people editing the same line of one shared note: the person who notices the change complains loudly, and the person who made it has already moved on.
saying these in an interview costs you the question
- Re-runs until green and moves on
- Raises a wait timeout so the failure stops appearing
- Debugs only the case that reported red
- Calls it noise without ever reducing the worker count
- Concludes the suite simply cannot be run in parallel