skip to content

Simply running a multi-threaded test in a loop usually finds nothing. How would you design a stress harness that actually surfaces an interleaving bug?

level: middleimportance: must knowfreq 48%

answer

  1. staggered start = never overlapped
  2. barrier release each iteration
  3. shrink body so window dominates
  4. assert an invariant, not a crash
  5. log the seed or the failure is worthless

basics

~20 s

Shrink the code under test so the bad window is a large fraction of it, start all threads from a barrier so they collide instead of running sequentially, perturb timing with randomized delays and varied thread counts, check a real invariant rather than a crash, and record the seed and observed state so a failure is reproducible.

solid answer

~60 s

A naive loop fails because each iteration re-runs the same alignment: thread A finishes its short body before thread B is even scheduled, so the dangerous window never overlaps. Fix that deliberately. - **Narrow the target.** Test the smallest operation that can break. The chance of hitting a 50-instruction window inside a 50,000-instruction test is negligible; alone it is substantial. - **Synchronize the start.** Release all threads from a barrier per iteration so they enter the critical region together instead of staggered. - **Perturb the schedule.** Randomized short delays or yields at varied points, run with different thread counts, and on different core counts — a bug that needs true parallelism will never appear on one core. - **Assert an invariant, not a crash.** "Sum of balances is constant", "every enqueued element is dequeued exactly once". Many interleaving bugs produce wrong state, not exceptions. - **Make failures usable.** Fix the seed per iteration, log it, and dump the offending state — an unreproducible one-in-a-million failure that is not captured is nearly worthless. Then run it long, and treat any failure as real, never as flakiness.

code

text · 12 lines
text
threads = spawn N workers once      // not per iteration

for iter in 1..M:
    seed = iter
    reset(sharedObject)
    barrier.await()                 // all threads released together
    each worker:
        maybe_delay(random(seed))   // perturb the alignment
        op(sharedObject)
    barrier.await()
    if not invariantHolds(sharedObject):
        report(seed, iter, dump(sharedObject)); fail

go deeper

for a junior

Say that threads must actually run at the same time, that the tested operation should be small, and that the test should check a property rather than just not crashing.

for a middle

Give the concrete mechanics: reuse threads with a barrier, randomize delays and thread counts, assert an invariant, log the seed.

for a senior

Add configuration variety (core counts, instrumented vs. fast runs), invariant design such as conservation and exactly-once, and the rule that a rare failure is a bug and blocks.

for a principal

Place it in the pipeline — short scoped stress per commit, long seeded soaks nightly — and be explicit that sampling raises probability without proving absence, which is when you escalate to controlled or exhaustive exploration.

## Why the naive loop fails Running `for i in 1..100000 { runConcurrentTest() }` feels like coverage and usually is not. Three effects conspire: 1. **Staggered start.** Creating threads is expensive relative to a small test body, so the first thread often completes before the last one starts. The threads never actually overlap, and the interleaving space explored is one point. 2. **Repetitive alignment.** When they do overlap, the scheduler and the code produce the same relative phase every iteration. A hundred thousand runs of the same alignment is one sample, repeated. 3. **Window dilution.** The bug needs thread B to arrive within a window of a few instructions. If the test body is long, the probability of arriving in that window per iteration is tiny. Every design rule below attacks one of these. ## Rule 1: make the window a large share of the test Strip the test to the smallest operation that can exhibit the bug. Not "process an order end to end" but "two threads call `reserve()` on this counter". If the dangerous region is ten instructions out of thirty, collisions happen constantly; if it is ten out of a hundred thousand, you are waiting for lightning. Shrinking the body is by far the highest-leverage change and is usually where a failing harness turns into a passing one. ## Rule 2: force the collision Allocate the threads once, outside the loop, and use a **barrier** so every iteration begins with all participants released simultaneously. This removes the staggered-start effect entirely and is the single most common omission in hand-written stress tests. A spin phase just before the barrier release also helps ensure each thread is genuinely running on its core rather than being woken from a park. ## Rule 3: perturb deliberately Once threads overlap, you still need to vary *where* the overlap lands. - Insert short randomized delays or yields at a few candidate points, chosen from a per-iteration seed. - Vary the number of threads: 2 finds pairwise races; 8 finds them faster; a count exceeding the core count adds preemption at arbitrary points. - Vary the machine shape: a bug needing genuine parallelism cannot appear when everything is time-sliced onto one core, and some bugs only appear under preemption. Run both. - Vary the operation mix and the data: contended keys versus spread keys change which paths collide. Note the tension with detection tooling: heavy instrumentation slows every operation and can serialize threads, shrinking the very variety you are engineering. Run both an instrumented configuration and a fast uninstrumented one. ## Rule 4: check an invariant Many interleaving bugs corrupt state quietly. A harness that only detects crashes will miss a lost update entirely. Define a property that must hold and check it: - **Conservation**: total across accounts is unchanged after all transfers. - **Uniqueness**: each produced item is consumed exactly once, no duplicates, no losses. - **Monotonicity**: a sequence number never goes backwards. - **Linearizability-flavoured checks**: record each operation's start and end and verify some sequential ordering explains the observed results. Check at the end of each iteration, and if the check is cheap, mid-flight as well. ## Rule 5: engineer the failure report A rare failure that cannot be reproduced or understood costs more than it gives. Before the run, derive everything random from a **per-iteration seed** and log that seed. On failure, dump the seed, thread count, operation log and the exact invariant violation. The goal is that the next person can re-run just that iteration. ## Rule 6: treat failure as failure The cultural rule matters as much as the technical ones. A stress harness that fails once in three thousand CI runs will be labelled flaky and retried away. Establish before the first failure that a failing invariant is a real bug and blocks; route the failure with its seed to an owner. Otherwise the harness quietly becomes a noise generator you have taught the team to ignore. ## Where it goes in the pipeline Stress runs are wall-clock hungry and unbounded in value, so they usually do not belong on the per-commit path. Run a short, tightly-scoped stress suite per commit for the primitives that matter, and a long soak nightly with many seeds and configurations. Keep the harness itself in the repository beside the primitive it exercises, so it is re-run when that code changes. ## The honest limitation Even a well-built harness samples. It raises the probability of hitting a bad interleaving by orders of magnitude but never proves that none exists. When you need that stronger claim, the small, stripped-down test you just built is exactly the input a controlled scheduler or an exhaustive explorer needs.

  • Why is starting the threads from a barrier each iteration so important?
    Thread creation and wake-up costs are large compared with a small test body, so without a barrier the first thread frequently completes before the last one starts and the threads never overlap at all. A barrier releases them at the same instant, which is the precondition for any interleaving to occur. It also lets you reuse threads across iterations instead of paying creation cost per iteration.
  • Your stress test fails roughly once every 2,000 CI runs. What do you do?
    Treat it as a real bug and preserve the evidence: the seed, thread count and invariant violation should already be logged so the iteration can be re-run in isolation. Then narrow it — shrink the operation sequence and increase perturbation until the reproduction rate is high enough to debug. What you must not do is add a retry or mark it flaky, since the failure is reporting a genuine interleaving the production system can also hit.

saying these in an interview costs you the question

  • Creating fresh threads inside the loop, so they never overlap
  • Testing a large end-to-end flow instead of the smallest breakable operation
  • Detecting only crashes and exceptions rather than checking an invariant
  • Adding sleeps to "stabilize" a failing stress test
  • Marking a rare stress failure as flaky and retrying it
  • Not recording the seed, making failures unreproducible

context