skip to content

questions

4

A worker thread spins in a loop reading an ordinary boolean 'stop' variable while another thread sets it to true and then exits. Sometimes the worker never leaves the loop. Explain how that is possible even though the write definitely executed.

level: juniorimportance: must knowfreq 58%

answer

  1. no edge = no obligation to ever see it
  2. loop-invariant hoist: load once into a register
  3. store buffer delay on the writer side
  4. sleep/print 'fix' = perturbation, not a fix
  5. flag AND the data it guards both need the edge

basics

~20 s

Nothing orders the write against the read, so the reader has no obligation to observe it. The compiler may load the variable once into a register and loop on that copy, and the write may sit in the writer's store buffer. Publish the flag through a synchronization edge.

solid answer

~50 s

There is no happens-before edge between the two threads, so the memory model permits the reader never to see the new value. Two concrete mechanisms produce it. First, the compiler: a plain variable that the loop body never modifies is loop-invariant and can be hoisted into a register before the loop, turning `while (!stop)` into `if (!stop) loop forever`. Second, the hardware: the writer's store may sit in a store buffer, and the reader may have older data in flight. Neither is a bug - both are legal because the program never asked for ordering. The fix is to create the edge: mark the flag with the language's synchronized-visibility qualifier so the write is a release and the read an acquire, or read and write it under the same mutex, or use a proper cancellation primitive. Adding sleeps, logging, or a delay only perturbs the compiler's decisions - it hides the bug rather than fixing it.

code

text · 8 lines
text
source:              legal transformation:
while (!stop) {        r = stop
    work()             if (!r) {
}                          while (true) work()
                       }

No write to `stop` inside the loop and no synchronization,
so the load is loop-invariant and may be hoisted out.

go deeper

for a junior

Say clearly that shared mutable state read by one thread and written by another needs synchronization, and that without it the reader may never observe the change.

for a middle

Name both mechanisms - register hoisting by the compiler and store buffering by the CPU - and describe the release/acquire pair or mutex that fixes it.

for a senior

Point out that timing-based 'fixes' only change probability, that the data guarded by the flag needs the same edge, and prefer a blocking cancellation primitive over a spun flag.

for a principal

Treat it as an API design issue: cancellation and completion signalling should come from a reviewed primitive so application code never has to reason about the visibility of a raw flag.

## What the code assumes The loop `while (!stop) { work() }` with a plain shared boolean assumes that a write performed by another thread will eventually become visible to this one. No mainstream memory model makes that promise for unsynchronized variables. The promise exists only where there is a happens-before edge, and here there is none: the writer performs a plain store, the reader performs plain loads, and nothing pairs them. ## Mechanism one: the compiler Optimizers work on a single-thread abstraction. Within one thread nothing in the loop body modifies `stop`, so the value is loop-invariant and can be loaded once into a register before the loop. The generated code becomes, in effect: ``` r = stop if (!r) { while (true) work() } ``` That transformation is correct for a sequential program, and the language rules permit it because a concurrent modification without synchronization is a data race - a case the compiler is entitled to ignore. This is the mechanism that most often makes the loop hang forever rather than merely late. ## Mechanism two: the hardware Even if the load is re-executed each iteration, the writer's store may sit in its store buffer for a while before becoming globally observable, and the reading core may have older data in flight. On strongly ordered CPUs this window is short - microseconds - which is exactly why the bug so often reproduces only under a different compiler, optimization level, architecture, or load. ## Why timing changes hide it People frequently report that adding a print statement, a short sleep, or an unrelated lock 'fixed' it. Those calls are opaque to the optimizer or contain synchronization of their own, which forces the variable to be reloaded. The program is still racy; only the probability changed. Treating that as a fix is the classic wrong turn. ## The correct fixes All of them work by creating an edge: - **Mark the flag** with the language's visibility qualifier - a release-store paired with an acquire-load. The write is then guaranteed to become visible in finite time and cannot be cached in a register across iterations. - **Use the same mutex** for reading and writing the flag. Unlock happens-before the next lock, so the reader is obliged to see the write. - **Use a cancellation primitive** - an interruption request, a cancellation token, closing a channel the worker waits on. These are already synchronized and additionally let a blocked worker wake up, which a spun flag does not. ## A second, quieter bug The same missing edge affects the data the flag is meant to guard. Code often waits for the flag and then reads a result the other thread wrote before setting it. Without an edge the reader can see `stop == true` and still see the old result, because the two plain writes may become visible in either order. Once the flag is a proper release/acquire pair, all writes preceding the release are visible after the acquire, and both problems disappear together. ## Cost note A spin loop that constantly reloads a shared variable also burns a core and generates cache-line traffic. If the wait may be long, block on a condition variable or channel instead; the visibility fix and the efficiency fix usually point at the same construct.

  • Someone reports that adding a log line inside the loop made the hang disappear. What do you tell them?
    That the race is still there. A logging call is an opaque function call and usually takes a lock internally, so the compiler can no longer keep the flag in a register and is forced to reload it. The observable behaviour changed but no ordering guarantee was added, so the bug will return with a different compiler, a different optimization level, or a lucky inlining decision. The fix must be an explicit synchronization edge.
  • Is a spin loop on a properly synchronized flag a good way to wait for work?
    It is correct but usually wasteful: it occupies a core, denies the scheduler that capacity for real work, and keeps a shared cache line hot. Spinning is justified only for very short, predictable waits, ideally with a backoff hint. For anything longer, block on a condition variable, semaphore or channel, which parks the thread and provides the ordering edge for free.

saying these in an interview costs you the question

  • Saying 'the write will show up eventually, it is just slow' - without an edge there is no eventual guarantee
  • Blaming CPU caches only, missing that compiler hoisting of the load is the more common cause
  • Claiming a sleep, a print, or a yield in the loop fixes it
  • Synchronizing the flag but leaving the data it guards unsynchronized, then being surprised by a half-visible result
  • Thinking the flag needs a lock for atomicity - a single boolean store is atomic; the missing property is ordering

context

open as a page

In a shared-memory concurrent program, what does it mean to say that one action happens-before another, and what does that relationship guarantee about what a thread can see?

level: middleimportance: must knowfreq 55%

basics

~20 s

Happens-before orders actions: program order inside a thread, plus synchronization edges across threads (a release paired with a matching acquire). If A happens-before B, every write made before A is visible to B. It is transitive, not chronological.

open as a page

Several languages offer a qualifier that marks a shared variable so its reads and writes participate in the memory model - Java's and C#'s volatile, or a C++ atomic used with release/acquire ordering. What does such a marking guarantee, what does it not guarantee, and how does C's volatile differ?

level: middleimportance: must knowfreq 48%

basics

~20 s

It guarantees the variable is really read and written each time, that a write becomes visible to a later read in finite time, and that the write acts as a release and the read as an acquire, so earlier writes are visible too. It does not make read-modify-write atomic. C's volatile gives none of this.

open as a page

Handing another thread a reference to an object you just finished building is a classic hazard in shared-memory concurrency. Why is that hazard a property of the shared-memory model rather than of the object, and how do message-passing designs — actor mailboxes, CSP-style channels, ownership transfer — remove it by construction?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Shared memory lets a reference reach another thread with no ordering edge, so its fields can look unwritten. Message passing makes the send/receive pair itself the edge and the only route to the value; copy or move semantics also remove aliasing.

open as a page