skip to content

A thread writes several plain (non-volatile) fields, then starts a worker thread that reads them; later the parent calls `join()` and reads fields the worker wrote. Which of these reads are guaranteed correct, and by which rule?

level: middleimportance: should knowfreq 45%

answer

  1. start(): everything before the call is visible to the child
  2. join(): everything the child did is visible after return
  3. writes after start() are unguarded
  4. timed join may return without termination
  5. interrupt is the same family of edge

basics

~20 s

Both are guaranteed. Thread.start() happens-before every action in the started thread, so the worker sees everything written before the start. All actions in a thread happen-before a successful return from join() on it, so the parent sees everything the worker wrote.

solid answer

~60 s

Two of the enumerated rules cover exactly this shape: - **Start edge:** the call to `t.start()` happens-before every action performed by `t`. Everything the parent wrote *before* that call — plain fields included — is visible to the worker with no volatile or locking involved. - **Join edge:** all actions in `t` happen-before another thread returns successfully from `t.join()` (or observes `t.isAlive()` as false). So after `join()` returns, the parent sees every write the worker made. The traps are at the edges of those two rules. Fields the parent writes *after* calling `start()` are not covered — that is a race. Reads the parent performs *before* `join()` returns are not covered. And `join()` with a timeout that expires does not mean the thread terminated; you must check `isAlive()` or the thread's own completion signal, otherwise you have no edge. This is why fork-then-join code is correct without any synchronized block, and why the same code silently breaks the moment someone replaces `join()` with a sleep or a plain polling flag.

code

java · 10 lines
java
int config = 7;                 // ordered before start -> visible to worker
Worker w = new Worker();
w.start();
config = 8;                     // RACE: after start(), no edge covers this write

w.join();                       // untimed join -> edge from all of w's actions
use(w.result);                  // guaranteed to see w's writes

// w.join(1000);                // timed: may return with w still alive -> no edge
// if (!w.isAlive()) { ... }    // termination must be confirmed

go deeper

for a junior

Know that starting a thread makes prior data visible to it and that joining makes the thread's results visible to you, so simple fork/join code needs no extra synchronization.

for a middle

Name both rules explicitly, walk the transitivity chain through program order, and identify the write-after-start and read-before-join gaps.

for a senior

Add the timed-join subtlety, the interruption edge, and why replacing the join with sleeping or plain-flag polling turns correct code into a race that may only fail once the JIT hoists the read.

for a principal

Generalize to lifecycle APIs: any handoff primitive should document the edge it establishes, and code review should look for data written outside the window an edge actually covers rather than for missing keywords.

## The two rules involved The specification lists, among the actions that create happens-before edges: - A call to `start()` on a thread happens-before any action in the started thread. - All actions in a thread happen-before any other thread successfully returns from a `join()` on that thread (equivalently, before any action that detects the thread has terminated, such as seeing `isAlive()` return false). Combined with program order and transitivity, these cover the entire fork/join lifecycle without any explicit synchronization. ## Walking the guarantee Parent thread: 1. writes `config = ...` (plain field) 2. calls `worker.start()` 3. later calls `worker.join()` 4. reads `result` (plain field the worker wrote) Step 1 happens-before step 2 by program order. Step 2 happens-before every action in the worker by the start rule. By transitivity, step 1 happens-before the worker's read of `config` — so no volatile, no lock, and no `AtomicReference` is needed to hand data into a freshly started thread. On the way back: the worker's write of `result` happens-before the worker's last action by program order, all of the worker's actions happen-before the parent's successful return from `join()` by the join rule, and the parent's return from `join()` happens-before step 4 by program order. Transitivity again gives the parent a guaranteed view of `result`. ## The three ways this is gotten wrong **Writing after `start()`.** The start edge orders only what preceded the call. A parent that starts the worker and *then* sets `config` has no edge for that write; the worker may read the old value indefinitely. This is a common bug when initialization is split across a builder and a `start()` call. **Reading before `join()` returns.** Polling a plain `done` flag, or sleeping for "long enough", produces no edge at all. Elapsed time is not a rule. The read may return a stale value forever, and in compiled code a hoisted read of a non-volatile flag can turn a polling loop into an infinite loop. **Timed join.** `join(millis)` returns whether or not the thread finished. If it returned because the timeout expired, the thread is still running and no edge exists; the calling thread must check `isAlive()` (or loop) to know it actually terminated. Treating a timed join like an untimed one is a subtle race. ## Related lifecycle edges The interruption rule sits in the same family: if the parent sets a plain field and then calls `worker.interrupt()`, the field write is visible to the worker at the point it detects the interrupt, whether by catching `InterruptedException` or by observing `isInterrupted()`. This is why a cancellation reason written just before an interrupt can be read safely by the cancelled thread. Higher-level constructs give the same shape with better ergonomics. Submitting a task to an executor orders everything before the submission with the task's execution, and `Future.get()` returning orders the task's actions before the caller's subsequent reads. Those are not extra rules — they are the same edges established internally by the library's own locks and volatile fields. ## Why it matters in interviews The question separates candidates who memorized "use volatile or synchronized" from those who can reason with the rule set. Adding `volatile` to every field passed into a worker thread is harmless but reveals the misunderstanding; recognizing that the start and join edges already cover it, and knowing precisely where their coverage ends, is the real signal.

  • Does replacing `join()` with a loop polling a plain boolean `done` field preserve the guarantee?
    No, on two counts. A plain field read gives no happens-before edge, so the parent may never observe `done` becoming true, and even if it does, the worker's other writes are unordered relative to that read. Making `done` volatile fixes both, because the volatile write-read pair creates the edge and transitivity carries the earlier writes with it.
  • Is `Thread.sleep()` long enough ever a substitute for `join()`?
    Never. The memory model has no rule mentioning elapsed time; sleeping creates no edge. It may appear to work because in practice caches propagate quickly, but a compiler that hoists a non-volatile read out of a loop can defeat it permanently, and the program is racy by definition regardless of observed behaviour.

saying these in an interview costs you the question

  • Adding `volatile` to fields written before `start()` on the belief they are otherwise unsafe.
  • Believing fields written after `start()` are still covered by the start edge.
  • Treating `join(timeout)` as equivalent to `join()` without checking `isAlive()`.
  • Substituting `Thread.sleep()` or a plain polling flag for `join()`.
  • Claiming the parent must synchronize on the thread object to read the worker's results.

context