N threads must all reach a shared rendezvous point before any continues, and one participant dies, times out, or is cancelled while the others are already waiting. What should happen to the remaining waiters, and why do rendezvous primitives expose an explicit "broken" state instead of letting them keep waiting?
answer
- Nth arrival can never come — unsatisfiable forever
- broken flag is per generation, wakes everyone
- one timeout or cancel breaks it for all
- action throws → break, state may be half-updated
- reset the barrier, restart the phase, never resume
basics
~20 sThe rendezvous can never complete, so leaving the others blocked is a permanent, silent hang. The barrier is marked broken and all current and future waiters fail immediately with a distinct error, so failure propagates to every participant instead of one loud error plus N−1 stuck threads.
solid answer
~50 sA rendezvous of N parties only completes on the Nth arrival, so losing one participant makes completion impossible forever. The alternatives are: block the other N−1 indefinitely (a silent hang with no stack trace anywhere useful), or **break** the barrier. Breaking means: set a broken flag on the current generation, wake every waiter, and make them — and anyone who arrives later at that generation — throw a distinct "broken barrier" error rather than a generic timeout. Cancellation or a timeout on *any* waiter breaks it for all, and if the optional barrier action throws, that too must break the barrier, because the phase transition it was supposed to perform did not happen. The design principle is **all-or-nothing failure propagation**: the participants are mutually dependent, so failure must be too. Recovery means resetting the barrier to a fresh generation and restarting the whole phase — you cannot resume mid-phase, because you do not know how much of the phase each survivor completed.
code
text · 18 linesstate: parties=N, arrived, generation, broken
await(timeout):
lock()
g = generation
if broken: unlock(); throw BrokenBarrier
arrived += 1
if arrived == parties:
try: action() // phase transition, single-threaded
catch e: break_all(); throw e
nextGeneration(); wakeAll(); unlock(); return
while generation == g and not broken:
if not cond.wait(timeout): // expiry
break_all(); unlock(); throw Timeout
unlock()
if broken: throw BrokenBarrier
break_all(): broken = true; wakeAll() // every waiter fails, not just onego deeper
Know that a rendezvous needs all N parties, so losing one blocks the rest forever, and that a good implementation fails everyone loudly instead of hanging.
Name the four triggers (cancel, timeout, throwing action, explicit reset), explain that breakage is per generation and affects late arrivers too, and that the error type is distinct from a timeout.
Argue the all-or-nothing propagation principle, walk the thread-dump diagnosis (N−1 waiting, find the missing party, check party-size arithmetic), and explain why reset requires restarting the phase.
Treat the barrier as a consistency boundary: pair it with phase-boundary checkpoints, bounded waits derived from expected phase time, and an escalation policy that fails the job rather than degrading silently — and question whether a global rendezvous is needed at all.
## Why one lost participant is fatal A barrier of party size N is a *mutual* dependency: the release condition is "N arrivals in this generation". Every waiter's progress is conditional on every other party's arrival. Remove one party — it crashed, it was cancelled, it timed out, it exited its loop early, it deadlocked elsewhere — and the release condition becomes unsatisfiable. The remaining N−1 threads are not slow; they are permanently blocked. This is qualitatively worse than a single task failing, because the failure is *invisible*. The dead thread may have logged an error, but the N−1 survivors just sit in a wait state, holding whatever resources they hold. In production this presents as "the job stopped making progress", with no exception at the point of the hang. ## The broken state The fix is to make the barrier itself fail loudly. A barrier carries per-generation state: an arrival count, a generation identity, and a **broken** flag. Any of the following breaks the current generation: 1. A waiting thread is interrupted or cancelled. 2. A waiting thread's bounded wait expires. 3. The barrier action (the code run once when the barrier trips) throws. 4. Someone explicitly resets the barrier while parties are waiting. On breakage the barrier wakes every waiter, and each of them — plus every party that arrives at that same generation afterwards — completes with a distinct *broken barrier* error rather than returning normally. The distinct error type matters: a generic timeout says "I waited too long", which invites a retry; "the barrier is broken" says "this phase is dead, do not retry the wait, unwind". ## Why timeout-on-one breaks it for all A subtle but important rule: if one participant does a bounded wait and it expires, the barrier breaks for everyone, not just for the impatient thread. If it did not, that thread would leave the rendezvous while the barrier still expected N arrivals, and the arithmetic would be silently wrong for the rest of the computation — the next generation would trip one arrival early or never trip. Consistency of the party count is what forces the all-or-nothing rule. ## Why a throwing barrier action must break it The barrier action is the phase transition: swap buffers, check convergence, publish the round's aggregate. If it throws halfway, shared state is in an unknown, possibly half-updated condition. Releasing the N waiters to run the next phase over that state would turn a clean failure into corrupted results. So the action's failure is propagated by breaking the barrier and delivering the action's exception to the waiters. ## Recovery: reset, restart the phase, do not resume A broken barrier can be *reset* — discard the current generation, clear the broken flag, start a fresh generation with zero arrivals. But reset only fixes the primitive, not the computation. The survivors are at unknown, differing points in the phase: some finished their slice, some were mid-write, some never started. There is no safe "continue from where we were". Recovery therefore means: - fail the entire phase (or the entire job) outward, - restore or discard the phase's mutable state, - if you retry, retry the whole phase from a known-good snapshot, with a fresh full set of participants. This is why long-running bulk-synchronous systems checkpoint at phase boundaries: the barrier gives you a consistent point at which state is known-good, and any breakage rolls back to the last such point. ## Diagnosing a stuck rendezvous When a barrier-based system stops progressing, the diagnostic sequence is: 1. Dump all threads. N−1 sitting in the barrier's wait is the signature. 2. Find the missing party: it either died (look for an unhandled exception on a thread whose cleanup did not run), is blocked somewhere else (a lock cycle, a slow I/O call), or exited its loop early on a condition the others do not share. 3. Check party-size arithmetic: a barrier constructed for N but joined by N−1 threads never trips even in the happy path. This is a common off-by-one when the coordinating thread does or does not count itself as a party. Preventive design: give every barrier wait a bound derived from the expected phase duration, make it break on expiry, and log which parties had arrived at the moment of breakage — the set of *missing* parties is the actual diagnosis. ## Contrast with a latch None of this applies to a countdown latch, because a latch has no mutual dependency: a worker that dies without counting down leaves the coordinator waiting, but no worker is waiting on any other worker. That is why latches have no broken state, and why the standard mitigation there is simply a bounded await plus counting down in a cleanup path.
- Why should one participant's timeout break the barrier for everyone rather than just letting that thread give up?Because the party count must stay consistent. If the impatient thread simply left, the barrier would still expect N arrivals while only N−1 participants remain, so the next generation would either never trip or trip with the wrong set of parties. Breaking for all keeps the invariant intact and converts an ambiguous partial state into one clear, shared failure.
- After a barrier is reset, can the surviving workers resume the phase they were in?No. Reset restores the primitive, not the computation: the survivors are at unknown and differing points in the phase, and any shared phase state may be half-updated. You must fail the phase outward and restart it from a known-good snapshot with a full set of participants — which is why bulk-synchronous systems checkpoint at barrier boundaries.
A climbing team roped together at a checkpoint: if one member turns back, the rest cannot simply stand there forever. The rope is cut deliberately and everyone is told the attempt is off — a loud, shared abort beats silent, indefinite waiting.
saying these in an interview costs you the question
- Saying the surviving threads should just keep waiting because the failed party might come back.
- Letting one participant time out and walk away while the others continue to expect N arrivals.
- Reporting breakage as a generic timeout, so callers retry the wait instead of unwinding the phase.
- Assuming a barrier action that throws can be ignored and the waiters released anyway.
- Claiming reset lets workers resume mid-phase, rather than restarting the phase from a consistent point.