You have one coordinator thread that must not proceed until N independent worker tasks have each finished. Describe how a one-shot countdown latch solves this, and what happens to a thread that waits on a latch whose count has already reached zero.
answer
- counter fixed at construction, only counts down
- zero is terminal — late await returns immediately
- countDown in the cleanup path, always
- latch of 1 = starting gun for all workers
- arrive-and-continue, not arrive-and-wait
basics
~20 sA latch is created with the count N. Each worker decrements it once when done; the coordinator blocks in await until the count reaches zero, then continues. The count never resets, so any wait after zero returns immediately instead of blocking.
solid answer
~50 sA countdown latch holds one non-negative counter fixed at construction time. Two operations: `countDown` decrements by one and never blocks, and `await` blocks until the counter is zero. For "wait for N tasks", create the latch with N, hand it to the workers, have each worker count down in a finally-style cleanup path so it fires even on failure, and have the coordinator call `await` (ideally with a timeout). When the last worker decrements, every waiter is released at once. The key property is that a latch is **one-shot and terminal**: zero is an absorbing state. A thread that awaits after the count already hit zero returns immediately rather than blocking — which makes the latch usable as a "has initialization finished?" gate for threads that arrive late. If you need the same rendezvous again for a second round, you need a *new* latch (or a reusable cyclic barrier instead).
code
text · 12 lineslatch = Latch(N)
worker(i):
try:
result[i] = compute(i)
finally:
latch.countDown() // fires even if compute() throws
coordinator:
if not latch.await(30s):
fail("only " + (N - latch.count) + " of " + N + " workers finished")
combine(result) // all writes to result[] are visible herego deeper
Know the two operations, the wait-for-N-tasks shape, and that the count never goes back up. Say explicitly that awaiting an already-open latch returns immediately.
Add the failure-path discipline (count down in a cleanup block, await with a timeout) and the start-gate trick with a latch of one. Be able to contrast arrive-and-continue with a barrier's arrive-and-wait.
Talk about the visibility guarantee (work before countDown is visible after await), about turning hangs into diagnosable timeouts, and about when a fixed N is the wrong model because task count is dynamic.
Frame it as a choice among completion-signalling mechanisms: a latch is fine for a fixed, known fan-out but does not compose or report partial failure. At scale you usually want a completion abstraction that carries results and errors, with the latch reserved for simple lifecycle gates.
## What a countdown latch is A countdown latch is one of the simplest coordination primitives: a single non-negative integer counter, fixed at construction to some N, plus two operations. - **countDown()** — decrement the counter by one, if it is not already zero. It never blocks. When the decrement takes the counter from one to zero, all threads currently blocked in `await` are released. - **await()** — block until the counter is zero. If it is already zero, return immediately. That is the whole primitive. There is no way to increase the count, and no way to reset it. Zero is an *absorbing* state: once reached, the latch stays open forever. ## The canonical use: waiting for N tasks ``` latch = new Latch(N) for i in 1..N: spawn(worker_i): try: do_work(i) finally: latch.countDown() coordinator: latch.await(timeout) aggregate_results() ``` Two details matter more than they look. **Count down on every exit path.** If a worker throws, returns early, or is cancelled without decrementing, the counter never reaches zero and the coordinator waits forever. The decrement belongs in a cleanup block that runs unconditionally — not at the bottom of the happy path. **Await with a timeout.** An unbounded wait converts one lost decrement into a permanently stuck thread with no diagnostic. A bounded wait lets you log "N of M workers reported in" and fail loudly. ## The mirror-image use: a start gate Because `await` releases *all* waiters simultaneously, a latch initialized to **1** works as a starting gun. Every worker awaits the gate latch; the coordinator does its setup and then counts down once, releasing all workers at nearly the same instant. This is the standard way to reduce startup skew when you want to measure contention: without it, the first worker may finish before the last one has even started. A benchmark harness often uses two latches at once — a start gate of 1 and a done latch of N — which is the clearest illustration that a latch is a directional, one-shot signal rather than a general rendezvous. ## Arrive-and-continue vs arrive-and-wait The defining asymmetry of a latch is that the counting party and the waiting party are different roles. `countDown` is *arrive-and-continue*: a worker signals its arrival and keeps going. `await` is pure waiting: the waiter contributes nothing to the count. A cyclic barrier, by contrast, is *arrive-and-wait* — the same call both registers your arrival and blocks you — so every participant is symmetric. This is why a latch cannot express "all N workers meet here before any continues": each worker would have to count down *and* await, and the last one to count down would be released along with everyone else, which happens to work once, but the latch cannot be used for a second round. ## One-shot and terminal — the consequences 1. **Late waiters do not block.** A latch is therefore a durable record that an event has happened, not an edge-triggered notification you can miss. That makes it the natural primitive for "service is initialized" or "configuration is loaded" gates: components starting up at any later time can await and proceed immediately. 2. **No reuse.** Round two needs a fresh latch. Code that tries to reset a latch is either using the wrong primitive (it wants a cyclic barrier or a phaser) or is about to introduce a race between resetting and awaiting. 3. **Extra decrements are harmless but hide bugs.** Counting down more than N times typically saturates at zero rather than going negative, so a double-decrement bug shows up as the coordinator proceeding *early* — with partial results — rather than as a crash. Treat "latch released but results incomplete" as a symptom of exactly this. ## Ordering guarantees Every reasonable latch implementation guarantees that everything a worker did *before* its `countDown` is visible to a thread that returns from `await`. Without that guarantee the primitive would be useless: the coordinator would be released and then read half-written results. Conceptually, each `countDown` *happens-before* the return of each `await` — the latch is both a control signal and a data-visibility boundary. This is why you do not need any additional synchronization around results that each worker wrote to its own slot before counting down. ## When it is the wrong tool If you need repeated rounds, use a cyclic barrier. If you need to bound *concurrent* access rather than wait for completions, you want a semaphore. If the number of tasks is not known until they start spawning, a fixed N is fragile — a dynamically registering phaser or a completion-counting mechanism fits better.
- What happens if one worker throws an exception before reaching its countDown call?The counter never reaches zero, so the coordinator blocks forever on an unbounded await — a silent hang with no error surfaced. The fix is to put countDown in a cleanup block that runs on every exit path, and to record the failure separately so the coordinator can distinguish "all finished" from "all reported in, some failed". Using a bounded await as a second line of defence turns a hang into a diagnosable timeout.
- Can you reuse a latch for a second round of the same N tasks?No. The count only decreases and zero is terminal, so a second round's awaits return immediately without waiting for anything. You either construct a new latch per round or switch to a cyclic barrier or phaser, which are designed to reset for the next generation. Code that attempts to reset a latch is a strong signal that the wrong primitive was chosen.
A latch is a turnstile that is locked until a fixed number of tickets have been dropped in — and once it unlocks, it stays unlocked. Everyone already queued walks through together, and anyone who shows up an hour later walks straight through without stopping.
saying these in an interview costs you the question
- Saying you can reset or increment a latch to reuse it for the next round.
- Putting countDown only on the success path, so a failing worker hangs the coordinator forever.
- Claiming a thread that awaits after the count reached zero will block until someone counts down again.
- Confusing a latch with a semaphore — a semaphore's permits go up and down to limit concurrency; a latch only counts down to signal completion.
- Adding extra locks around results because they "might not be visible", not knowing the latch itself provides the ordering guarantee.