skip to content

A parked caller's connection is lost and re-established on another node after a failover — why might it wake with nothing?

level: seniorimportance: should knowfreq 36%

answer

  1. the wait lives on the connection
  2. no connection, no registration
  3. a loop, not a single call
  4. re-check the entry before re-parking
  5. no redelivery, no replay

basics

~20 s

A wait is state on the connection that registered it. When that connection dies the registration dies too, nothing on the new node knows the caller was waiting, and a value written across the gap may be gone or already taken.

solid answer

~40 s

The caller asked one node to hold its request open, and that promise lives on that connection. A failover, a restart, or a proxy moving the caller breaks it, and the registration goes too — the new node has no record that anyone was waiting. What crossed the gap depends on the store: where a write is acknowledged before a replica holds it, the value may never have reached the node now serving; where acknowledgement waits for a replica, it survives, and some stores decide this per write. A value written after the move goes to whichever caller re-parked first. So a waiter is a loop, not a call: reconnect, re-read the entry before parking again, and treat an empty return and a lost connection alike. This call carries no acknowledgement, redelivery or replay.

go deeper

for a junior

Recall that a wait is not a durable subscription. It belongs to one connection, and if the connection goes away the caller can come back to find nothing, even though something was written.

for a middle

Explain the three distinct causes — the registration dying with the connection, a value that may not have crossed to the node now serving, and another caller re-parking first — and why none of them is an error condition.

for a senior

Demonstrate the loop: reconnect, re-check the entry, park again with a finite deadline, tolerate duplicates, and keep the authoritative state in the keyspace so a missing wakeup costs latency rather than correctness.

for a principal

Draw the line for the organisation: workloads that lose work when a wakeup is missed do not belong on this mechanism at all, and the choice between it and a durable messaging system is made on that criterion, not on convenience.

## Where a wait actually lives When a caller parks, nothing durable has been created. The server has recorded, against one connection, that this caller is waiting on this key until a **server-side wait deadline** passes. It is per-connection state on one node. That single fact explains everything that follows. Break the connection and the registration is gone. There is no handle to hand over, and the node the caller lands on next — after a failover, a restart, a rolling upgrade, or a proxy placing it somewhere new — has never heard of the wait. Nor could it: the caller is not there to be woken. ## The three ways the wakeup goes missing 1. **The registration dies with the connection.** Even if the value is written a millisecond later on a node that is still healthy, nobody is registered to receive it from the caller that just moved. 2. **The value may not have crossed.** Here the stores in this class genuinely differ, and the difference matters: where a write is acknowledged to its author before any replica holds it, a value written just before the old node went away may simply not exist on the new one. Where the store acknowledges only once a replica has the write, it does exist. Some stores let the caller ask for the stronger behaviour per write. Never assert one of these as the model. 3. **Someone else took it.** A value written after the move goes to whichever caller is registered at that moment. If several callers park on the same key and one of them reconnects faster, it wins. Nothing redelivers it to the caller that was not there. ## Write the waiter as a loop, not as a call The whole point is that a parked caller must be able to wake up with nothing and carry on correctly. That means: - **Re-check the entry after reconnecting, before parking again.** The value may already be sitting there, written while the caller was away, with nobody registered to be told. A waiter that re-parks immediately can sit through a full deadline with the answer already in the keyspace. - **Treat an empty return and a lost connection as the same outcome.** Both mean *I do not know*, and both lead back to the top of the loop. Neither is an error worth an alert. - **Keep the server-side wait deadline finite.** An unbounded wait means a lost wakeup is never surfaced, because nothing ever returns. A bounded wait puts a ceiling on how long the caller believes a broken promise. - **Never let correctness depend on the wakeup arriving.** Use it for latency — to learn quickly, rather than to learn at all. The state in the keyspace is what is authoritative, and the caller must be able to reconstruct its position from it. - **Expect duplicates as well as gaps.** A caller that re-parks after a reconnect may be answered for something it had already partly handled before the connection broke; the handling has to tolerate that. ## What this mechanism is not This is the boundary that gets crossed most often on this subject, usually by accident, so state it plainly. | Property | Provided by a wait-with-deadline call | |---|---| | Wakeup when a value arrives, if you are connected | Yes | | Acknowledgement that the caller handled what it received | No | | Redelivery when the caller dies mid-handling | No | | Replay of what arrived while the caller was away | No | | A record that survives the caller's disconnection | No | A parked caller is not a broker consumer and this call is not a cheap message queue. Those guarantees are the business of durable messaging systems, which keep state independent of any connection for exactly this reason. If losing a wakeup would lose work, the design needs a mechanism built to survive a disconnection, not a wait with better handling around it. ## Designing around it - **Keep the authoritative state in the keyspace**, and use the wait only as a fast notification that it changed. A loop of re-check-then-park is then correct whatever happens to connections. - **Bound the deadline against how stale the caller may be.** If a lost wakeup can go unnoticed for a full deadline, the deadline is the freshness guarantee you are actually offering. - **Assume connections move.** Failover is the obvious case; rolling restarts, proxy reconfiguration and ordinary network faults produce the same result far more often. - **State the store's acknowledgement behaviour in the design.** Whether an acknowledged write is guaranteed to be on the node you land on is a property of the store you chose, and a design that silently assumes the stronger one will be wrong on half of this class. And the capability itself is not universal: a substantial part of this class offers no wait call at all, where the equivalent question becomes how stale the re-asking interval lets a caller be.

  • Why must the caller re-read the entry before parking again after a reconnect?
    Because the value may already be there. It can be written during the gap with no registration left to fire, so a caller that re-parks immediately waits out a full deadline with the answer sitting in the keyspace. Re-check first, then park only if there is still nothing.
  • What changes if the store acknowledges a write only after a replica holds it?
    The second failure mode narrows: a value the writer was told it had written is present on the node taking over, so the caller's re-check can find it. The lost registration and the race with other waiters remain, so the loop is still required — only the window of vanished values closes.

saying these in an interview costs you the question

  • Assumes a wait registration follows the caller to the new node.
  • Treats an empty return as proof that nothing was ever written.
  • Expects redelivery of a wakeup the caller was not there to receive.
  • Re-parks after reconnecting without re-checking the entry first.
  • Relies on a wakeup arriving for correctness rather than for speed.
  • Asserts that an acknowledged write is always present on the node taking over.