A caller requests a thirty-second server-side wait while its client library caps every operation at five seconds — what happens?
answer
- four clocks, not one
- the shorter deadline is the real one
- caller-side against server-side wait
- abandoned call, discarded connection, connect cost
- polling with extra steps
basics
~20 sThe inner deadline wins. Every wait longer than five seconds is abandoned by the caller, which usually discards the connection and retries, so the design degrades into polling at five-second granularity plus a connect cost per attempt.
solid answer
~50 sFour separate clocks are in play — the connect cost, the pool-acquire wait, the caller's deadline on one operation, and the server-side wait deadline — and here the third is shorter than the fourth. Whenever the value does not arrive within five seconds the caller's deadline fires first: the library gives up on a call the server still considers live, and since a reply may still arrive on that connection it usually cannot be reused, so it is discarded and re-established. The design becomes polling at five-second granularity at a connect cost per poll, with connection churn and a stream of failures that are not failures. The fix is ordering: the caller's deadline on that one operation must exceed the server-side wait deadline plus a round-trip time and slack, which usually means giving the wait call its own connection or pool.
go deeper
Recall that a call which waits has two deadlines around it — the caller's own and the one handed to the server — and that the shorter of the two is the one that actually takes effect.
Explain the loop the mismatch produces: the caller gives up first, the connection usually cannot be reused, a new one is established, and the design becomes polling at the caller's deadline with a connect cost attached.
Show that you would separate the wait onto its own connection or pool rather than loosening one global deadline, and that you treat an empty return at the deadline as the normal path, not an error to alert on.
Own the relationship as a written rule between two settings that different teams configure: state the ordering, state that the wait deadline must be finite, and require the connect cost to be measured so the failure is visible when someone breaks it.
## Four clocks, not one The word people reach for here collapses four independent numbers with four different failure modes: 1. **The connect cost** — what it takes to establish a new connection to the store. 2. **The pool-acquire wait** — how long a caller will wait for a connection from the pool before giving up. 3. **The caller's deadline on one operation** — how long the caller will wait for a reply to a call it has already sent. 4. **The server-side wait deadline** — the value the caller hands the store, telling it how long to hold the request open before answering empty. The first three are the caller's and are usually set once, for everything. The fourth is an argument of this particular call. A wait-with-deadline call is the one place where the third and the fourth meet, and getting their order wrong turns a working design into a churn machine. ## What happens when the inner deadline is the shorter one With a five-second caller-side deadline on one operation and a thirty-second server-side wait deadline requested: 1. The caller issues the call. The server registers the waiter and holds the request open. 2. If the value arrives in under five seconds, everything works, and the bug stays hidden. This is why the mistake reaches production. 3. If it does not, the caller's deadline fires at five seconds. From the caller's point of view the operation failed. 4. The library abandons the call. It cannot simply reuse the connection, because the server still owns that request and may write a reply onto it at any moment; a subsequent call would read the wrong reply. So the usual handling is to discard the connection. 5. The server discovers nothing at that instant. Its side of the wait ends only when it notices the peer is gone — which may be when it tries to answer. 6. The caller opens a new connection, paying the connect cost, and parks again. Run that loop and the caller has re-implemented re-asking on a schedule, at five-second granularity, at a connect cost per ask, with connection churn on the store thrown in. It is worse than either of the two honest designs. | | What you asked for | What you get | |---|---|---| | Trips for a thirty-second wait | one | about six, each with a connect cost | | Delay after the value is written | one round-trip time | one round-trip time, if it lands inside a live window | | Connections used | one | a new one every five seconds | | What the caller logs | nothing | a stream of failures that are not failures | ## Getting the layering right - **Order the two deadlines explicitly**: the caller's deadline on that operation must be greater than the server-side wait deadline plus one round-trip time plus slack. Write the relationship down next to both numbers, because the two are usually configured in different places by different people. - **Give the wait call its own connection or pool.** A single caller-side deadline applied to every operation cannot simultaneously be tight enough for a microsecond lookup and loose enough for a thirty-second wait. Separating them also keeps waiters from taking slots the request path needs. - **Keep the server-side wait deadline finite.** An unbounded wait removes the reconnect loop but replaces it with a slot occupied indefinitely and a lost wakeup that nothing ever surfaces. A bounded wait in a loop is the shape you want: it returns empty, the caller re-checks and re-parks. - **Do not read an expiry as an error.** An empty return at the deadline is the normal outcome of the call and belongs in the loop, not in the error path or the alert. - **Make the connect cost visible.** If it is not measured, this failure mode presents as a vague latency complaint with a perfectly healthy store behind it. ## What varies How the client library handles an abandoned call is not a constant: some discard the connection, some attempt to drain the pending reply and reuse it, some expose a separate deadline for calls known to wait. None of that is a property of the store, and none of it should be assumed from memory — check the library in front of you. On the store's side, whether an expired wait is distinguishable from a wait cut short by a lost peer also varies. What does not vary is the arithmetic: if the caller stops waiting before the server does, the server's wait deadline was never in effect, and whatever you configured it to is decoration.
- Why can the client library usually not reuse the connection after abandoning a wait?Because the server still owns the request. It may write the reply onto that connection at any moment, and the next call sent down it would read the wrong reply. Discarding the connection is the safe handling, which is exactly why abandoning a wait costs a connect cost each time.
- Why not simply ask for an unbounded server-side wait?It trades one failure for a worse one. The reconnect loop disappears, but a connection slot is now occupied indefinitely, and a wakeup lost to a broken connection is never surfaced because nothing ever returns. A bounded wait inside a loop gives the same behaviour with a floor under both problems.
saying these in an interview costs you the question
- Configures one caller-side deadline for every operation and parks on that connection.
- Believes the server shortens its wait to match the caller's deadline.
- Reads each abandoned wait as a store failure and adds retries on top.
- Removes the server-side wait deadline so the wait can never expire.
- Ignores the connect cost paid on every torn-down and re-established wait.