skip to content

Eighteen workers park on wait-with-deadline calls against a shared pool of twenty connections, request-path lookups stall, and the store reports no load — why?

level: seniorimportance: must knowfreq 52%

answer

  1. the pool is exhausted, not the store
  2. waiters hold slots, not execution
  3. count slots against waiters
  4. the gap shows in pool-acquire wait
  5. separate pool for waiting callers

basics

~20 s

The scarce resource is the pool, not the store. Eighteen parked callers hold eighteen slots for their whole wait while spending no server processing time, so the store is genuinely idle and the request path queues for two connections.

solid answer

~50 s

A parked caller occupies a connection slot from the moment it issues the call until the value arrives or the server-side wait deadline passes, and in that time it costs the server nothing. So eighteen waiters remove eighteen of twenty slots permanently — a waiter that times out simply re-parks — and the request path is left with two. If each lookup takes roughly half a millisecond of the caller's wall time, mostly one round-trip time within the zone, two slots serve at most about four thousand lookups a second with zero slack, and everything above that turns into pool-acquire wait. Meanwhile the server's service time, its operations per second and its processor use are all flat, because the store really is idle. The remedy is to stop the two workloads sharing: give the waiting callers their own pool, cap concurrent waiters below the pool size, and keep the server-side wait deadline finite.

go deeper

for a junior

Recall that a waiting caller keeps its connection for the whole wait. If the number of waiters approaches the size of a shared pool, other work in the same process has nowhere to run.

for a middle

Explain why the store's signals stay flat: waiters are registered rather than executed, so service time and operations per second never move while the caller's pool empties. Do the slots-against-demand arithmetic.

for a senior

Diagnose it under pressure — separate the pool-acquire wait from the operation, reason with the high percentile instead of the mean, and reach first for a separate pool for the waiting callers rather than a bigger store.

for a principal

Make it a standing rule rather than an incident: waiting workloads do not share connections with serving ones, concurrent waiters are capped, wait deadlines are bounded, and a trip and slot budget exists per request path before anyone parks anything.

## Where the scarce resource actually is The instinct on a slowdown is to look at the store. Here the store is innocent and its own signals will say so convincingly. A caller parked on a wait-with-deadline call holds a **connection slot** for the entire wait and spends **no server processing time**; the wait was registered, not executed. Eighteen such callers take eighteen of the pool's twenty slots and keep taking them, because a waiter whose deadline expires re-checks and parks again. That leaves two slots for everything else in the process that shares the pool. ## The arithmetic Say the request path needs about four thousand lookups per second, the tier is one **cross-zone hop** away, and each lookup costs microseconds of the **server's service time** plus one **round-trip time** — call it half a millisecond of the **caller's wall time** end to end. - Two slots, each able to carry one call at a time, complete at most `2 / 0.0005 s = 4000` lookups per second, and only with no slack whatsoever. - Every lookup above that rate does not get slower on the store. It waits for a connection. The **pool-acquire wait** becomes the dominant term in the request's latency, and once the pool's own acquire limit is reached, callers start failing before they send anything at all. - Adding store capacity changes none of these numbers. The store was never involved. ## Why every store-side signal is flat | Signal | What it shows | Why |---|---|---| | Server service time per operation | Unchanged, microseconds | Lookups still execute exactly as fast | | Operations per second | Unchanged or lower | Waiters issue one operation and then nothing | | Processor use on the store | Flat | Registered waits consume no cycles | | Connected clients | Elevated | The only store-side signal that moves | | Caller wall time on lookups | Elevated, in the tail first | The added time is spent before the call is sent | This is also why the throughput ceiling is the wrong number to quote. *The store handles a hundred thousand operations a second, so it cannot be the bottleneck* answers a capacity question with a capacity fact, when the question being asked is about fairness between two workloads sharing one pool. And it is why the mean is the statistic guaranteed to miss it: at first only the calls unlucky enough to arrive with no free slot pay, and the mean absorbs them. The high percentile moves immediately, and it moves in the pool-acquire wait, which most instrumentation never separates from the operation. ## Diagnosing it from the caller's side 1. **Record the pool-acquire wait separately from the operation.** If acquire time is large while the operation's own time is unchanged, the store is not your problem. 2. **Compare the server's service time with the caller's wall time.** A widening gap with a flat server number localises the cost to your side of the wire. 3. **Count parked callers against pool size.** This is the actual capacity equation, and it is rarely on a dashboard anywhere. 4. **Watch the high percentile, not the mean**, and state the window. ## Remedies, in the order worth trying 1. **Give the waiting callers their own pool.** Two workloads with opposite occupancy profiles should not compete for the same slots. This is the standard remedy and it fixes the problem rather than deferring it: the request path then cannot be starved by waiters at all, whatever they do. 2. **Cap concurrent waiters** below the number of slots they are allowed. A cap makes the worst case arithmetic rather than emergent. 3. **Keep the server-side wait deadline finite**, so any one slot's occupancy is bounded and a waiter that has silently lost its wakeup re-checks instead of sitting forever. 4. **Stop parking.** Where the waiting population is large, re-asking on a schedule holds a slot only for each round trip, and the interval you pick is a latency cost you can see. ## What does not fix it - **A bigger store, or a faster one.** Nothing about the store is saturated. - **Switching execution models.** Waiters occupy no execution under either a store that executes one operation at a time or one that serves requests from a thread pool, so neither model helps. Where a server dedicates a thread per connection, changing models can even add a second limit to trip over. - **A shorter caller-side deadline on each lookup.** Starved callers then fail faster, which is sometimes desirable, but the same number of requests still fails. - **More retries.** Retries need connections. They queue behind exactly the same two slots.

  • From the caller's side, how do you tell a starved pool from a slow store?
    Record the pool-acquire wait separately from the operation and compare the server's service time with the caller's wall time. A large acquire time with an unchanged operation time and a flat server number puts the cost on your side of the wire; counting parked callers against pool size confirms it.
  • Does the store's execution model change this outcome?
    No. A parked caller occupies no execution under either a store that executes one operation at a time or one that serves requests from a thread pool, so neither relieves the pool. The one variation is a server that dedicates a thread per connection, where waiters also consume a server-side thread budget.
  • Why does a finite server-side wait deadline help even though waiters re-park?
    It bounds the occupancy of any single slot and forces a periodic re-check. That makes the waiting population measurable, gives the pool regular moments where slots are released, and puts a ceiling on how long a caller whose wakeup went missing sits there believing it is still waiting.

saying these in an interview costs you the question

  • Blames the store because the request path slowed down.
  • Cites the store's operations-per-second ceiling to dismiss a fairness problem.
  • Reports the mean lookup latency as evidence that nothing is wrong.
  • Adds more waiters because the store's own signals look idle.
  • Measures the operation but never the wait for a pooled connection.
  • Expects a faster or larger store to relieve the starved request path.