skip to content

An in-memory store takes 40,000 calls per second, each holding a connection 0.4 ms of wall time - how many connections does the pool need?

level: middleimportance: must knowfreq 62%

answer

  1. concurrency, not rate
  2. how long a slot stays held
  3. arrival rate times hold time
  4. microsecond calls need few connections
  5. placement multiplies it, traffic does not

basics

~20 s

About sixteen, plus a margin. Pool size follows concurrent calls in flight, which is arrival rate times hold time - 40,000 per second times 0.0004 seconds - not the request rate, so microsecond-scale calls need few connections.

solid answer

~40 s

The number of connections in use at any instant is the arrival rate multiplied by how long each call holds one: 40,000 per second times 0.4 ms gives 16 in flight, so a pool of roughly 16 plus a margin covers it. Rate alone tells you nothing, because the same rate over a hold time of 2 ms would need 80. That is the arithmetic peculiar to this tier: hold time is dominated by one round trip, so it is small, and a small pool serves a very large throughput. The lever is hold time, not traffic - moving the tier a zone further away, or holding the connection while the caller deserialises the reply, multiplies the pool needed at unchanged traffic. Oversizing buys nothing and consumes slots against the store's connection ceiling.

go deeper

for a junior

Remember that a pool holds connections for reuse and that a few of them can serve a great many calls per second, because each call gives its connection back within a fraction of a millisecond.

for a middle

Do the arithmetic aloud: arrival rate times hold time gives calls in flight, and that is the pool. Then name what is inside hold time, including the caller's own work before the connection is released.

for a senior

Show that undersizing looks nothing like a slow store: the caller's wall time rises with the server's service time flat, the tail moves before the mean, and the added time sits ahead of the request on the wire.

for a principal

Treat hold time as a design variable the organisation controls - placement, call shape and what callers do before releasing - and set per-instance pool maxima that a fleet at full scale can still fit under the tier's ceiling.

## The number that actually sets pool size Pool size is a concurrency question, not a throughput question. The count of calls in flight at any instant is the arrival rate multiplied by how long each call holds a connection - the relationship known as **Little's law**. For the numbers given: 40,000 calls per second x 0.0004 seconds of hold time = **16 calls in flight**. So the pool needs about 16 connections, plus a margin for jitter and for bursts above the average rate. Call it 24. A service handling a very large request rate against this tier needs a strikingly small pool, and that is not a trick of the example - it is the defining property of a tier whose calls are measured in fractions of a millisecond. The rate on its own answers nothing. The same 40,000 calls per second needs 4 connections if the tier is co-located and 80 if it is a zone away. ## What hold time actually covers Hold time is the interval between taking a connection out of the pool and putting it back, which is usually longer than people assume: - the **round-trip time** carrying the request out and the reply back; - the store's **service time**, normally the smallest part; - reading and parsing the reply on the caller; - **any caller-side work done before the connection is returned** - deserialising into domain objects, logging, or worse, making another call downstream while still holding it. This is the most common sizing defect and the cheapest one to fix: release first, then process. - for a **wait-with-deadline call**, the entire wait. A parked caller does no server work at all, but the slot it holds is a slot no one else can use, so its demand must be counted apart from ordinary calls rather than blended into this arithmetic. ## Hold time is the lever At a constant 40,000 calls per second: | Call shape and placement | Hold time | Calls in flight | Pool with margin | |---|---|---|---| | Co-located single-key read | 0.1 ms | 4 | about 8 | | Same-zone single-key read | 0.4 ms | 16 | about 24 | | Same-zone read, 1 ms of caller work before release | 1.4 ms | 56 | about 70 | | Cross-zone hop single-key read | 2 ms | 80 | about 100 | | A parked caller, waiting up to five seconds | up to 5,000 ms | one slot per parked caller | budget separately | Two readings of that table matter. First, **placement multiplies the pool** - the same traffic needs five times the connections a zone away, which is a topology decision showing up as a capacity number. Second, **the caller's own code is in the hold time**, so a change with nothing to do with the store can double the pool it needs. ## Why a bigger pool does not buy more - On **a store that executes one operation at a time**, extra connections buy no extra execution: the operations queue at the server regardless of how many sockets delivered them. - On **a store that serves requests from a thread pool**, extra connections buy real parallelism, but only up to its worker count; past that they queue too. - Every connection, busy or idle, consumes a slot against **the connection ceiling** and carries per-connection buffers on the server. - An oversized pool **hides the caller's real concurrency**. If something makes calls slow, a large pool lets one instance open every connection it is permitted, turning a latency problem into a ceiling problem. - Per-instance pools **multiply by fleet size**, so a generous per-instance number is never a local decision. ## How undersizing announces itself 1. The **caller's wall time** rises while the **server's service time** stays flat - the store is not the thing that changed. 2. The added time appears **before the request goes out**, as the pool-acquire wait; a trace shows the gap between issuing the call and the first byte on the wire. 3. The **mean barely moves while a high percentile (the 99th) moves a lot**, because waiting for a slot is a queue and queues punish the tail first. 4. At the limit the pool-acquire wait exceeds **the caller's deadline on one operation**, and the call fails without the store ever seeing it. ## Where designs differ The arithmetic is stable but its unit is not. Some **client libraries multiplex many in-flight operations over a single connection**, in which case a connection is not a unit of concurrency at all and the number you are sizing is in-flight operations rather than sockets. Others hand one connection to one operation at a time, which is what the calculation above assumes. Establish which model the library uses before quoting a pool size, and never inherit a library default as though it were a measured number.

  • Why does moving the tier one zone away change the pool size when the traffic has not changed?
    Because hold time, not traffic, sets concurrency. A cross-zone hop adds milliseconds of round-trip time to every call, and each call therefore occupies its connection several times longer. At a fixed arrival rate the number in flight scales directly with that hold time, so the same workload suddenly needs several times the connections - and several times the slots against the store's connection ceiling.
  • What is the cheapest sizing mistake to fix in a service whose pool keeps running out?
    Work done while the connection is still held. Deserialising the reply into domain objects, logging it, or calling something else before returning the connection all inflate hold time, and concurrency scales with hold time directly. Returning the connection as soon as the bytes are read often shrinks the pool requirement more than any tuning of the pool itself.

saying these in an interview costs you the question

  • Sizes the pool from request rate alone, ignoring how long a call holds it.
  • Assumes more connections always buy more parallel execution at the store.
  • Adopts the client library's default pool size and calls it sized.
  • Counts a parked caller's slot as if it were a sub-millisecond call.
  • Says an oversized pool is harmless because idle connections do nothing.