One call keeps an in-memory store busy for 200 milliseconds: who is delayed if the server executes one operation at a time, and who if it serves requests from a thread pool?
answer
- who is behind it, not who called it
- the execution model decides the radius
- one at a time versus a worker pool
- capacity falls by one worker's share
- delayed callers' own operations stayed fast
basics
~20 sWhere the server executes one operation at a time, every other caller waits the full 200 ms and then the backlog drain. Where a thread pool serves requests, one worker is lost and only callers needing the same lock block.
solid answer
~50 sThe answer depends entirely on the execution model, which is why the first move is to establish it. On a store that **executes one operation at a time**, the expensive call is not interruptible: every operation that arrives during those 200 ms queues behind it, and after it finishes the store still has the backlog to drain, so callers of ordinary microsecond operations see tens to hundreds of milliseconds of wall time. On a store that **serves requests from a thread pool**, one worker of several is occupied for 200 ms: capacity falls by roughly that worker's share, and callers block only where they contend — for the same entry's lock, or for a shared structure the store locks internally. The second model divides the damage rather than removing it: enough concurrent expensive calls to occupy every worker reproduces the first model exactly.
go deeper
Take away the shape: one expensive call can make other people's fast calls slow, because they are waiting for their turn rather than being slow themselves. Knowing that much beats guessing at the mechanism.
Explain the queueing: work arriving during the stall piles up and must drain afterwards, so the delay outlasts the expensive call. Be able to multiply arrival rate by stall duration to say how many callers were touched.
Ask which execution model before answering, and give both branches with the contention caveat for the multi-threaded one. Then say what the delayed callers' own numbers will and will not show, and why retries make it worse.
Frame it as shared fate: on a store executing one operation at a time, worst-case latency for everyone equals the most expensive call any one team may issue. Decide whether that is acceptable, or whether the tier needs a bound or a boundary.
## The question underneath the question "A call takes 200 ms — what happens to everyone else?" has no single answer, and answering it as though it does is the most common way to be wrong about this tier. The blast radius of one expensive operation is decided by how the server executes operations, and both designs are real and in production use. Establish the model first; the arithmetic follows from it. One thing is true in both models and worth saying up front: the delayed callers' **own** operations were never slow. Each still executed in microseconds of **server service time**. What they accumulated was waiting, which shows only in the **caller's wall time**. ## Where the server executes one operation at a time The expensive call runs to completion before anything else runs. It is not pre-empted, not time-sliced and not interleaved. Suppose ordinary operations arrive at 30,000 per second and each takes about ten microseconds, leaving the store around 30% busy. During the 200 ms stall: 1. roughly **6,000 operations arrive and queue**; 2. a caller whose operation arrived at the start of the stall waits close to the full 200 ms, one arriving at the end waits almost nothing, and the average wait across them is around **100 ms**; 3. once the expensive call finishes, the backlog still has to drain — 6,000 operations at ten microseconds is about **60 ms** of additional catch-up, during which newly arriving callers queue behind the backlog too. So one caller's single call spent the store's entire capacity for a fifth of a second, and the cost landed on several thousand unrelated callers. ## Where a thread pool serves requests Here the expensive call occupies one worker. With four workers, capacity drops by roughly a quarter for 200 ms; with sixteen, by a sixteenth. Unrelated callers whose operations land on free workers proceed at normal speed. What they do *not* escape is contention. A multi-threaded store of this class still has to make each operation atomic, and it does that with locking rather than with a single execution thread. The **granularity of that locking varies**: some lock per entry, some lock a partition of the key index, some take a coarser lock — and the coarser it is, the more the second model behaves like the first. Callers touching the same entry the expensive call is walking are the ones most likely to stop. The model also converges on the first one under load. Four concurrent expensive calls on four workers leaves nothing serving, and the head-of-line picture returns in full. | | Executes one operation at a time | Serves requests from a thread pool | |---|---|---| | Who is delayed | every caller with an operation queued behind it | callers contending for the same entry or shared structure; everyone once workers run out | | How long | the full duration, plus the backlog drain | a share of throughput for the duration | | Capacity during the stall | effectively zero | reduced by roughly one worker's share | | What it takes to hurt | one expensive caller is enough | enough concurrent expensive calls to occupy the pool | | What the delayed caller's own service time shows | microseconds, unchanged | microseconds, unchanged | ## What neither model changes - **The work is delayed, not lost.** Every queued operation eventually runs, which is why counters of completed work recover. - **The caller's deadline on one operation may fire first.** A caller that allows 50 ms for a call it queued behind a 200 ms operation gives up, and from its point of view the store failed — with nothing in the store's own numbers to corroborate it. - **Retrying makes it worse.** Callers that abandon and re-issue add arrivals to a store that is already behind. - **Placement is additive, not causal.** A cross-zone hop adds its round-trip time to every one of these waits, but it did not create the stall. ## The design consequence On a store that executes one operation at a time, the tier's worst-case latency for **every** caller is the duration of the most expensive operation any single caller is allowed to issue. That sentence is the reason the expensive class matters at all, and it is what makes a bound on such calls a shared-fate decision rather than one team's optimisation. On a thread-pool store the same sentence holds with "every" replaced by "a share of", which is better but is not immunity — and knowing which sentence applies to your store is the whole of the answer.
- Does a delayed caller experience this as its own operation being slow?No, and that is what makes it hard to trace. Its operation still executed in microseconds; the extra time was spent queued, before execution began. The delay lives in the caller's wall time and in nothing the server records about that operation, so a caller reporting a slow call and a server reporting a fast one are both telling the truth.
- Four expensive calls arrive together on a store with four worker threads — what changes?Every worker is occupied, so capacity for that period is effectively zero and unrelated callers queue exactly as they would on a store executing one operation at a time. The thread pool bought a factor, not an exemption: it converts one caller's stall into a fraction of capacity, and that fraction disappears once concurrent expensive calls match the worker count.
- Does it matter whether the tier is a single node or partitioned across several?Yes: on a partitioned tier the stall is contained to the node that holds the entry being walked, so callers whose keys live elsewhere are untouched. That narrows the radius but does not shrink it for anyone sharing that node, and a caller whose request needs keys on several partitions is exposed to the slowest of them.
A supermarket checkout. With one till open, the customer with three hundred items delays everyone in the queue for as long as the scan takes, whatever the person behind is buying. Open four tills and that customer occupies one of them: the queue moves at roughly three-quarters speed, and only shoppers who need the single price-check terminal he is using actually stop. Send in four three-hundred-item customers at once and all four tills are busy — the one-till picture is back.
saying these in an interview costs you the question
- Says one expensive call only affects the caller that made it.
- Asserts that every caller waits, without asking how the server executes operations.
- Thinks a thread pool removes the problem rather than dividing it.
- Cites a throughput ceiling to argue that one call cannot delay anyone.
- Believes the delayed callers' own operations were slow at the server.
- Recommends retrying the abandoned calls, which adds load to a store already behind.