An in-memory store stalls for 200 ms a few times a minute, yet its mean server service time and operations-per-second counter look normal; why?
answer
- millions of samples, one big one
- delayed work is still counted
- queueing is not service time
- affected fraction picks the percentile
- measure it on the caller's clock
basics
~20 sAveraging buries a rare stall: one 200 ms sample among millions of microsecond samples barely moves the mean, and the delayed operations are still counted, only later. The waiting shows in the callers' wall time, at a high percentile.
solid answer
~50 sThree separate effects hide it. First, **averaging**: at 30,000 operations per second the store takes 1.8 million samples a minute, so three 200 ms samples raise the mean by well under a microsecond. Second, the **counter measures completed work, and the work was delayed rather than lost** — over a one-minute window the total is unchanged; only at one-second granularity do you see a notch and a catch-up spike. Third, and most important, the **queueing wait is not in the server's numbers at all**: every delayed operation still executed in microseconds, so a distribution of service times is structurally blind to the time its callers spent waiting for a turn. That time exists only in the **caller's wall time**, and which percentile carries it is arithmetic — the affected fraction here is about 1%, so the 99th percentile sits right at the edge and a higher one shows the stall in full.
go deeper
Remember that an average can hide a rare bad event completely, because it is divided by everything else that happened. A store looking fine on average does not mean every caller was served quickly.
Explain the dilution with numbers: a few long samples among millions of short ones cannot move a mean, and delayed work still gets counted when it finally runs. Name the caller's wall time as the place to look.
Show the structural point — a service-time distribution cannot contain queueing time — and derive the percentile from the affected fraction instead of reaching for the 99th by habit. Say which measurement you would add.
The decision is what the tier's access contract promises and on whose clock it is measured. Set the statistic and the window before the incident, and require caller-side measurement so a dashboard cannot certify a tier its users are abandoning.
## What the server's numbers actually record A store of this class typically reports how long it spent executing operations and how many it completed. Both are honest, and neither is capable of showing what an expensive call did to everyone else. Understanding *why* is the point of this question, and it is the difference between believing a dashboard and reading one. The key distinction is between **server service time** — what the store spent executing one operation — and the **caller's wall time**, which is the service time plus the round-trip time plus any time the operation spent waiting for its turn. The damage from an expensive call is almost entirely in that last term, and the last term appears in neither of the server's two numbers. ## Why the mean cannot move Take a store handling 30,000 operations per second with a typical service time of about ten microseconds, stalled three times a minute by a 200 ms call. - Samples in the window: 30,000 x 60 = **1.8 million**. - Extra execution time contributed by the stalls: 3 x 200 ms = **600 ms**. - Effect on the mean: 600 ms spread over 1.8 million samples is about **0.33 microseconds** — a mean of 10.0 microseconds becomes 10.33. The mean is a ratio with a very large denominator, and rare events cannot survive it. This is not a flaw in how the mean was computed; it is what a mean is for. ## Why the throughput counter stays flat The queued operations were not dropped. They ran late, and they were counted when they ran. Over a one-minute window the store completed the same total it always does. At **one-second granularity** the shape appears: the second containing the stall is short by a few thousand operations and the following second is long by roughly the same amount, because the backlog drained into it. A counter is a window-length question, and at the window most dashboards default to, the notch and the spike cancel. ## Which percentile carries it, and how to pick it The affected fraction is arithmetic, not intuition: > operations delayed per window = stall duration x arrival rate x stalls per window Here: 0.2 s x 30,000 per second x 3 = **18,000 delayed operations** out of 1.8 million, which is **1%**. That number tells you exactly where to look: | Statistic | What it shows | Why | |---|---|---| | Mean server service time | essentially nothing | 600 ms diluted across 1.8 million samples | | Operations per second, one-minute window | nothing | delayed work is still completed work | | Operations per second, one-second window | a notch then a spike | the stall, then the backlog draining | | 99th percentile of callers' wall time | the edge of the effect | exactly 1% of operations were delayed | | Higher percentiles of callers' wall time | the full stall | the worst-delayed callers waited the whole duration | If the stall happened once a minute instead of three times, the affected fraction would be about 0.33% and the 99th percentile would be clean while a higher one carried everything. This is why "watch p99" is advice, not a rule: the percentile you need is the one above the affected fraction, and the affected fraction is something you calculate. ## Two clocks, one of which is blind - **Server-measured service time** is necessary for identifying *which* operation is expensive — it is where the 200 ms sample lives. - **Caller-measured wall time** is the only place the consequences live, because the consequence is waiting, and waiting is not executing. A team that measures only the first concludes the tier is healthy while its callers are timing out; a team that measures only the second knows something is wrong but not which call. The pair is the instrument. ## What varies between stores What is exposed differs, and assuming your store's reporting is the class's reporting is its own error. Some stores keep a record of individual calls whose execution exceeded a threshold, which is the fastest way to identify the offending operation when it exists. Some expose per-operation service-time statistics; some expose very little. **None** of them, in any store of this class, contains the waiting time of the callers that queued — that measurement can only be taken on the caller's side. Under a thread-pool execution model the same blindness applies, with a smaller affected fraction to find: capacity dipped rather than stopped, so the delayed set is smaller and the percentile you need is correspondingly higher.
- How do you choose which percentile to alert on?Calculate the affected fraction: stall duration times arrival rate times stalls per window, divided by operations per window. The percentile you need sits above that fraction. At 1% affected, the 99th percentile is on the boundary and a higher one carries the signal; at 0.1%, the 99th shows nothing at all and only a far higher percentile does.
- The store reports a fast call and the caller reports a slow one. Who is right?Both. The caller's operation executed in microseconds once it started; the rest of its wall time was spent queued, plus the round-trip time. The two numbers measure different intervals, and the gap between them is precisely the quantity that an expensive call inflates for everybody else.
saying these in an interview costs you the question
- Concludes the tier is healthy because its mean service time is flat.
- Hunts for the delay in server-side service time, where queueing never appears.
- Takes a flat operations-per-second counter as proof nothing was delayed.
- Picks the 99th percentile without checking what fraction was affected.
- Reports the delayed callers as slow operations rather than as waiting ones.