skip to content

Your anomaly-scoring fleet was sized so its 12,000-reading peak just fits the ceiling; nothing is dropped, yet p99 misses the 250 ms deadline - why?

level: seniorimportance: must knowfreq 57%

answer

  1. throughput is not latency
  2. queue wait plus service time
  3. the u over one-minus-u factor
  4. small per-instance queues, not one pool
  5. target is measured from the deadline

basics

~20 s

Queue wait. Sizing for throughput only checks that arrivals stay under the ceiling; a reading's latency is queue wait plus service time, and mean wait scales with utilisation as u / (1 - u), so a fleet at saturation queues heavily while its throughput still looks fine.

solid answer

~50 s

Throughput sizing and latency sizing ask different questions. Throughput sizing asks whether arrivals stay below the ceiling, and at 100% utilisation the answer is still technically yes - which is why nothing is dropped. But a reading's latency is **queue wait plus service time**, and queue wait is governed by utilisation: Kingman's approximation puts mean wait proportional to `u / (1 - u)`, so it is 1.5 service times at 60% utilisation and 9 at 90%, heading to infinity as utilisation approaches 1. A fleet sized to exactly meet peak sits on the steep part of that curve. The fix is to make the utilisation target an *output* of the deadline rather than an input: measure p99 against offered load, find the utilisation at which measured p99 still fits inside 250 ms, and size the fleet so peak lands there.

go deeper

for a junior

Remember that a reading's latency is the time it waits for a free slot plus the time it is worked on, and that only the second part is fixed. A fleet can be fast enough in total and still too slow per reading.

for a middle

Explain the shape: mean queue wait rises with utilisation as u over one minus u, so it is modest at 60% and severe near saturation. Name service-time variability as the multiplier on top.

for a senior

Show the measurement. Draw the p99-against-offered-load curve for this tier, pick the utilisation where the curve still clears the deadline, and size from that - rather than defending a round number after the fact.

for a principal

Make the ownership explicit. The utilisation target is where a detection deadline is traded against machine count, so it belongs with whoever owns the deadline, and it has to be re-derived whenever the served artifact changes.

## Two sizing questions that look like one "Can the fleet handle peak?" is two questions wearing one coat. - **Throughput sizing** asks: is the offered arrival rate below the fleet's ceiling? It compares two rates and produces the saturation count. A fleet sized this way drops nothing, and its dashboards look healthy. - **Latency sizing** asks: does a reading get scored inside its deadline at the 99th percentile? It compares a duration to a budget, and it produces a **utilisation target**. The scenario in the question is a fleet that passed the first test and was never given the second. That is the single most common capacity mistake on a scoring tier, because the first test is the one that is easy to compute. ## Where the milliseconds actually go End-to-end, a reading's time in the scoring tier is: ``` latency = queue wait + service time ``` Service time is fixed by the work: 40 ms of feature assembly and forward pass. Queue wait is the time the reading spends waiting for a free worker slot, and it is not fixed by anything - it is an emergent property of how full the tier is. At low utilisation a reading almost always finds a free slot and queue wait is close to zero, so latency is essentially service time. At high utilisation it almost never does. ## Kingman's formula and the shape of the blow-up **Kingman's formula** approximates the mean wait in a single-server queue as ``` E[Wq] ~= ( u / (1 - u) ) x ( (ca^2 + cs^2) / 2 ) x S ``` where `u` is utilisation, `S` is mean service time, and `ca^2` and `cs^2` are the squared coefficients of variation of inter-arrival times and service times. Three things follow, and all three matter in a design round: 1. **The utilisation term dominates near saturation.** `u / (1 - u)` is 1 at 50%, 1.5 at 60%, 4 at 80%, 9 at 90% and 99 at 99%. Going from a 60% target to a saturated fleet is not a modest degradation; it is a different regime. 2. **Variability is a multiplier, not a detail.** A tier whose per-reading work varies - larger payloads, cold feature lookups, an occasional retry - has a high `cs^2`, and it pays the utilisation penalty several times over. 3. **Service time sets the units.** Everything is measured in service times, which is why a 40 ms tier tolerates a wait factor that a 5 ms tier would not. An instance with several worker slots does better than this single-server formula predicts - pooling slots helps - but the `u / (1 - u)` shape survives, and that shape is the answer to the question. ## The queue that matters is the instance's It is tempting to reason about the fleet as one pool of 800 worker slots, and a single 800-slot queue really can run at high utilisation with little waiting. But that is not the system. The tier is a hundred separate queues of eight slots each, and a reading that lands on a busy instance waits there regardless of idle slots elsewhere. The concurrency that governs the queueing behaviour is the **per-instance** slot count, which is small, and small queues degrade at much lower utilisation than large ones. | | throughput sizing | latency sizing | |---|---|---| | the test | offered rate vs ceiling | measured p99 vs deadline | | what it produces | the saturation count | the utilisation target | | what it ignores | queue wait, variability, the tail | nothing, if measured under load | | how it fails | silently: nothing is dropped | visibly: the deadline is missed | ## Turning the deadline into a utilisation target 1. Load one instance, or a small representative slice of the fleet, with production-shaped readings at a series of offered rates. 2. At each rate, record the utilisation and the **p99** of queue wait plus service time - not the mean, which stays flat far longer than the tail does. 3. Find the highest utilisation at which measured p99 still sits inside the 250 ms deadline with margin for the inputs you do not control. 4. Use that utilisation as the divisor in the sizing arithmetic, and record it with the model version it was measured against. 5. Re-measure when the model artifact changes, because it changes both `S` and `cs^2`. The discipline is that the target is measured, not chosen. A round number carried over from another service is a guess about a latency curve nobody has drawn. ## The mean is not the tail One last trap: the mean latency of a saturated tier can sit comfortably inside a deadline that its p99 misses by a wide margin. Wait times are long-tailed, and the deadline is defined at a percentile. Report the percentile the operator's window is specified at, and size against that one - looking at the mean is how a tier passes review and fails in production.

  • After re-sizing, measured p99 at 60% utilisation is 210 ms against the 250 ms deadline. What does moving the target to 75% buy and cost?
    It buys machines: the fleet goes from 100 instances to `60 / 0.75 = 80`, a fifth fewer. It costs tail latency: the wait factor moves from 1.5 to 3, so mean queue wait roughly doubles and the p99 rises with it, very likely through 250 ms. The honest move is to re-measure at 75%, not to extrapolate the mean and hope.
  • Two scoring tiers both run at 60% utilisation but one has a much worse p99. What differs?
    Variability, and the size of the queue it hits. Kingman multiplies the utilisation factor by the average of the squared coefficients of variation, so a tier whose per-reading work varies widely - mixed payload sizes, occasional cold feature lookups - waits far longer at the same load. Fewer worker slots per instance makes it worse again, because a small queue pools less.
  • The scoring tier's mean latency is 55 ms and its deadline is 250 ms. Why is that not reassuring?
    Because the deadline is defined at a percentile and the mean is not close to it. Queue wait is long-tailed, so a tier whose mean sits at little more than its 40 ms service time can still have a p99 several hundred milliseconds out. Always read the percentile the deadline is written at.

A checkout lane that is busy 60% of the time has a short queue behind it; the same lane busy 95% of the time serves the same customers per hour, but almost all of a customer's time is now spent waiting to reach it rather than being served.

saying these in an interview costs you the question

  • Says the fleet is fine because no readings were dropped.
  • Assumes latency stays flat until the ceiling is actually reached.
  • Reads the mean latency and assumes the p99 tracks it proportionally.
  • Calls queue wait negligible because the work itself is only 40 ms.
  • Believes more instances at the same utilisation will fix the tail.
  • Treats the whole fleet as one large queue rather than many small ones.