skip to content

A pick-path service is quiet mid-shift but every handheld requests at wave release — what must its latency clause state?

level: seniorimportance: must knowfreq 55%

answer

  1. a percentile needs a population
  2. state the load beside the number
  3. peak at wave release, not shift average
  4. name the burst window too
  5. rate times service time gives concurrency

basics

~20 s

The arrival rate and the burst window the percentile must hold at, not just the percentile. A latency number without a load attached is met at the shift average and missed at the minute the business actually cares about.

solid answer

~40 s

A percentile is a statement about a population of requests, so it only becomes a requirement once the envelope says which population. Here the two populations are wildly different: a shift averages a few tens of requests per second, while wave release puts hundreds of handhelds into the path within a minute or two. The clause has to read something like "250 ms p99 at 600 requests per second sustained for four minutes after a wave releases, against a 40 per second shift average". Without the load and the window, every downstream team picks its own and reports success. The paired numbers also imply the concurrency the design must carry — by Little's Law, arrival rate times time spent in the service — while how many machines provide that concurrency is sized separately.

go deeper

for a junior

Recall that a latency promise needs a load beside it, and that traffic which averages out to a small number can still arrive all at once.

for a middle

Explain why a percentile aggregated over a whole shift can look healthy while a burst was badly served, and convert an arrival rate into requests in flight.

for a senior

Show you write the peak, the window and the measurement point into the clause, and that you can recognise a green aggregate dashboard hiding a bad four minutes.

for a principal

Decide deliberately whether one promise covers peak and trough or whether the envelope carries two, and own the cost consequence of pinning everything to the peak.

## A latency without a load is not a requirement `p99 < 250 ms` sounds like a commitment and is not one. A percentile is computed over a set of requests, and nothing in that phrase says which set. Measured across a whole shift — mostly idle aisles, a few tens of requests per second — it is comfortably met by a design that collapses in the four minutes everyone is actually waiting. The pick-path load is not smooth and the shape is not incidental: - **Waves release together.** A wave is a batch of orders assigned at once, so the handhelds all ask for a first location within the same minute. - **The burst coincides with the business value.** The minutes after release are exactly when routing decides how the next hour of walking goes. - **The trough is long and flat.** A daily or shift-level average, divided out, understates the peak by an order of magnitude and looks reassuring while doing it. ## What the clause has to name 1. **The percentile** — `p99`, with `p50` reported as a diagnostic rather than as the promise. 2. **The measurement point** — at the handheld, so the caller's round trip is inside the number. 3. **The arrival rate it holds at** — the peak, stated as requests per second, not the average. 4. **The window** — how long that rate is sustained, because a two-second spike and a four-minute plateau are different engineering problems. 5. **The behaviour outside the window** — whether the same percentile is promised mid-shift or only a looser one. A clause carrying all five is testable. One carrying only the first is an opinion that everyone can honestly claim to have met. ## From a rate to what the design must carry The pair of numbers does real work immediately. **Little's Law** states that the average number of items in a system equals the arrival rate multiplied by the average time each spends inside it. Take a peak of 600 requests per second and roughly 130 ms spent inside the service — the part of the budget left after the caller's network and the fixed overheads: > concurrency = 600 requests/second x 0.130 seconds = **about 78 requests in flight** That single figure is what the serving design has to accommodate: 78 simultaneous feature lookups, 78 scores being computed, 78 responses being ordered and serialised. It changes the conversation from "is the model fast enough" to "what does this path look like with dozens of requests inside it at once", which is where queueing appears and where the tail the p99 promises is actually won or lost. Note also which direction the dependency runs: if utilisation rises, time in the system rises, which raises concurrency further. That feedback is why a peak-load clause without headroom in the latency split is unsafe, and it is why the two clauses are written together. ## Where this clause deliberately stops The envelope states the load the promise holds at. It does **not** decide how many machines provide the capacity, what the autoscaling policy is, or what it costs to run at peak — those are sizing questions, answered once the architecture exists and answered by a different part of the design. Nor does it decide whether the burst is absorbed by scoring on the request or by having answers ready before the wave releases; that is a serving-topology choice, constrained by this number rather than made by it. ## The failure mode this prevents The recognisable one is a green dashboard during a bad shift. The service reports `p99 = 210 ms` over the day and every team believes the envelope holds. Aggregated over a shift, the four-minute burst is a rounding error in the sample: even if every request during wave release took 900 ms, those requests are far too few to move a daily 99th percentile. The pickers experience nothing but those minutes. The fix is not a better dashboard. It is a requirement that named the load and the window in the first place, so that the percentile is measured over the population that matters and a breach is visible as a breach.

  • Why is "p99 under 250 ms" on its own something no team can design against?
    Because a percentile is only defined over a population of requests, and the phrase never says which one. One team can honestly measure it across a whole shift and pass, while another measures it during wave release and fails, with neither having made an error. Stating the arrival rate, the window and the measurement point together is what turns the number into a testable commitment.
  • What does the pair of numbers imply about how many requests are in flight at once?
    Little's Law gives it directly: concurrency equals arrival rate multiplied by time in the system. At 600 requests per second with roughly 130 ms spent inside the service, about 78 requests are in flight simultaneously. That concurrency is what the serving path has to carry without queueing into the tail; how many machines supply it is a sizing exercise done separately.
  • Should the envelope promise the same percentile mid-shift as at wave release?
    It should say explicitly, and the honest answer is often no. A single promise pinned to the peak makes the quiet hours over-provisioned, while a single promise pinned to the average is broken exactly when it matters. Stating one percentile at the peak load and, if useful, a separate one off-peak keeps both the requirement and its cost visible.

Promising a five-minute wait at the checkout is easy to keep at ten in the morning and meaningless unless the promise names the Saturday afternoon queue.

saying these in an interview costs you the question

  • Stating a latency target with no arrival rate attached.
  • Using the shift average as the load the p99 holds at.
  • Assuming traffic arrives smoothly because the daily total is modest.
  • Treating a four-minute burst as noise rather than the requirement.
  • Confusing requests per second with requests in flight at once.
  • Reading a daily aggregate percentile as proof the peak was fine.