skip to content

In a warehouse pick-path service, what does a 250 ms p99 budget cover beyond the model's inference time?

level: juniorimportance: should knowfreq 52%

answer

  1. measured where the user waits
  2. the clock starts at the handheld
  3. inference is one line item
  4. network, lookup, inference, post-processing, serialisation
  5. p99 at the caller, not the slice

basics

~20 s

A p99 budget is measured end to end at the caller, so it also covers the handheld's network round trip, request handling, feature lookup, post-processing and serialisation. Inference is one line item inside that total, not the total itself.

solid answer

~40 s

The envelope promises a wait to whoever is waiting, and here that is a picker holding a handheld terminal. So the clock starts when the request leaves the device and stops when the ordered pick list is rendered, and everything in between is inside the budget: the wireless round trip across the warehouse floor, request parsing, the feature lookup from the online store, the model's forward pass, the post-processing that orders the remaining picks, and serialisation. Inference is usually a minority of the budget, not the whole of it. The percentile matters as much as the number: `p99` is the slow request a picker meets several times a shift, while a mean is met by requests nobody notices. A number quoted for the inference stage alone is a measurement, not a promise.

go deeper

for a junior

Recall that a latency budget is measured end to end where the caller waits, and that the model's forward pass is only one stage inside it alongside network, lookup and post-processing.

for a middle

Explain why a percentile rather than a mean, and be able to name the stages of a request path and roughly what each costs, including the ones that run no model at all.

for a senior

Show you insist on the measurement point being written down, and that you can read a tail-only breach as a minority path rather than as a general slowdown.

for a principal

Frame the clause as an accountability boundary: an unsplit budget with no stated measurement point is a number every team can meet separately while the person holding the device still waits.

## What the number is actually measured against In a warehouse pick-path service, a picker's handheld terminal asks for the next location and the service returns the remaining picks of the wave in the order it wants them walked. Before anyone chooses a model, the design round writes an envelope: how fast a prediction must return, how stale its inputs may be, and what one prediction may cost. The latency clause is the first of the three, and `250 ms p99` is only a requirement once it says **where the clock starts and stops**. The default that matters is the **caller**. The handheld sends a request, the reply arrives, the screen updates. Every millisecond between those two events belongs to the budget — including the ones no one in the room owns, such as the wireless hop across a building full of steel racking. A budget measured at the service's own ingress quietly excludes exactly the part of the path a warehouse is worst at. ## The line items inside one request | stage | share of a 250 ms budget | who can move it | |---|---|---| | wireless round trip, handheld to service | ~60 ms | the network and site teams | | request handling, auth, deserialisation | ~10 ms | the service owner | | feature lookup from the online store | ~30 ms | the feature pipeline owner | | inference | ~60 ms | whoever authors the model | | post-processing and serialisation | ~40 ms | the service owner | | unallocated headroom | ~50 ms | nobody, deliberately | Two things follow from reading that table rather than a single number: - **Inference is one line among six.** Even a generous slice leaves more than half the budget to stages that have nothing to do with the model. - **Post-processing is not free.** Ordering the remaining picks, applying aisle constraints and serialising the list is real work on the request path, and it is the stage most often forgotten because no model runs in it. - **The stages have different owners.** A budget that is not split is a budget nobody can be held to. ## Why the tail rather than the mean A percentile is a statement about a population of requests. `p99` says one request in a hundred may exceed the number; the mean says nothing about any individual request at all. - A picker who makes roughly 300 requests over a shift meets the 1% tail about **three times** — three visible stalls, not a rounding error. - Request paths are usually bimodal: a warm cache hit and a cold lookup are different code paths with different costs, and a mean lands in the empty space between them, describing neither. - The tail is where retries, a cold replica, a pause on a busy host and a full queue all land. Those are the events an operator has to design for, and the mean is blind to every one of them. ## What the latency clause has to say out loud 1. **The measurement point** — at the handheld, not at the service's ingress; if both are reported, say which one is the promise. 2. **The percentile** — `p99` here, with the p50 reported alongside as a diagnostic rather than as the commitment. 3. **The load and the window it holds at** — a percentile measured over a quiet mid-shift hour is a different and much easier promise than the same percentile at wave release. 4. **What a breach costs** — a picker standing still is the unit of harm, and naming it is what stops the number being negotiated away later. ## The failure this clause prevents The common one is the inference-only number. A team measures its model server at 40 ms p99, reports that the envelope is met, and the picker still waits most of a second. Nothing in that report was false; it simply measured one stage of six. When the feature lookup turns out to cross a slow link, or the post-processing sorts a list far longer than anyone sized for, the budget was never being watched at the place it was promised. The opposite failure is over-reading the clause. The envelope states the **total and the measurement point**; it does not decide whether the answer is computed on the request or looked up from something precomputed earlier, and it does not decide how many machines serve it. Those are the next conversations, and they are constrained by this number rather than contained in it.

  • Why is the p99 measured at the handheld rather than at the service's own ingress?
    A server-side percentile excludes the wireless round trip and any queueing before the request is accepted, which on a warehouse floor is exactly where time is lost. The picker experiences the caller-side number, so that is the one the envelope promises. Measuring at ingress is still useful, but as a component of the budget rather than as the commitment.
  • If the p99 breaches the budget while the p50 stays comfortable, what does the split tell you to look at?
    A tail-only breach points at a minority path rather than at the median cost of anything: a cache miss, a cold replica, a retry, a queue that only forms under burst. Because the envelope allocated a slice per stage, each stage can be measured at its own p99, and the one whose tail moved is the one that owns the breach.

It is the difference between a restaurant promising the grill takes four minutes and promising the food reaches your table in fifteen. Only one of those is a promise to the person sitting down.

saying these in an interview costs you the question

  • Quoting the model's inference time as the service's latency.
  • Writing an average latency target instead of a percentile.
  • Assuming the network inside a warehouse is free.
  • Treating post-processing as too small to budget.
  • Measuring at the server and ignoring the caller's round trip.