skip to content

In a warehouse pick-path service, how would you split a 250 ms p99 budget across its stages before choosing a model?

level: middleimportance: must knowfreq 64%

answer

  1. reserve before you allocate
  2. network, overhead and headroom come off first
  3. the remainder splits three ways
  4. the inference slice is a ceiling
  5. parts must reconcile with the total

basics

~20 s

Reserve first what you cannot spend — the caller's network round trip, fixed request handling and explicit headroom — then allocate the remainder across feature lookup, inference and post-processing. The inference slice becomes a ceiling any candidate model must fit.

solid answer

~40 s

I start from the parts that are not mine to optimise. Of 250 ms at the handheld I reserve roughly 60 ms for the wireless round trip, 10 ms for request handling and parsing, and 50 ms of unallocated headroom for retries, a cold replica or a pause on a busy host. That leaves 130 ms to spend, and I split it about 30 ms for the feature lookup, 60 ms for inference and 40 ms for ordering the remaining picks and serialising them. The point of doing this **before** a model exists is that the 60 ms becomes a stated ceiling on the request path — a constraint on how many features may be fetched and how heavy the scorer may be, measured on the hardware it will actually run on.

code

pseudocode · 15 lines
pseudocode
budget_p99_ms = 250

reserve wireless_round_trip    = 60   // not ours to optimise
reserve request_handling_parse = 10   // fixed per-request overhead
reserve headroom               = 50   // retry, cold replica, host pause

remaining = budget_p99_ms - (60 + 10 + 50)      // 130 ms to spend

allocate feature_lookup   = 30        // one batched read
allocate inference        = 60        // ceiling for any candidate
allocate post_processing  = 40        // order remaining picks, serialise

assert feature_lookup + inference + post_processing == remaining   // 130

model_ceiling_ms = inference          // stated before a model is chosen

go deeper

for a junior

Recall the order of operations: reserve what you cannot control, keep some headroom, then divide what is left among lookup, inference and post-processing so the parts reconcile with the total.

for a middle

Be able to produce a ledger with real numbers and defend each line, and explain why the inference allocation has to be measured on target hardware rather than on an idle developer machine.

for a senior

Show that you write the split before the model exists, and that you can explain why stage percentiles do not add and what headroom is actually absorbing.

for a principal

Treat the split as the contract that lets several teams work in parallel without renegotiating the promise, and decide deliberately how much of the total is left unallocated against operational risk.

## Start from what you cannot spend A latency budget is not divided fairly; it is divided in order of who can move it. In a pick-path service, whose caller is a handheld terminal on a warehouse floor, three costs are effectively fixed by the time the design round starts: - the **wireless round trip** in both directions, which belongs to the site and network teams; - **request handling** — accepting the connection, authenticating, parsing the payload; - **headroom**, which is not a stage at all but a deliberate refusal to allocate the last slice. Only what survives those reservations is available to the stages the design actually controls. ## A worked ledger for 250 ms | stage | allocation | notes | |---|---|---| | wireless round trip | 60 ms | reserved, not allocated; the service cannot shorten it | | request handling and parse | 10 ms | fixed overhead per request | | feature lookup | 30 ms | one batched read for the picker, the wave and the candidate bins | | inference | 60 ms | the ceiling handed to whoever chooses the model | | post-processing | 40 ms | ordering the remaining picks, applying aisle constraints, serialising | | headroom | 50 ms | retries, a cold replica, a pause on a busy host | | **total** | **250 ms** | reserved 120, remaining 130, allocated 130 | The arithmetic is the discipline: 60 + 10 + 50 reserved is 120, leaving 130, and 30 + 60 + 40 spends exactly that. A split whose parts do not reconcile with the stated total is not a budget, it is a wish. ## Why this happens before the model exists Writing the split afterwards inverts the design. Once a candidate has been trained, the conversation becomes "can we make the budget fit the model", and the honest answer is usually to widen the budget. Written first, the split is a **specification**: 1. The **inference allocation is a ceiling** measured on the hardware the service will run on, at the batch size the peak actually produces — not a laboratory figure on an idle machine. 2. The **feature-lookup allocation caps the fan-out**: 30 ms is one batched read, not four sequential ones, which constrains how many signals a candidate may consume before anyone has tried to consume them. 3. The **post-processing allocation caps the candidate list**: ordering the remaining picks costs time proportional to how many there are, so a wave size that blows 40 ms is a requirement problem, not a tuning problem. None of that names a model family, and it should not. The envelope states a ceiling; which kind of model fits underneath it is a modelling decision made later and elsewhere. ## Why the per-stage numbers do not simply add A tempting shortcut is to measure each stage's p99 and add them. That sum is a **conservative sketch rather than an identity**: - When stage tails are independent, they rarely coincide in the same request, so the sum overstates the true end-to-end p99. - When a single cause — a saturated host, a full queue, a garbage-collection pause, a failing replica — slows several stages at once, the tails coincide precisely in the moment that matters, and the sum stops being pessimistic. This is why the ledger carries explicit headroom instead of allocating all 250 ms. Headroom is the budget line that absorbs the correlated case, and a budget with none holds only in a quiet test. ## What the split deliberately does not decide The envelope constrains the design without making it: - It does **not** decide whether the answer is computed when the request arrives or read from something computed earlier; a precomputed answer simply makes the inference slice nearly free and moves the cost elsewhere. - It does **not** decide how many machines are needed, which is a sizing exercise done once the architecture and the peak load are both known. - It does **not** promise anything about the freshness of the values the lookup returns; that is a separate clause of the same envelope. What it does is hand every downstream decision a number it has to live inside — which is the entire purpose of writing an envelope in the first ten minutes rather than the last.

  • What does the 60 ms inference allocation constrain before a single candidate has been trained?
    It becomes a stated ceiling on the request path, measured on the hardware the service will run and at the batch size the peak produces. That caps how many signals may be fetched and scored, how deep an ensemble may be, and how large a scorer may be loaded. Which kind of model then fits underneath the ceiling is decided later and by someone else; the envelope only publishes the number.
  • Why reserve headroom explicitly instead of allocating all 250 ms to stages?
    Because stage percentiles do not add to the end-to-end percentile. Independent tails rarely coincide, so the sum is usually pessimistic — but a shared cause such as a full queue or a cold replica makes them coincide exactly under load. Headroom is the line item that absorbs the correlated case, and a budget without it holds only when nothing is going wrong.
  • The wave size doubles and post-processing now exceeds its 40 ms. Is that a tuning problem?
    No — it is a requirements problem surfacing in the right place. Ordering the remaining picks costs time proportional to how many remain, so the envelope has to say what wave size it was written for. Either the stated peak wave size changes and the split is rewritten, or the post-processing stage returns a bounded prefix of the route rather than the whole remainder.

saying these in an interview costs you the question

  • Splitting the budget evenly across stages because it looks tidy.
  • Giving every millisecond to stages and reserving no headroom.
  • Forgetting the caller's network round trip entirely.
  • Choosing the model first and discovering the budget afterwards.
  • Assuming post-processing is free because no model runs in it.
  • Measuring the inference ceiling on idle hardware at batch size one.