A bidder's 75 ms internal p99 budget is decode 5, feature fetch 25, scoring 25, post-processing 10 and 10 ms reserve, and the measured p99s are 4, 38, 22 and 9 - which stage is over its line?
answer
- compare each stage to its own line
- the total can hide one breach
- other stages were quietly under budget
- the reserve absorbed the difference
- feature fetch is 13 ms over
basics
~20 sFeature fetch, at 38 ms against a 25 ms line. The measured total still fits only because the other three stages came in under theirs, and the 10 ms reserve is down to 2 ms.
solid answer
~40 sCompare each stage to its own line, not the total. Feature fetch is 38 ms against a 25 ms commitment, so it is 13 ms over. Decode, scoring and post-processing are each a millisecond or two under, and the measured p99s add to 73 ms inside a 75 ms budget, so the end-to-end number looks green. That reading is the trap: the lines are per-stage commitments, not a shared pool. The overrun has been paid for out of the reserve, which is now 2 ms instead of 10, so the next request on which any other stage runs normally-slow has nothing left to absorb it. Report the fetch stage as breached and treat the healthy total as a coincidence, not a pass.
code
json · 15 lines{
"exchange_deadline_ms": 100,
"network_round_trip_ms": 25,
"internal_budget_ms": 75,
"percentile": "p99",
"stages": [
{ "stage": "decode", "line_ms": 5, "measured_ms": 4 },
{ "stage": "feature_fetch", "line_ms": 25, "measured_ms": 38 },
{ "stage": "scoring", "line_ms": 25, "measured_ms": 22 },
{ "stage": "post_processing", "line_ms": 10, "measured_ms": 9 }
],
"reserve_ms": 10,
"reserve_remaining_ms": 2,
"breached": ["feature_fetch"]
}go deeper
Know that a budget is read stage by stage against each stage's own line, and that a healthy total can hide a breach. Name the over-budget stage and the size of the overrun in milliseconds.
Explain why the total fits anyway - other stages under-ran and the reserve absorbed the rest - and why per-stage p99s are commitments rather than a pool that any stage may draw down.
Say what you would do next: confirm the measurement point, decide whether the overrun is structural, and re-cut the budget explicitly with the other owners rather than quietly living off the reserve.
Frame the reserve as insurance the organisation is spending without deciding to, and set the rule for when a line may be re-cut, who agrees to it, and what the budget review looks at each quarter.
## The promise a budget line makes An exchange gives every bidder the same contract: a response that arrives after the deadline is discarded, so a late bid and no bid are the same event. That single external number - here 100 ms measured at the exchange, wire to wire - is useless as an engineering instruction, because nobody can work on "100 milliseconds". A **latency budget** breaks it into line items, one per stage on the request path, and each line becomes a commitment the stage's owner defends at a stated percentile. In this bidder the public network round trip takes 25 ms, leaving **75 ms inside the service**, and that 75 ms is cut into decode 5, feature fetch 25, scoring 25, post-processing 10 and a 10 ms reserve. ## Reading the sheet | stage | line (ms) | measured p99 (ms) | verdict | |---|---|---|---| | decode | 5 | 4 | inside, 1 ms spare | | feature fetch | 25 | 38 | **13 ms over** | | scoring | 25 | 22 | inside, 3 ms spare | | post-processing | 10 | 9 | inside, 1 ms spare | | reserve | 10 | 2 left | nearly consumed | | total | 75 | 73 | fits today | One stage is breached. The other three are inside their lines by a millisecond or three, and those spare milliseconds plus 8 ms of reserve are exactly what is paying for the fetch overrun. ## Why the total is the wrong first check - The lines are **commitments, not a pool**. A stage that overruns is in breach even when the request still fits, because the fit depended on other teams being lucky. - The three under-running stages have no obligation to stay under. The moment any of them returns to its line the request path is at 77 ms against a 75 ms budget. - The **reserve is insurance, not capital**. Reading 2 ms left as "still positive" converts an emergency buffer into baseline spend without anyone deciding to. - A stage p99 measured at the caller includes its own queueing; a stage that is over its line for structural reasons will keep being over on the next traffic peak, when every other stage is also worse. ## Percentile arithmetic, stated precisely Summing per-stage p99s does not produce the end-to-end p99. Under independence the end-to-end p99 sits **well below** the sum, because all four stages hitting their own tails on the same request is rare. That is why designing against the sum is a deliberately conservative choice: it buys the specific guarantee that **any one stage can hit its tail and the request still fits**. The conservatism evaporates when the stages share a cause - one saturated host, one congested network path, one process-wide pause - because then slow moments coincide and the sum stops being a safe ceiling. So treat the sum as a design rule and the measured end-to-end p99 as the truth. ## What to do when a stage is over its line 1. **Check the measurement point.** A stage p99 taken inside the stage excludes the wait to get into it; the number that matters is the one the caller sees. 2. **Decide whether the overrun is structural or transient.** A wider fan-out, a bigger feature vector or a new dependency is structural; one bad host for an hour is not. 3. **Fix the stage or re-cut the budget explicitly.** Re-cutting means taking milliseconds from a named other line, with that line's owner agreeing - never from the reserve by default. 4. **Re-state the reserve.** If the reserve is now the fetch stage's overflow, the budget has silently lost its insurance and the timeout rate will rise on the next peak. ## What the reserve is actually for The reserve covers what the per-stage lines do not model: the coincidence of two stages being slow on the same request, the cost of abandoning a stage and emitting a degraded response instead, and the small fixed costs - serialisation, the outbound hop - that never get their own row. Sizing it is a judgment call, but spending it is not; once it is routinely consumed by one stage's overrun, the budget has stopped describing the system.
- Why does the end-to-end p99 usually come in below the sum of the per-stage p99s?Because under independence it is rare for several stages to hit their own tails on the same request, so the sum is a conservative ceiling rather than a prediction. Budgeting against it buys one specific property: any single stage can reach its tail and the request still fits. When stages share a cause - a saturated host, a common network path - the slow moments coincide and the sum stops being a safe ceiling.
- The fetch overran only because one shard was slow. Does that change the verdict on the line?No. The line is defended at the stage's p99 as its caller experiences it, whatever the internal cause. Attribution changes the remedy, not the reading: the stage is in breach, and the owner now has a specific fix to make rather than a vague one.
- If the reserve is sitting at 2 ms, what is the honest way to report the budget?As one breached line and no remaining insurance, not as a 73-of-75 pass. The reserve exists for coincidence and for the cost of emitting a degraded response; once a single stage's overrun consumes it, the next ordinary fluctuation in any other stage pushes the request past the deadline.
saying these in an interview costs you the question
- Calls the budget healthy because 73 ms is under 75 ms
- Treats per-stage lines as one shared pool any stage may draw from
- Adds the stage p99s and calls the sum the end-to-end p99
- Compares stage means against lines that were stated at p99
- Assumes the external deadline can simply be renegotiated upward
- Re-baselines the reserve downward instead of reporting the breach