skip to content

Why does a real-time bidder state its response-time budget at p99 rather than as a mean response time?

level: middleimportance: must knowfreq 68%

answer

  1. value falls off a cliff
  2. a late bid is worth zero
  3. the mean cannot see a cliff
  4. the percentile is a forfeit budget
  5. ten slots turn 1% into 10%

basics

~20 s

Because the exchange discards a late bid entirely, so value falls off a cliff at the deadline and a mean cannot see a cliff. A percentile names the share of impressions the bidder is willing to forfeit.

solid answer

~50 s

The deadline is a cliff, not a slope: a bid at 99 ms is worth its full value and a bid at 101 ms is worth zero, because the exchange has already closed the auction. A mean summarises the middle of the distribution and says nothing about how often you fall off the cliff - two systems with the same mean can lose 0.1% and 5% of their impressions. Stating the budget at p99 makes the loss explicit: it commits to at most 1% of impressions missing the deadline, and at a few hundred thousand impressions per second that 1% is thousands of forfeited auctions every second. The percentile is therefore a business choice about acceptable forfeit, and it is also the shape of the exchange's own rule, since bidders whose timeout rate stays high get called less often or dropped from the pool.

go deeper

for a junior

Remember that a response past the deadline is discarded, so it is worth nothing rather than slightly less. That cliff is why the budget is written as a percentile and not as an average.

for a middle

Explain what a stated percentile commits you to as a share of forfeited impressions, and give a case where a change improves the mean while making the tail worse.

for a senior

Show how you would choose the percentile: price a forfeited impression, price the headroom that buys another nine, and pair the percentile with a timeout rate measured where the contract is measured.

for a principal

Own the forfeit as a business number rather than an engineering one, and be able to defend spending headroom on the tail instead of on model quality when the two compete for the same milliseconds.

## The deadline is a cliff, not a slope Most latency conversations assume a slope: slower is worse, a bit slower is a bit worse. A bidding path does not work that way. The exchange collects bids until its deadline, runs the auction and serves the winner; a response that arrives afterwards is thrown away. The value of a bid as a function of its latency is therefore flat right up to the deadline and **zero immediately after it**. Nothing about a mean can describe a cliff. A mean of 35 ms is compatible with a tidy distribution that never exceeds 60 ms and with a bimodal one that is fast most of the time and 400 ms on one impression in twenty; the first forfeits nothing and the second forfeits 5% of its inventory. ## What a percentile actually commits you to Designing so that the p99 equals the deadline is a statement that **1% of impressions will be forfeited**. Choosing the percentile is choosing the size of that loss, and at a high request rate it is a large absolute number. | budget stated at | share past the deadline | forfeited impressions per second at 300k rps | |---|---|---| | p50 | 50% | 150,000 | | p95 | 5% | 15,000 | | p99 | 1% | 3,000 | | p99.9 | 0.1% | 300 | The table also shows why the mean is not merely imprecise here but structurally wrong: it does not appear in this table at all, because it answers a question nobody is asking. ## Why the mean can move in the opposite direction A change that makes the common path cheaper and the rare path more expensive improves the mean and degrades the tail at the same time. A lookup that returns immediately when the entry is present and does substantially more work when it is absent has exactly this shape: most requests get faster, the unlucky minority gets slower, the mean falls and the timeout rate rises. If the team watches the mean it will ship that change and congratulate itself. - **The mean hides multimodality**, and serving paths are multimodal by construction: hit and miss, warm and cold, fast shard and slow shard. - **The mean is not robust**: a handful of very slow requests can drag it up, so a mean that looks bad may be a tail problem in disguise. - **Percentiles do not average across machines.** The p99 of a fleet is not the mean of the per-host p99s; aggregate the raw distribution, not the summaries. ## The tail is the common case at scale A publisher page with ten ad slots issues ten independent bid requests. If each has a 1% chance of missing the deadline, the chance that **at least one slot on the page is unfilled by this bidder** is 1 - 0.99^10, about 9.6%. The per-request tail is rare; the per-page tail is not. Any unit of user-visible work that fans out multiplies a small per-request percentage into a large per-unit one, which is the argument for budgeting at a further-out percentile than instinct suggests. ## Choosing the percentile, and what it costs Each additional nine is bought with headroom: lower utilisation, a narrower fan-out, a cheaper scorer, more replicas - resources that could otherwise have bought accuracy. The honest procedure is: 1. Price a forfeited impression, which for a bidder is roughly the expected value of winning it. 2. Price the headroom that moves the budget out one more nine. 3. Pick the percentile where the second exceeds the first, and write it into the budget as the stated percentile for every stage line. ## What p99 still does not tell you A percentile says how often you are late, never how late. A p99 of 95 ms against a 100 ms deadline is comfortable if the 99th-to-100th percentile band runs to 110 ms and alarming if it runs to 2 seconds, because the second case means a whole class of requests is failing for a different reason. So the p99 is paired in practice with the **timeout rate** - the fraction actually past the deadline - and with one further-out percentile to expose the shape of the very end of the distribution. Both numbers have to be measured where the contract is measured: at the exchange, wire to wire, including the network legs and any queueing at the bidder's own edge. An in-process p99 that excludes those is a number about the code, not about the promise.

  • When would you budget at p99.9 rather than p99?
    When a forfeited impression is expensive relative to the headroom that buys the extra nine, or when each user-visible unit fans out into many requests, so a 1% per-request miss becomes a near-certain per-page miss. The cost is real: the extra nine is paid for with lower utilisation, a narrower fan-out or a cheaper scorer, all of which could have bought accuracy instead.
  • Mean latency fell after a change, but the timeout rate rose. What kind of change does that?
    One that makes the common path cheaper and the rare path more expensive - a lookup that short-circuits when the entry is present and does extra work when it is absent, for example. Most requests get faster, so the mean improves, while the unlucky minority moves past the deadline. Watching the mean makes this look like a win.
  • Where should the p99 be measured, inside the bidder or at the exchange?
    At the exchange, because the contract is wire to wire: it includes both network legs, connection setup effects and any queueing at the bidder's own edge. A service that measures only its in-process p99 can report a comfortable number while the exchange is timing it out and reducing the share of requests it sends.

Arriving one minute late for a train and thirty minutes late are the same outcome: you are not on the train. The average arrival time tells you nothing about how often you miss it.

saying these in an interview costs you the question

  • Treats a late response as merely slow rather than worthless
  • Quotes a mean or median as the latency SLA
  • Assumes p99 is roughly the mean plus a small margin
  • Thinks adding replicas removes a tail caused by a slow dependency
  • Averages per-host p99s to get a fleet p99
  • Reports an in-process p99 for a wire-to-wire contract