A bidder's feature fetch now finishes 12 ms under its line - how do you decide whether a heavier scorer may spend those milliseconds?
answer
- milliseconds have a price
- a late bid scores zero
- marginal lift per millisecond is concave
- slack is insurance, not capital
- compare expected value, not accuracy
basics
~20 sPrice the milliseconds against what they buy. The accuracy gain must beat the impressions forfeited to a higher timeout rate, because a bid past the deadline is worth zero, and headroom is only spendable if it holds at peak.
solid answer
~40 sTreat it as an expected-value trade, not a free upgrade. Suppose the heavier scorer lifts the value of a won impression by 1.5% but, by consuming the slack that was absorbing the fetch tail, moves the timeout rate from 0.5% to 2%. Expected value per impression goes from 0.995 to 0.98 x 1.015 = 0.9947 - slightly worse, so the upgrade loses. A cheaper variant that takes only 3 ms more, lifts value 1.0% and holds the timeout rate near 0.7% gives 0.993 x 1.010 = 1.0029, which wins. Two checks decide whether the headroom is even real: is the fetch improvement structural rather than an off-peak trough, and was the scorer's cost measured at p99 under production concurrency rather than alone on an idle host?
go deeper
Remember that spare milliseconds in one stage are not automatically free for another, because a response that misses the deadline is discarded and earns nothing at all.
Explain the diminishing return of extra scoring time, and compute expected value as one minus the timeout rate times the value per served impression rather than comparing accuracy.
Check whether the headroom is structural and read at the budget's percentile, measure the scorer under production concurrency, and be ready to choose the cheaper variant over the heavier one.
Set the rule that model investments are priced in expected value per impression, and be willing to spend recovered milliseconds on reserve when the tail is the binding constraint.
## A millisecond has a price and a value Once a budget exists, every proposal to make the model better is a proposal to buy milliseconds, and milliseconds on this path have a market price. The value side is the metric lift the extra scoring time buys. The price side is what those milliseconds were previously doing - almost always absorbing some other stage's tail. Spending them is therefore not free even when the sheet shows them as unused. ## The value side is concave Extra scoring time buys accuracy at a sharply diminishing rate: more features, a wider model or an extra cascade stage each help, and each helps less than the last. | extra scoring time | lift in value per won impression | |---|---| | +3 ms | +1.0% | | +6 ms | +1.3% | | +12 ms | +1.5% | | +20 ms | +1.7% | The shape of that curve is the whole argument. If the first few milliseconds buy most of the available lift, the cheap variant is usually the right buy and the expensive one is usually a loss once its tail cost is counted. ## The price side: what the milliseconds were doing - **Slack is insurance.** The reserve exists so that a stage overrun still produces an on-time response; consuming it converts a rare degraded response into a frequent forfeited one. - **A discarded bid is worth zero**, so the timeout rate multiplies the value of every impression you do serve. - **Another stage's line is someone's commitment.** Taking milliseconds from it is a negotiation, not an accounting entry. - **The scorer's cost is not what a benchmark says.** Measured alone on a quiet host, a scorer looks cheaper than it is; under production concurrency it queues behind its neighbours, so use its p99 in situ. ## Working the trade | option | scoring p99 | timeout rate | value per won impression | expected value per impression | |---|---|---|---|---| | current scorer | 25 ms | 0.5% | 1.000 | 0.995 | | heavier scorer | 37 ms | 2.0% | 1.015 | 0.9947 | | cheaper variant | 28 ms | 0.7% | 1.010 | 1.0029 | Read the last column, which is simply (1 - timeout rate) x value. The heavier scorer gains 1.5% on the impressions it still serves and loses 1.5 points of impressions outright, which nets to a small **loss**. The cheaper variant gains less per impression and costs almost nothing in forfeits, so it wins. This is the accuracy-per-millisecond argument in its usable form: compare expected value, not accuracy. ## Is the headroom even real? 1. **Structural or transient?** A fetch that got faster because the fan-out was permanently narrowed has genuinely released milliseconds. A fetch that looks fast at 04:00 has not; at peak it returns to its line and the scorer's new appetite becomes a timeout. 2. **Correlated or independent?** If the fetch is fast on exactly the requests where scoring is fast, the spare time is not available when it is needed. The transfer is only safe if the fetch is under its line on the same requests that the scorer's tail lands on. 3. **Measured where?** The headroom must be read at the same percentile as the budget. A stage that is 12 ms under at the median and on its line at p99 has released nothing. ## When the milliseconds simply do not exist The interesting case is a target the budget cannot fund: the lift needs 20 ms and there are 12. The move then is not to overrun but to obtain a **cheaper scorer that approximates the expensive one** - and the useful direction of that request is worth stating precisely. The budget names the requirement: a scorer whose p99 fits in the milliseconds available, losing no more than a stated fraction of the heavy model's lift. The millisecond gap sets the compression target and the accuracy floor; how that smaller scorer is produced is a modelling exercise, and the budget's job is to hand it a number rather than a preference. One variant deserves its own note. If the scorer no longer fits on one host and is split across two, the scoring line stops being a single item: it becomes compute plus an in-datacentre round trip, with a second tail of its own inside the same stage. That is a real option, but it is priced the same way - the lift the split unlocks has to beat the added p99 and the extra failure surface, judged on expected value per impression rather than on model quality alone.
- The lift you want needs 20 ms and the budget has 12. What is left to try?Ask for a cheaper scorer that approximates the heavy one, stated as a requirement: a p99 that fits the available milliseconds while giving up no more than a named fraction of the lift. The gap in milliseconds is what sets that target and that floor. If no such scorer exists, the question stops being a budget question and becomes a question about where the prediction is produced at all.
- The scorer no longer fits on one host and must be split across two. What does that do to the scoring line?It turns one line item into compute plus an in-datacentre round trip, and puts a second tail inside the same stage, so the line's p99 is now a maximum over two things rather than one measurement. It is worth doing only if the accuracy the split unlocks beats the added p99 and the extra failure surface, priced as expected value per impression.
- Would you ever hand the 12 ms back as reserve instead of spending it?Yes, when the timeout rate is already at or near what the exchange tolerates. The tail is a hard constraint - past it the impression is worth nothing - while accuracy is a soft one, so buying back headroom can be the higher-value purchase even though it shows up on no model metric.
saying these in an interview costs you the question
- Spends every free millisecond on model size by default
- Compares scorer options on offline accuracy alone
- Benchmarks a scorer on an idle host and calls that its production cost
- Treats an off-peak trough's headroom as permanent capacity
- Assumes accuracy gains scale linearly with scoring time
- Forgets that the milliseconds were insurance against another stage's tail