Why must a bidder's scoring timeout fire well before the exchange's response deadline rather than at it?
answer
- the deadline is not the abort point
- reserve the time to answer
- subtract the work after scoring
- compute remaining, do not hardcode
- abort at 55, not at 75
basics
~20 sBecause everything after scoring - post-processing, building the response and the hop back - still has to fit inside the deadline. Aborting at the deadline leaves no time to answer at all, turning a slow bid into no bid.
solid answer
~50 sThe deadline is not the abort point. Work the abort point backwards: the deadline minus the post-scoring work, the serialisation and outbound hop, and a margin. With 75 ms inside the bidder and 20 ms reserved for everything after the scorer, the scorer must be abandoned by about 55 ms elapsed. Set the timeout at the deadline and the request runs until there is no time left to emit anything, so a bid that was merely slow becomes silence - which scores zero. The stronger form is a remaining-time checkpoint rather than a fixed constant: recompute `remaining = deadline - elapsed - reserve` before the scorer starts, and if the scorer's p99 no longer fits, go straight to the cheaper path. That way an overrun in feature fetch shortens the scorer's allowance instead of blowing the response.
code
pseudocode · 18 linesINTERNAL_BUDGET_MS = 75 // what is left after the network legs
RESERVE_MS = 20 // post-processing 10 + response build and hop 10
SCORER_P99_MS = 25 // measured, under production concurrency
on bid_request(req):
start = now()
features = fetch_features(req, timeout_ms = 25)
elapsed = now() - start // 42 after a 4 ms decode and a 38 ms fetch
remaining = INTERNAL_BUDGET_MS - elapsed - RESERVE_MS
if remaining < SCORER_P99_MS: // 13 < 25, so this branch fires
return respond(contextual_bid(req), degraded = true)
score = run_scorer(features, timeout_ms = remaining)
if score is TIMED_OUT:
return respond(contextual_bid(req), degraded = true)
return respond(price_from(score), degraded = false)go deeper
Remember that work still happens after the scorer finishes, so the scorer cannot be allowed to run until the deadline. Give up early enough to build and send a response.
Derive the abort point by subtracting the post-scoring work and the outbound hop from the deadline, and explain why a recomputed remaining time beats a constant taken from the budget sheet.
Show the failure a late abort produces: requests run the full budget, concurrency climbs, queueing worsens, and a fully loaded bidder returns nothing at all while the exchange reduces its traffic.
Set the rule that every request path must be able to produce some response inside the deadline, and make the degraded path's own cost a budgeted line rather than an assumption.
## Two clocks, not one There are two different times in play and confusing them is the whole mistake. The **end-to-end deadline** is the external promise: the moment after which the exchange has closed the auction and any response is discarded. A **per-stage timeout** is an internal decision about how long one stage may run before the service gives up on it. The second must be chosen so the first is still keepable *after* it fires. A stage timeout set equal to the deadline guarantees that when it fires there is no time left to do anything with the fact that it fired. ## Where the abort point sits Work backwards from the deadline through every piece of work that still has to happen once the scorer is abandoned. | item after scoring | ms | |---|---| | post-processing (pacing check, price shading, creative choice) | 10 | | response build, serialisation, outbound hop, margin | 10 | | **reserve total** | **20** | | internal budget | 75 | | **latest safe abort point** | **55 ms elapsed** | So the scorer gets whatever is left between the end of feature fetch and 55 ms elapsed - not 25 ms because the sheet says 25, and never "until the deadline". ## A fixed timeout against a remaining-time checkpoint | | fixed per-stage timeout | remaining-time checkpoint | |---|---|---| | input | a constant from the budget sheet | deadline, elapsed time, reserve | | behaviour after an earlier overrun | unchanged, so the overruns add up | shrinks the later stage automatically | | behaviour when an earlier stage is fast | leaves the spare time unused | hands the spare time to the scorer | | failure mode | the response misses the deadline | the response is degraded but on time | The checkpoint is the reason the bidder can absorb a bad fetch at all: after a 38 ms fetch and a 4 ms decode, 42 ms are gone, remaining is 75 - 42 - 20 = **13 ms**, and the full scorer's 25 ms no longer fits, so the cheaper path runs immediately rather than starting a scorer that cannot finish. ## Why a degraded answer beats no answer here - A discarded response wins nothing, so its expected value is exactly **zero**. - A bid formed without the personalised score still wins some auctions at some price, so its expected value is positive even if much lower. - Any positive number beats zero, which means the budget's job is not only to make the good path fast but to **guarantee that some path can always finish**. - The degraded path has its own cost, and that cost must live inside the reserve - a fallback that itself needs 15 ms does not fit in a 10 ms margin. ## Sizing the reserve Size it at the **p99 of the work after scoring**, not its mean. The reserve is consumed precisely on the requests where things are already going badly, so a reserve sized at the average of the post-scoring path fails in the case it exists for. The same logic applies to the margin for the outbound hop: it is the tail of that hop that decides whether the response lands, not its typical value. ## The failure this prevents When every stage timeout is set at the deadline, a slow dependency produces a characteristic pattern: 1. Requests run the full budget rather than aborting early, so in-flight concurrency rises. 2. Higher concurrency makes queueing delay worse, so more requests run the full budget. 3. Every one of them returns nothing, so the bidder is fully loaded and contributing no bids at all. 4. The exchange, seeing a bidder whose responses are consistently late, reduces or stops sending it requests - so the recovery is slower than the incident. Aborting early breaks the loop at step one: the work stops before it becomes useless, the concurrency stops climbing, and every request still produces a response the exchange can count. The budget is not only an allocation of milliseconds; it is the mechanism that decides when to stop spending them.
- Feature fetch overran by 13 ms. Should the scorer's allowance shrink, or should the request be abandoned?Shrink it. Recompute the remaining time and either run a path that fits or degrade immediately; abandoning the request outright is strictly worse, because a degraded response still has positive expected value while silence has none. Abandonment is only right when even the degraded path no longer fits inside what is left.
- Should the reserve be sized at the mean or the p99 of the work that follows scoring?At the p99. The reserve is drawn on exactly when the request is already in trouble, and the post-scoring path is not immune to the same pressure - the same busy host that slowed the scorer slows serialisation and the outbound hop. A reserve sized at the mean fails in the case it was created for.
- The fetch finishes early and the checkpoint sees more time than expected. What should the bidder do with it?Either run a fuller scorer variant that fits the remaining time, or bank the milliseconds as extra reserve. Which one is right is a value judgment, not a mechanical one: spending it raises the timeout rate on the requests where the fetch is not early, so the accuracy gained has to beat the impressions forfeited.
saying these in an interview costs you the question
- Sets the scoring timeout equal to the end-to-end deadline
- Forgets serialisation and the outbound hop when sizing the reserve
- Uses a fixed stage timeout that ignores time already spent
- Treats returning nothing and returning a degraded bid as equivalent
- Sizes the reserve at the mean of the post-scoring path
- Assumes a timed-out request costs the system nothing