skip to content

A login submit gives risk scoring a 120 ms p99 budget; how do you divide it across feature fetch, model call and decision emit?

level: middleimportance: should knowfreq 62%

answer

  1. p99, not average
  2. parts must sum to the total
  3. one absolute deadline, passed down
  4. fan-out costs the slowest branch
  5. keep a reserve for the fallback

basics

~20 s

Give transport 20 ms and the scorer 100 ms: about 45 ms for a parallel feature fetch, 25 ms for the model call, 10 ms for deciding and handing off the emit, and a 20 ms reserve so the fallback still fits inside the same budget.

solid answer

~50 s

I would spend the 120 ms as roughly 20 ms transport in and out plus 100 ms inside the scorer, split 45 / 25 / 10 with a 20 ms reserve. Feature fetch is the largest single stage because it is a fan-out over the account, device and network keys, and its cost is the p99 of the **slowest** lookup, not the sum. The model call on a small tabular model is cheap; deciding and handing the emit to an asynchronous path is in-process. The reserve is the part people forget: if a stage blows its slice, the fallback path itself still has to run and answer inside 120 ms, so I set one **absolute deadline** at arrival and have each stage check `remaining - reserve` before it starts rather than giving each stage an independent timeout.

code

pseudocode · 15 lines
pseudocode
budget   = 100ms                      // scorer's share; 20ms of the 120ms is transport
deadline = arrivalAtScorer + budget
reserve  = 20ms                       // must survive to run the fallback and answer

features = fetchInParallel([accountKey, deviceKey, networkKey], timeout = 45ms)
if features.timedOut or now() > deadline - reserve:
    return fallback(reason = "feature-fetch")

score = scoreModel(features, timeout = 25ms)
if score.timedOut or now() > deadline - reserve:
    return fallback(reason = "model")

action = applyDecision(score)         // <= 10ms, in process
emitAsync(action, score)              // after the answer, never inside the budget
return action

go deeper

for a junior

Know that the risk score is produced inside the login request and therefore has a fixed millisecond allowance, and that the allowance is a tail number rather than an average.

for a middle

Be able to lay out the stages, give each a number, show the numbers summing to the total, and explain why the feature fetch is parallel.

for a senior

Demonstrate the deadline discipline: one absolute deadline, a reserve for the degraded path, and the guarantee that an answer always arrives on time even when it is a worse answer.

for a principal

Treat the budget itself as negotiable with the product owner, and be explicit about what the organisation buys with each extra 20 ms of login latency.

## What the 120 ms actually is It is a **p99 for the whole risk decision**, measured from the login service's call to its answer, and it is a slice carved out of a larger budget for the credential submit as a whole. Two things follow immediately: - The number is a **tail**, not a mean. A design that averages 40 ms and has a 300 ms p99 has failed this budget, because the p99 is the experience of a meaningful number of real logins every minute. - The scorer does not own all 120 ms. Serialisation, the network hop in each direction and queueing at the scorer's own listener take real time, and the budget has to name them or they get spent silently. ## A worked split | stage | budget | why it gets that much | |---|---|---| | transport in and out | 20 ms | two hops plus serialisation; not the scorer's to spend | | feature fetch | 45 ms | parallel fan-out over the account, device and network keys | | model call | 25 ms | a small tabular model; the forward pass is not the problem | | decide and hand off emit | 10 ms | in-process; the emit itself leaves the critical path | | reserve | 20 ms | so the fallback still answers inside the same 120 ms | The scorer's own share is **45 + 25 + 10 + 20 = 100 ms**, and 100 + 20 transport = the stated **120 ms**. If your parts do not sum to your total, the budget is decoration. ## Deadlines, not timeouts The common mistake is to give each stage an independent timeout. Three stages at 45, 25 and 10 ms mean a worst case of 80 ms of work *after* whatever the earlier stages already overspent, and the total drifts past the budget without any single stage looking wrong. Instead, compute one **absolute deadline** when the request arrives and pass it down: - each stage starts only if `now() < deadline - reserve`; - each stage's own timeout is `min(its slice, deadline - reserve - now())`; - the first stage that cannot start within that bound stops and returns the degraded decision. This gives the property the budget exists for: **the answer is always inside 120 ms**, and the only thing that varies is how good the answer is. ## Why feature fetch gets the biggest slice The scorer needs counters and attributes for several entity keys at once. Two facts shape the slice: - **Fan-out latency is the slowest branch, not the sum** — three parallel lookups at 8 ms each cost about 8 ms, not 24 ms, so parallelism is not optional here. - **Every extra key widens the tail.** With three independent lookups each at a 1-in-100 chance of being slow, the chance that *at least one* is slow is close to 3-in-100. The tail of the fan-out is worse than the tail of any single store call, which is why the slice is generous and why some systems hedge or simply drop a late feature and score without it. ## Where inline budgets are lost 1. **Synchronous logging or emitting.** Writing the decision record before answering puts a storage write inside the budget. Emit asynchronously after the response. 2. **A retry hidden inside a stage.** One retry inside the feature fetch quietly doubles that stage and eats the next one's slice, unless the retry is bounded by the same deadline. 3. **Measuring the mean.** Dashboards showing average scoring latency hide exactly the population the p99 budget is about. 4. **No reserve.** The system times out at the deadline with nothing left to run the fallback, so the login path gets no answer at all rather than a degraded one. ## What you trade when the budget binds If the split will not fit, the levers are all about doing **less** inline, in roughly this order of preference: - fetch fewer entity keys, or fetch the expensive one only when the cheap features are already suspicious; - move a heavy aggregate off the request by having the stream keep it warm instead of computing it inline; - use a smaller model, since on this path the forward pass is rarely the dominant cost anyway; - widen the budget, which is a product conversation about login latency and not an engineering decision the scorer's owner can take alone. What you do **not** do is let the scorer run long. An inline risk decision that misses its deadline is not a slow decision; it is a decision the login path will make without you.

  • Three feature lookups run in parallel at 8 ms each. Why is the stage's p99 still worse than any single store's p99?
    The stage waits for the slowest branch, so it is slow whenever *any* one of the three is slow. With three roughly independent calls, the chance of at least one tail event is close to three times a single call's. Fan-out improves the mean and degrades the tail, which is why the slice is sized on the fan-out's p99 and not the store's.
  • Why set an absolute deadline rather than a timeout per stage?
    Per-stage timeouts compose by addition, so overspend early in the chain is invisible until the total is blown. An absolute deadline computed at arrival makes every stage account for what has already been spent, and lets the scorer guarantee an answer inside the budget while degrading what the answer is based on.
  • Where should the decision record be written, given the 10 ms emit slice?
    Hand it to an asynchronous path and answer the login service immediately; the 10 ms covers forming the record and enqueueing it, not durably storing it. A synchronous write puts a storage system's tail inside a hard inline budget, and storage tails are exactly the kind of latency this budget cannot absorb.

saying these in an interview costs you the question

  • Budgeting against the average scoring latency instead of the p99
  • Per-stage timeouts that sum to more than the total budget
  • Adding the parallel feature lookups together instead of taking the slowest
  • Writing the decision record synchronously before answering
  • Leaving no reserve, so a timeout produces no answer at all
  • Assuming the model forward pass is the dominant cost on this path