Why does a synchronous policy gate need a stated p99 latency budget and a hard timeout?
answer
- the caller is the one waiting
- budget comes from the caller's deadline
- the tail, not the average
- target and cut-off are different numbers
- timeout strictly below the caller's deadline
basics
~20 sA synchronous gate runs inside someone else's request, so its time is added to theirs. The p99 target bounds the tail that callers actually feel, and the hard timeout caps the worst case so the caller's own deadline stays predictable.
solid answer
~50 sEvery synchronous decision is time the caller spends waiting — a deploy, a provisioning request, a pipeline step. So the budget is derived top-down: take the caller's deadline, subtract transport and serialisation, and what is left is what the engine may spend. State it as a percentile rather than a mean, because the tail is what people experience: at ten thousand decisions a day, a p99 of 800 ms is a hundred slow deploys every day. The hard timeout is a different number from the target — it is the point where you stop waiting at all. Without one, a wedged or overloaded engine becomes an unbounded hang inside the caller. Set it strictly below the caller's own deadline, so the caller, not the engine, is the one that decides what happens next. What the caller then does when the timeout fires is a separate design decision from choosing the number.
go deeper
Be ready to say in one line what a latency budget is — how long the decision may take before the caller feels it — and that a call with no timeout can hang the caller indefinitely.
Explain where the number comes from: the caller's deadline minus transport and serialisation, stated as a percentile because the tail is what people experience, with the timeout as a separate, smaller, enforced cut-off.
Show that you measure at the caller rather than trusting engine-side timing, and that you can defend a specific pair of numbers for a specific caller instead of asserting that the gate 'should be fast'.
Own the framing that decision latency is a cost charged to every team on the path, so the budget belongs to the platform running the gate and is negotiated with its callers, not chosen by whoever wrote the most recent rule.
## What "synchronous" actually costs A policy decision is *synchronous* when something is waiting on the answer before it can proceed: a deployment that will not continue, a provisioning request that is not yet applied, a pipeline step that has not yet gone green. The moment a decision is on that path, its latency stops being an engine statistic and becomes part of somebody else's user experience. That is the whole reason this leaf exists: a rule that is correct but slow is still a production problem, and on call you will be paged for the second one, not the first. ## Where the number comes from A latency budget is not invented for the engine; it is *derived from the caller*. Work top-down: 1. Start with the caller's own deadline — how long the deploy tool, the provisioning API or the pipeline step is willing to wait in total before it gives up or before a human notices. 2. Subtract everything that is not evaluation: serialising the input, network transport in both directions, queueing at the engine, deserialising the answer. 3. What is left is the evaluation budget. On a small input those overheads are noise. On a large one — an inventory of several thousand listeners handed to a gate that checks each one negotiates at least TLS 1.2 — serialisation and parsing can easily be the largest term, and the engine's own "evaluation took 6 ms" is then almost irrelevant to the number the caller sees. ## Target and timeout are two different numbers People conflate these constantly, and the distinction is worth being crisp about in an interview: - The **target** is an objective you observe and track — "p99 under 300 ms". It is measured continuously, it can be missed without anything breaking, and missing it is a signal to go and look. - The **timeout** is enforced. It is the point at which the caller abandons the call. It is not a goal, it is a bound on damage. The timeout must be strictly smaller than the caller's own deadline, with headroom for whatever the caller does next. If the gate's timeout is longer than the caller's deadline, the gate's timeout is dead code: the caller gives up first, and the engine keeps burning CPU on an answer nobody will read. A callee's timeout can never extend a caller's patience. A missing timeout is worse than a generous one. Without a cut-off, an engine that is merely slow — garbage collecting, cold after a restart, or handed an input ten times larger than usual — propagates as an indefinite hang in every caller at once. What the caller *does* when the timeout fires (proceed, refuse, retry, fall back) is a genuinely separate design decision; the point here is that there must be a bounded moment at which that decision gets made. ## Why a percentile rather than a mean or a maximum A mean hides exactly the cases that generate complaints. If 99% of decisions take 20 ms and 1% take four seconds, the mean is about 60 ms and looks excellent, while one deploy in a hundred stalls visibly. A maximum, at the other extreme, is set by the single worst outlier you ever recorded; it cannot be tracked or defended, and one cold start ruins it forever. A high percentile — p95, p99, sometimes p99.9 if the call volume is large — is the number that is both defensible and measurable. The enforced maximum role is played by the timeout. ## Measure at the caller Engine-side timing tells you about evaluation only. The number that matters is the one the caller observes, because it includes serialisation, transport, queueing and the response trip. Instrument the call site, and keep the engine's own timing as a *breakdown* of that total. A gap between the two is normal and informative: it tells you the fix is in what you send, not in what the rules do. ## Does a non-blocking decision need a budget? If it is on the synchronous path, yes. Whether the answer blocks anything is independent of whether the caller waits for it. A gate that only records violations but is still called inline still spends the caller's time. The honest fix for such a check is to take it off the request path entirely — at which point it has no latency budget at all, only a freshness expectation. ## The shape of a good answer "Our provisioning path allows 2 s end to end. The gate gets 400 ms as a p99 target and a 700 ms hard timeout, measured at the caller. Evaluation is typically 15 ms; the rest is shipping the inventory, which is why our first optimisation was to send less of it." Concrete numbers, derived from a caller, measured in the right place.
- Where should you measure decision latency — at the engine or at the caller?At the caller, keeping the engine's own timing as a breakdown. The caller's number includes serialising the input, transport, queueing and deserialising the answer, which on a large input is often most of the cost. An engine reporting 8 ms of evaluation while the caller waits 900 ms is a normal and highly informative discrepancy — it says the fix is in what you send, not in the rules.
- Why state the budget as a percentile rather than as a maximum?A maximum is set by whatever the worst outlier ever was, so it cannot be tracked or defended and a single cold start ruins it. A percentile is a target you can measure continuously and alert on. The enforced maximum role belongs to the hard timeout, which is a bound on damage rather than an objective.
- Does a decision that only warns still need a latency budget?Yes, if the caller waits for it. Whether the answer blocks anything is independent of whether it is on the request path. If the result is only recorded, the right move is to take the call off the path entirely — then there is no latency budget at all, only a freshness expectation on the record.
saying these in an interview costs you the question
- Quoting an average decision time and ignoring the tail
- Setting the gate's timeout equal to or above the caller's deadline
- Treating a missing timeout as harmless because the engine is usually fast
- Measuring latency inside the engine only, never at the caller
- Treating the budget as a property of the engine rather than of the caller