skip to content

Rules in Production

Once a rule blocks real work the engine is production: someone must explain a denial, hold a latency budget, and admit when a rule has stopped earning its place. Interviewers probe all three.

on this pageshow

explore

questions

16

Why does a synchronous policy gate need a stated p99 latency budget and a hard timeout?

level: juniorimportance: must knowfreq 62%

answer

  1. the caller is the one waiting
  2. budget comes from the caller's deadline
  3. the tail, not the average
  4. target and cut-off are different numbers
  5. timeout strictly below the caller's deadline

basics

~20 s

A synchronous gate runs inside someone else's request, so its time is added to theirs. The p99 target bounds the tail that callers actually feel, and the hard timeout caps the worst case so the caller's own deadline stays predictable.

solid answer

~50 s

Every synchronous decision is time the caller spends waiting — a deploy, a provisioning request, a pipeline step. So the budget is derived top-down: take the caller's deadline, subtract transport and serialisation, and what is left is what the engine may spend. State it as a percentile rather than a mean, because the tail is what people experience: at ten thousand decisions a day, a p99 of 800 ms is a hundred slow deploys every day. The hard timeout is a different number from the target — it is the point where you stop waiting at all. Without one, a wedged or overloaded engine becomes an unbounded hang inside the caller. Set it strictly below the caller's own deadline, so the caller, not the engine, is the one that decides what happens next. What the caller then does when the timeout fires is a separate design decision from choosing the number.

go deeper

for a junior

Be ready to say in one line what a latency budget is — how long the decision may take before the caller feels it — and that a call with no timeout can hang the caller indefinitely.

for a middle

Explain where the number comes from: the caller's deadline minus transport and serialisation, stated as a percentile because the tail is what people experience, with the timeout as a separate, smaller, enforced cut-off.

for a senior

Show that you measure at the caller rather than trusting engine-side timing, and that you can defend a specific pair of numbers for a specific caller instead of asserting that the gate 'should be fast'.

for a principal

Own the framing that decision latency is a cost charged to every team on the path, so the budget belongs to the platform running the gate and is negotiated with its callers, not chosen by whoever wrote the most recent rule.

## What "synchronous" actually costs A policy decision is *synchronous* when something is waiting on the answer before it can proceed: a deployment that will not continue, a provisioning request that is not yet applied, a pipeline step that has not yet gone green. The moment a decision is on that path, its latency stops being an engine statistic and becomes part of somebody else's user experience. That is the whole reason this leaf exists: a rule that is correct but slow is still a production problem, and on call you will be paged for the second one, not the first. ## Where the number comes from A latency budget is not invented for the engine; it is *derived from the caller*. Work top-down: 1. Start with the caller's own deadline — how long the deploy tool, the provisioning API or the pipeline step is willing to wait in total before it gives up or before a human notices. 2. Subtract everything that is not evaluation: serialising the input, network transport in both directions, queueing at the engine, deserialising the answer. 3. What is left is the evaluation budget. On a small input those overheads are noise. On a large one — an inventory of several thousand listeners handed to a gate that checks each one negotiates at least TLS 1.2 — serialisation and parsing can easily be the largest term, and the engine's own "evaluation took 6 ms" is then almost irrelevant to the number the caller sees. ## Target and timeout are two different numbers People conflate these constantly, and the distinction is worth being crisp about in an interview: - The **target** is an objective you observe and track — "p99 under 300 ms". It is measured continuously, it can be missed without anything breaking, and missing it is a signal to go and look. - The **timeout** is enforced. It is the point at which the caller abandons the call. It is not a goal, it is a bound on damage. The timeout must be strictly smaller than the caller's own deadline, with headroom for whatever the caller does next. If the gate's timeout is longer than the caller's deadline, the gate's timeout is dead code: the caller gives up first, and the engine keeps burning CPU on an answer nobody will read. A callee's timeout can never extend a caller's patience. A missing timeout is worse than a generous one. Without a cut-off, an engine that is merely slow — garbage collecting, cold after a restart, or handed an input ten times larger than usual — propagates as an indefinite hang in every caller at once. What the caller *does* when the timeout fires (proceed, refuse, retry, fall back) is a genuinely separate design decision; the point here is that there must be a bounded moment at which that decision gets made. ## Why a percentile rather than a mean or a maximum A mean hides exactly the cases that generate complaints. If 99% of decisions take 20 ms and 1% take four seconds, the mean is about 60 ms and looks excellent, while one deploy in a hundred stalls visibly. A maximum, at the other extreme, is set by the single worst outlier you ever recorded; it cannot be tracked or defended, and one cold start ruins it forever. A high percentile — p95, p99, sometimes p99.9 if the call volume is large — is the number that is both defensible and measurable. The enforced maximum role is played by the timeout. ## Measure at the caller Engine-side timing tells you about evaluation only. The number that matters is the one the caller observes, because it includes serialisation, transport, queueing and the response trip. Instrument the call site, and keep the engine's own timing as a *breakdown* of that total. A gap between the two is normal and informative: it tells you the fix is in what you send, not in what the rules do. ## Does a non-blocking decision need a budget? If it is on the synchronous path, yes. Whether the answer blocks anything is independent of whether the caller waits for it. A gate that only records violations but is still called inline still spends the caller's time. The honest fix for such a check is to take it off the request path entirely — at which point it has no latency budget at all, only a freshness expectation. ## The shape of a good answer "Our provisioning path allows 2 s end to end. The gate gets 400 ms as a p99 target and a 700 ms hard timeout, measured at the caller. Evaluation is typically 15 ms; the rest is shipping the inventory, which is why our first optimisation was to send less of it." Concrete numbers, derived from a caller, measured in the right place.

  • Where should you measure decision latency — at the engine or at the caller?
    At the caller, keeping the engine's own timing as a breakdown. The caller's number includes serialising the input, transport, queueing and deserialising the answer, which on a large input is often most of the cost. An engine reporting 8 ms of evaluation while the caller waits 900 ms is a normal and highly informative discrepancy — it says the fix is in what you send, not in the rules.
  • Why state the budget as a percentile rather than as a maximum?
    A maximum is set by whatever the worst outlier ever was, so it cannot be tracked or defended and a single cold start ruins it. A percentile is a target you can measure continuously and alert on. The enforced maximum role belongs to the hard timeout, which is a bound on damage rather than an objective.
  • Does a decision that only warns still need a latency budget?
    Yes, if the caller waits for it. Whether the answer blocks anything is independent of whether it is on the request path. If the result is only recorded, the right move is to take the call off the path entirely — then there is no latency budget at all, only a freshness expectation on the record.

saying these in an interview costs you the question

  • Quoting an average decision time and ignoring the tail
  • Setting the gate's timeout equal to or above the caller's deadline
  • Treating a missing timeout as harmless because the engine is usually fast
  • Measuring latency inside the engine only, never at the caller
  • Treating the budget as a property of the engine rather than of the caller

context

open as a page

Your CI gate denied a build on policy. What is your first step to reproduce that denial?

level: juniorimportance: must knowfreq 62%

basics

~10 s

Capture the exact input document the engine evaluated and the version of the rules it used, then replay that pair on your own machine. Reproduce the decision before you start editing the job definition.

open as a page

A policy rule has not denied a single change in twelve months - is it safe to delete?

level: juniorimportance: must knowfreq 56%

basics

~20 s

Not yet. Zero denials is ambiguous: it can mean nothing in scope violated the rule, or that the rule never matched anything at all. Compare how often it was evaluated with how often it denied before deciding.

open as a page

Which platform changes can make a working policy rule stop enforcing with no error at all?

level: middleimportance: must knowfreq 66%

basics

~20 s

Anything that moves the shape the rule reads: a field renamed or relocated between API versions, a kind served under a new group-version, a field that gains a default so it is never absent, or an engine change to how the input document is built.

open as a page

Why keep a policy test fixture that your rule is expected to deny, and what does its sudden pass mean?

level: juniorimportance: should knowfreq 52%

basics

~20 s

A rule that stops matching anything denies nothing, and that silence looks exactly like success. A fixture the rule must deny turns the silence into a failing test: if it suddenly passes, enforcement has quietly stopped.

open as a page

A gate evaluates 40 rules against an inventory of 5,000 listeners — where does the time go?

level: middleimportance: should knowfreq 51%

basics

~20 s

Mostly in two terms that multiply: marshalling the whole inventory once per call, then every rule walking every listener. Cost tracks rule count times input size, so one extra rule is paid on every call and on every item.

open as a page

Your local replay allows a CI job the gate denied. What could differ between the two evaluations?

level: middleimportance: should knowfreq 47%

basics

~20 s

One of the decision's inputs differs. Either you replayed a rebuilt document rather than the recorded one, a different rule revision, or reference data that has since changed. Chase the difference; do not call the gate flaky.

open as a page

A policy gate caches allow decisions keyed on listener name and port. Why did a listener that dropped to TLS 1.0 still get an allow?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The cache key covered the object's identity, not the content the rule reads. Nothing in the name or port changed when the TLS minimum did, so the gate returned an old allow without evaluating anything.

open as a page

A policy gate denied every entry of a CI matrix job. How do you tell a rule defect from a real violation?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Diff the documents the engine actually evaluated against the shape the rule author assumed. If expansion produced entries where the field the rule reads is absent rather than wrong, the rule met an input it was never written for.

open as a page

Before a policy engine major upgrade, how do you prove your rules still deny what they denied before?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Replay a corpus of the estate's real stored objects through both the old and the new engine-and-rules pair, then diff the decisions. Deny-to-allow flips are enforcement regressions; allow-to-deny flips are the blast radius you are about to inflict.

open as a page

A rule defaults audit-log retention to 30 days; another denies anything under 90. How do you fix this?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Two rules own the same field with different intents: validation judges the value the defaulting rule just wrote. Give the field one authority - make the default compliant - and test the two rules together rather than separately.

open as a page

How do you retire a policy rule that other teams' pipelines import, without breaking them?

level: principalimportance: should knowfreq 35%

basics

~20 s

Treat the rule as a published interface with consumers. Decide whether the control is going or only this encoding, enumerate who imports it, announce a dated removal sized to the slowest consumer, then actually delete it - never leave an always-allow stub.

open as a page

Your audit-log retention rule duplicates a control your nightly benchmark run already checks - which one survives?

level: middleimportance: nice to knowfreq 29%

basics

~20 s

Usually both survive, because they are not the same control. The gate rule is preventive and stops a non-compliant change before it lands; the nightly benchmark is detective and reports live systems afterwards. Neither covers the other's blind spot.

open as a page

The synchronous gate's p99 latency budget is fixed and every team wants a rule in it — how do you allocate it?

level: principalimportance: nice to knowfreq 31%

basics

~20 s

Treat the budget as a finite shared resource with an owner. Require a measured latency cost per proposed rule, admit it to the blocking path only when the prevention is worth the spend, and default the rest to out-of-band checks.

open as a page

Every policy denial reaches your team as a ticket. How do you make denial triage self-service?

level: principalimportance: nice to knowfreq 33%

basics

~10 s

Remove your team's monopoly on explanation: publish the evaluated input, version the rules so the deciding revision can be fetched, and ship one command that replays both and names the rule and field path.

open as a page

Your policy engine major and your rule library both need upgrading. Which moves first, and why?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Neither, as a big-bang. Find a rule-library version both engine majors accept, ship and bake that, then move the engine, then adopt new-engine-only constructs. Two small windows beat one large one, and you must name the window where enforcement is weakest.

open as a page