The synchronous gate's p99 latency budget is fixed and every team wants a rule in it — how do you allocate it?
answer
- individually cheap, collectively expensive
- the budget needs a named owner
- measure the delta before merging
- blocking is a promotion, not a default
- more replicas do not shorten one decision
basics
~20 sTreat the budget as a finite shared resource with an owner. Require a measured latency cost per proposed rule, admit it to the blocking path only when the prevention is worth the spend, and default the rest to out-of-band checks.
solid answer
~50 sThe trap is that every request is individually reasonable — one more rule, two milliseconds — while the sum is what every caller pays, forever. So make the cost visible and attributable: require a measured latency delta on a representative input before a rule joins the synchronous path, and publish the running total against the budget so teams see what is left. Then bias the default. A new rule starts as an out-of-band check that reports, and is promoted to the blocking path only when someone can argue that catching this at the gate rather than an hour later is worth milliseconds charged to everyone. When the budget is full there are three honest moves: spend less per rule with narrower match predicates, spend less overall by trimming the input, or decide which existing rule the new one displaces. That last call belongs to the team that owns the budget, not the one requesting the rule.
go deeper
Know that every rule added to a blocking gate costs time on every request that passes through it, so the number of rules is limited by a budget rather than by how good each idea is.
Describe a promotion path: a new rule reports out of band first and moves onto the blocking path only after its measured cost and its prevention value have both been examined.
Show that you measure a candidate rule's tail cost against a representative input, and that you know the two levers that create budget — cheaper match predicates and a smaller projected input — before asking for more.
Own the allocation: who holds the budget, what displaces what when it is full, and how you keep refusing reasonable requests without becoming the reason teams route around the gate entirely.
## Why this is a governance problem, not a tuning problem The technical facts are simple and are covered elsewhere: decision cost grows with rule count multiplied by input size, and the budget is derived from what the caller can wait for. What makes this a lead's problem is the arithmetic of consent. Each individual request is modest and well-motivated, the requesting team pays none of the cost, and the people who pay — everyone deploying through the gate — are not in the room. That is a textbook commons, and commons do not resolve themselves by everyone being reasonable. ## Give the budget an owner and a number The first move is to say out loud that the synchronous path has a fixed capacity: "the gate has 400 ms of p99 and it is currently spending 310." A budget nobody has quantified cannot be defended, and the conversation degenerates into whether each individual rule is important — a debate the security team always wins and always regrets, because the endpoint is a gate so slow that teams route around it. Ownership matters as much as the number. The platform team that runs the gate holds the budget, publishes the spend, and makes the displacement calls. Requesting teams argue for value; they do not get to price their own rule. ## Make cost measured, not asserted Before a rule joins the blocking path, measure it: evaluate the current rule set and the set plus the candidate against the *same representative input*, repeatedly, and report the difference at the tail. Representative matters more than large — a rule that is free against an empty namespace and expensive against the biggest estate you own must be measured against the biggest estate you own. The output is a number attached to the change request, which converts "this feels cheap" into something reviewable. ## Bias the default toward not blocking The most effective structural decision is to make the *non-blocking* path the default and easy one. A new rule ships as an out-of-band check that evaluates the same estate on a schedule and reports what it finds. It costs nobody latency, it needs no negotiation, and it accumulates the evidence that decides the next question. Promotion to the blocking path then requires one argument, and it is not seniority: **what happens between the violation and the fix?** If a violation that reaches production for an hour is visible, bounded and recoverable, out-of-band detection is genuinely enough. If it is not — a listener that will negotiate TLS 1.0 with the public internet in the meantime — the case for spending budget is real and easy to make. Framing it this way also gives you a defensible no that is about consequence rather than about importance, which is the only kind of no that survives contact with a determined team. ## When the budget is full Three honest options, in order of preference: 1. **Reduce the cost per rule.** Narrow match predicates so rules exit immediately on calls they cannot apply to. Most rules are irrelevant to most calls, so this is usually the biggest untapped win. 2. **Reduce what everything is evaluated against.** Project the input to the fields the rules read and to the items that can be affected. This creates budget for every rule at once and changes no rule's logic. 3. **Displace.** Decide which existing rule leaves the blocking path to make room. This is uncomfortable and it is the point of having an owner. And two dishonest ones to name explicitly, because they get proposed in every such meeting. Raising the timeout does not create budget; it moves the pain to the caller and hides a breach. Adding engine replicas buys throughput and cuts queueing when the engine is saturated, but a single decision still takes as long as it takes — if the p99 is dominated by evaluating forty rules over a large document, capacity barely moves it. ## Re-check on a cadence, because the estate grows An allocation made against 500 listeners is not an allocation against 5,000. The budget is a function of an input that grows without anyone deciding to grow it, so the review is periodic and mechanical: re-measure the current spend against the current input, and treat input size as a tracked quantity in its own right. A gate that quietly slid over budget teaches teams that the gate is why deploys are slow, and that lesson is much harder to unteach than it is to avoid. ## What a strong answer sounds like A good candidate names the commons, names an owner, gives the promotion criterion in terms of consequence rather than importance, and knows the two levers that create budget without a fight. A weak one either approves everything (and ends up with a gate teams route around) or refuses everything (and ends up with a gate nobody adopts). The judgment being tested is whether you can keep the gate both meaningful and tolerable at the same time.
- How do you measure a proposed rule's contribution before it ships?Evaluate the current rule set, and the set plus the candidate, against the same representative input, repeatedly, and report the difference at the tail. Representative beats large: a rule that is free against a small namespace and expensive against your biggest estate has to be measured against the biggest estate. Attach the number to the change request so the cost is reviewed rather than asserted.
- Does adding engine replicas buy you more latency budget?It buys throughput and cuts queueing when the engine is saturated, but one decision still takes as long as it takes. If the p99 is dominated by evaluating forty rules over a large document, replicas move it very little. The levers that actually create budget are narrower match predicates and a smaller projected input.
- A team insists their rule must block rather than warn. What do you ask them?What happens between the violation and the fix. If a violation reaching production for an hour is visible, bounded and recoverable, an out-of-band check is enough. If it is not — a listener that will negotiate a weak protocol with the public internet meanwhile — the case for spending shared budget is real. That question decides promotion, not who is asking.
saying these in an interview costs you the question
- Approving rules one at a time with no running total
- Raising the timeout when the budget is exceeded
- Adding replicas to fix per-decision latency
- Letting each requesting team price its own rule
- Blocking by default because every rule feels important