A policy gate caches allow decisions keyed on listener name and port. Why did a listener that dropped to TLS 1.0 still get an allow?
answer
- the key was identity, not content
- hash what the engine actually read
- the rule version belongs in the key
- purity is the precondition for caching
- stale allow and stale deny differ
basics
~20 sThe cache key covered the object's identity, not the content the rule reads. Nothing in the name or port changed when the TLS minimum did, so the gate returned an old allow without evaluating anything.
solid answer
~50 sA decision cache is sound only when its key is a digest of everything the decision depended on: the exact input the engine evaluated, the version of the deployed rules, and the version of any external data those rules consulted. Keying on identity — name and port — caches per object rather than per content, so a change to a field the rule reads is invisible and the stale allow is served indefinitely. Fix the key first: hash the projected input itself, and fold the policy revision in so tightening a rule makes earlier entries unreachable rather than leaving old allows to age out. Two further constraints: a decision that reads anything outside its input, such as the clock or a live lookup, is not a pure function and should not be cached; and a stale allow is a control failure while a stale deny is an inconvenience, so they do not deserve the same TTL.
code
json · 13 lines{
"cacheKeyMaterial": {
"listener": "public-api-443",
"port": 443
},
"listenerAsEvaluated": {
"listener": "public-api-443",
"port": 443,
"protocol": "HTTPS",
"tlsMinimumVersion": "TLSv1.0",
"...": "..."
}
}go deeper
Know that caching a decision is only safe when the key changes whenever anything the decision looked at changes, and that a resource name or id is not that.
Explain the three parts of a sound key — the projected input, the policy revision, and the version of any external data — and why a decision that reads live state should not be cached.
Run the diagnosis: confirm from the key material which field is missing, size the exposure window from the logs, and choose deliberately between fixing the key, dropping the cache, and shortening the TTL in the interim.
Own the asymmetry. A stale allow is a control failure you may have to report; a stale deny is an inconvenience. That difference belongs in the TTLs, in what gets cached at all, and in whether you fund out-of-band sampling.
## What a decision cache actually assumes Caching a policy decision is memoisation, and memoisation is only valid for a **pure function**. The function here is: `decide(input, rules, external data) -> allow | deny + reasons` A cache entry is a claim that if you call `decide` again with the same arguments you will get the same answer. That claim is only true if the key covers *all three* arguments. Almost every stale-decision incident is a key that covers fewer than three. ## The defect in this scenario Keying on listener name and port keys on **identity**, not on **content**. The identity of a listener is exactly the thing that does not change when someone weakens its TLS configuration — that is what identity means. So the gate looked up `public-api-443`, found an allow recorded earlier, and returned it without evaluating a single rule. The rule set was correct, the engine was healthy, the gate reported an excellent p99 (a cache hit is fast), and a non-compliant listener sailed through. From the on-call chair this is a particularly unpleasant shape of incident: nothing is red, nothing is slow, and nothing is alerting. The same defect appears in a subtler form when the key *is* a content digest but is computed over a **projection** — the trimmed input someone introduced to cut latency. If the projection was built for the key rather than derived from what the rules read, any field outside it is invisible. The lesson is to derive one projection, feed it to the engine, and hash *that same projected document*. Never hash a hand-written subset. ## The three parts of a sound key 1. **The exact input the engine evaluated.** Hash the bytes you actually sent, after projection, in a canonical form (stable key ordering, stable number formatting) so that a semantically identical document does not produce two different keys. 2. **The policy revision.** Rules change. If the revision is not in the key, tightening a rule leaves every previously cached allow valid, and the new rule takes effect only as old entries expire — which is precisely when nobody is watching. Putting the revision in the key means publishing a new one makes all prior entries unreachable, and rolling back reuses them. That is far more reliable than trying to enumerate and purge affected entries. 3. **The version of any external data.** If rules consult a list of approved settings or an ownership table, that data is an argument to the function. Version it and include the version. ## When you should not cache at all If the decision reads anything the key cannot see, it is not cacheable. The classic cases are a rule that consults the current time (an expiring exception, a change-freeze window) and a rule that performs a live lookup during evaluation. Both make `decide` impure: the same arguments legitimately yield different answers at different moments, and a cache turns that from a feature into a bug. Either remove the impurity — pass the time or the lookup result *in* as part of the input, so it becomes part of the key — or do not cache those decisions. ## Allow and deny are not symmetric A stale **allow** is a control failure. Something that your policy forbids passed the gate, and if the gate is the evidence that the control operated, that evidence is now false for the affected window. A stale **deny** is an availability failure: somebody fixed their listener and the gate keeps refusing it, which is infuriating but visible, self-reported within minutes, and harmless to your security posture. That asymmetry should show up in the design: shorter TTLs on allows than on denies, or caching only the expensive negative path, or not caching allows at all when the decision guards something exposed. It also decides your incident severity — the two look identical in a dashboard and are not remotely the same problem. ## TTL is a bound on damage, not a fix A tempting response to this incident is "we will shorten the TTL to five minutes". That does not fix a wrong key; it means a non-compliant listener is allowed for up to five minutes instead of forever, and it makes the defect look self-healing so nobody chases it. Fix the key. Keep a TTL afterwards as a bound on the things a correct key genuinely cannot observe — clock drift, an external data source you version imperfectly, an operator change nobody told you about. ## Catching this before an auditor does The cheap control is **out-of-band sampling**: periodically take a small random slice of cached allows, re-evaluate them fresh away from the request path, and alert on any disagreement between the fresh answer and the cached one. It costs a fraction of full evaluation, it is not on anyone's latency budget, and it catches key defects, stale external data and policy-version bugs as a single class. Sizing the exposure window afterwards is then a question your logs can answer, which is the difference between a two-line incident note and a very bad conversation.
- Would a shorter TTL have prevented this?Only by accident, and only after the window had passed. A wrong key returns a wrong answer for as long as the entry lives, so a five-minute TTL still allows a non-compliant listener for five minutes and makes the defect look self-healing so nobody chases it. Fix the key. Keep the TTL afterwards as a bound on damage from the things a correct key genuinely cannot observe.
- How do you invalidate cached decisions when a rule is tightened?Make the deployed policy revision part of the key rather than trying to purge entries. Publishing a new revision changes every key, so all prior entries become unreachable and the first call after the change is evaluated fresh. It also gives you a clean rollback: reverting the revision makes the old entries reachable again, and you never have to enumerate which entries a rule change affected.
- How would you have caught this before an auditor did?Sample out of band: periodically re-evaluate a small random slice of cached allows away from the request path and alert whenever the fresh answer disagrees with the cached one. It costs a fraction of full evaluation, sits on nobody's latency budget, and catches key defects, stale external data and policy-version bugs as one class.
saying these in an interview costs you the question
- Keying a decision cache on a resource name or id
- Omitting the policy revision, so old allows outlive a tightened rule
- Caching decisions whose rules read the clock or a live lookup
- Giving allows and denies the same generous TTL
- Treating a shorter TTL as the fix for a wrong key