Your service caches authorization decisions to cut per-request evaluation cost — what must the cache key contain, and what does the TTL cost you?
answer
- one entry, one fully specified input
- no caller in the key, shared answers
- clamp to the next boundary
- the TTL is the revocation lag
- two staleness windows add up
basics
~20 sThe key must carry every input the rule read — caller identity, action, resource identity and version, and the rule-set version — or entries collide. The TTL is the worst-case lag on a revocation: a cached allow keeps answering until it expires.
solid answer
~50 sA stored decision is an answer to one fully specified input, so the key must identify that input: the caller (`sub`), the action, the resource identity **and its version**, plus the rule-set version. Drop the caller and two researchers share an answer, which is the catastrophic version of this bug. Drop the resource version and a lifted embargo is not honoured. Drop the rule-set version and a rule change takes effect only when entries expire. A rule that turns on a clock — "after the embargo date" — is cacheable only up to the instant it flips, so clamp the entry to the earlier of the TTL and the next boundary. And the TTL is not free: it is the worst-case lag between withdrawing an affiliation and the withdrawal biting. That lag adds to any staleness already frozen into the credential.
code
pseudocode · 21 linesfunction cachedDecide(subject, action, animal, ruleSet):
key = join("|", [
subject.sub, # without this, callers share answers
action,
animal.id,
animal.updatedAt, # a new embargo changes the key
ruleSet.version # a rule change changes the key
])
hit = cache.get(key)
if hit is not null:
return hit.decision
result = evaluator.evaluate(buildInput(subject, action, animal))
# a time-bounded rule must not outlive the instant it flips
lifetime = min(configuredTtl, result.nextBoundary - serverClock.now())
if lifetime > 0:
cache.put(key, result.decision, lifetime)
return result.decisiongo deeper
Remember that a stored decision answers one exact question. If the caller's identity is missing from the key, two different people share an answer — which is the bug that turns a speed-up into a leak.
List the key's parts and say what breaks when each is dropped, and explain why a rule mentioning a date cannot use a fixed lifetime. State plainly that the TTL is the lag before a withdrawal takes effect.
Bring the operational shaping: short lifetimes on allow answers, event-driven invalidation on the changes you can see, and the TTL as a backstop rather than the mechanism. Show that you measured whether the evaluation or the attribute read was the cost.
Own the total staleness budget — credential lifetime plus stored answer — as a product commitment about how fast a revocation bites for data this sensitive, and decide whether per-process caching is acceptable when it leaves an emergency revocation with no lever but waiting.
## What a stored decision actually is A decision cache does not store "this researcher may read precise fixes". It stores *the answer to one completely specified input*: this caller, this action, this animal in this state, under this rule set, at a moment when the clock had not yet crossed any boundary the rule mentions. Every part of that sentence has to be in the key, or the entry answers a question that was not asked. ## The key is the whole input - **Caller identity** (`sub`). Omit it and two researchers collide on one entry, so one inherits the other's allow. This is the failure that turns a performance optimisation into a disclosure. - **Action**. `read-coarse-cell` and `read-precise-fix` are different questions about the same animal, and a key that omits the action answers the wrong one. - **Resource identity and version**. The identity is obvious; the version — an `updatedAt`, a row version, an ETag-like marker — is what makes a newly imposed embargo take effect at once instead of at expiry. - **Rule-set version**. Without it, a rule change is invisible to every cached answer until the entry ages out, so a deployment appears to do nothing for a minute. - **Any environment attribute the rule reads** that is not the clock — a source network classification, for example — belongs in the key too, or in the decision not to cache that rule at all. A key that covers all of this has a lower hit rate than a coarse one, and that is the point: a high hit rate on a key that under-specifies the input is a high rate of wrong answers. ## Rules with a clock in them The embargo clause flips at a known instant with no write anywhere. A fixed 15-minute entry created five minutes before the embargo lapses will keep denying for ten minutes after the rule started saying yes. The fix is not to refuse to cache; it is to clamp the entry's lifetime: ``` lifetime = min(configuredTtl, timeUntilNextBoundaryTheRuleMentions) ``` That requires the evaluator to report the next instant at which its answer could change, which a rule language can compute from the comparisons it just made. Where it cannot, the honest options are a TTL short enough that the overshoot does not matter or no caching for that rule. ## What the TTL buys and what it bills | TTL | what it buys | what it bills | |---|---|---| | none | every answer current | full evaluation cost, and every live attribute read, on every request | | seconds | most of the saving | a lag measured in seconds on revocation | | minutes | a near-total saving on hot paths | a withdrawn affiliation keeps reading for that long | The honest way to state the cost is as a **lag**: a cached allow continues to answer until it expires, so the TTL *is* the worst case between withdrawing a grant and the withdrawal biting. Worse, that lag composes with the one from the previous step — an affiliation frozen into a credential with 20 minutes left, plus a 60-second entry, is roughly 21 minutes of worst-case exposure. Budget the sum. ## Allow and deny are not symmetric, but neither is free A stale **allow** is a disclosure that cannot be undone, which is the one this data cannot tolerate. A stale **deny** is a researcher who was granted access two minutes ago and is still being refused — annoying, and a support ticket, but recoverable. So the common shaping is a short TTL on allow answers, a longer one on denies, and event-driven invalidation on the changes you can actually observe: an affiliation withdrawn, a species flagged, an embargo altered. Treat the TTL as the backstop for the events you missed, not as the primary mechanism, because an invalidation path that is never exercised is a path that does not work. One more shape worth naming. A **per-process** cache is trivial to build and impossible to invalidate centrally — you can only wait it out — while a **shared** cache can be invalidated in one place but reintroduces the network hop the cache existed to avoid. Neither is wrong; the choice follows from whether your revocation story is "wait for the TTL" or "push the invalidation". ## When to cache nothing at all If the evaluator is embedded and answers in microseconds, a cache adds a staleness window in exchange for almost no latency. In that configuration the expensive part is usually the **live attribute read**, not the evaluation, and caching that read — with its own key and its own short lifetime — is the better-targeted optimisation. Measure which half actually costs you before adding a window you will have to defend after an incident.
- Do the credential's staleness and the cache's staleness add up?Yes, in the worst case they compose. An affiliation frozen into a credential with 20 minutes of life left, plus a 60-second stored allow, gives roughly 21 minutes between withdrawing the affiliation and the withdrawal taking effect. Budget the total against how fast a withdrawal must bite, then shorten whichever half is cheaper to shorten.
- When is the right amount of decision caching none at all?When the evaluator is embedded and answers in microseconds. The cache then buys almost no latency and costs a staleness window you must defend after any incident. In that setup the live attribute read is usually the expensive part, so cache that instead, with its own key and its own short lifetime.
- Why is a per-process cache harder to operate than a shared one?It cannot be invalidated centrally — you can only wait out the TTL on every instance, and an emergency revocation has no lever. A shared cache can be invalidated in one place but reintroduces the network hop the cache existed to avoid. The choice follows from whether your revocation story is "wait" or "push".
A photocopied staff phone list at every desk. It is right when printed and quietly wrong from the moment someone leaves, and how wrong it can be is exactly how long since the last reprint. Reprinting more often costs paper; reprinting less often costs wrong calls.
saying these in an interview costs you the question
- Keys entries by resource and action without the caller.
- Leaves the rule-set version out, so rule changes appear inert.
- Uses one fixed TTL for a rule that turns on a date.
- Calls a cached allow harmless because the rule was right once.
- Says a stale deny costs nothing at all.
- Adds a decision cache in front of an in-process evaluator without measuring.