skip to content

Your verifier must ask the issuer whether each access token is still valid — what do you cache, and for how long?

level: middleimportance: should knowfreq 37%

answer

  1. key on a digest, never the token
  2. never key on the subject
  3. clamp the TTL to the credential's own life
  4. a negative verdict can flip before validity starts
  5. the TTL is the revocation lag

basics

~20 s

Cache the decision you derived, keyed on a keyed digest of the token rather than the token itself, for the shorter of a configured TTL and the credential's own remaining life. The TTL you pick is exactly the revocation lag you are accepting.

solid answer

~50 s

The cache key is a keyed digest of the token string — never the raw token, and never the subject, because two credentials for one subject have different lifetimes and one may already be dead. The value is the decision you derived (principal, permissions, the answer's own expiry), so a hit produces the same object a live check would. The TTL is `min(configured TTL, time left on the credential)` and never outlives the token. Negative answers are worth caching too, so a replayed dead credential stops reaching the issuer on every attempt — with one exception: an answer that was negative only because the credential is not yet valid will turn positive, so cap negative entries short. Collapse concurrent lookups for the same token into one call. And be explicit about the price: between a revocation and the entry's expiry, your verifier still says yes.

code

pseudocode · 20 lines
pseudocode
on verify(token):
    k = keyed_digest(cache_secret, token)        # HMAC over the token; never the token itself
    hit = cache.get(k)
    if hit and hit.expires_at > now:
        return hit.decision

    answer = single_flight(k, () => issuer.introspect(token))
    if answer is error:
        return undecided(503)                    # errors are never cached as decisions

    if answer.active:
        ttl = min(config.positive_ttl, answer.exp - now)   # never past the credential's own expiry
        decision = principal_from(answer)
    else:
        ttl = config.negative_ttl                # short: a not-yet-valid credential will flip
        decision = reject(401, "invalid_token")

    if ttl > 0:
        cache.put(k, decision, expires_at = now + ttl)
    return decision

go deeper

for a junior

Recall that asking the issuer on every request is a network call on the hot path, and that caching the answer is how it is avoided — at the price of acting on slightly old information.

for a middle

Design the entry: a keyed digest as the key, the derived decision as the value, a TTL clamped to the credential's remaining life, and single-flight so a cold entry causes one lookup rather than many.

for a senior

State the budget out loud. Your TTL is the revocation lag; say what it is, how you would measure that it holds, and what a shared cache adds to the failure surface.

for a principal

Weigh the whole coupling: how much of your availability you are willing to hand to the issuer's uptime, what the lag buys you in traffic, and who signs for that number.

## What you are actually caching When verification means asking the issuer — a call to its introspection endpoint (RFC 7662) rather than a local signature check — the network hop is on the hot path of every single request. A venue's turnstiles during a two-minute arrival surge cannot each make a round trip per swipe. So you cache. The design has four parts: the key, the value, the lifetime, and the honesty about what the lifetime costs. Cache the **decision**, not the raw answer. A hit should produce the same principal object a live check would have produced, so that the two paths cannot drift apart. If the cache stores the wire response and the live path stores a mapped principal, a mapping change fixes one of them and not the other. ## The cache key - **Key on a keyed digest of the token string**, using a per-deployment secret. Two things follow: the cache's contents are not a pile of usable credentials if it is dumped or inspected, and the key is stable across instances if you ever share the cache. - **Never the raw token.** A cache key turns up in metrics labels, slow-lookup logs and debugging dumps. None of those are places a live credential should reach. - **Never the subject.** One subject can hold several credentials with different scopes and lifetimes, and revoking one must not take the others with it — or, far worse, a revoked one must not ride on another's cached yes. ## The value and its lifetime The entry holds the derived principal, the permission set, and an expiry computed as: > `expires_at = now + min(configured TTL, time remaining on the credential)` The second term is not optional. **A cache entry must never outlive the credential it describes.** If the answer reports the credential's own expiry, clamp to it; a cached yes that survives the token's expiry accepts an expired credential, which is a defect no TTL setting excuses. | Answer | Cache it? | For how long | |---|---|---| | Valid | Yes | `min(configured TTL, remaining life)` | | Not valid — expired, unknown or revoked | Yes | A short cap; the verdict will not change for this token string | | Not valid — not yet within its validity window | Briefly, or not at all | This verdict **will** flip when the window opens | | The issuer failed to answer | No | An error is not a decision and must not be stored as one | That third row is the edge that catches people. "A negative answer never becomes positive" is true for expiry and revocation and false for a credential whose validity window has not opened yet. Either cap negative entries at a few seconds, or key the cap on the reason when the answer tells you one. ## What a cached yes costs This is the sentence the whole design turns on: **for as long as an entry lives, a credential revoked a moment ago still works here.** The lag between a revocation and the last verifier noticing is the cache TTL, plus however long the revocation took to reach the issuer's own view. That number is a budget somebody signs for, and it is chosen against how much issuer traffic you are prepared to carry. It is also the number to quote when someone says the system supports immediate revocation. Sizing that budget, and the bookkeeping that makes a revocation fast on the issuer's side, are a separate subject; what belongs here is that your TTL **is** the lag, and that nobody should discover it during an incident. ## Where the cache lives | | Per instance, in memory | Shared across instances | |---|---|---| | Issuer load | Multiplied by the instance count | Close to one lookup per token | | Cold start | Every deploy starts cold | Survives a rolling deploy | | New dependency | None | One more thing whose outage stops verification | | What it holds | Digests, inside one process | Digests, in a store others can reach | Neither is the right answer everywhere. A fleet of sixty edge verifiers with modest traffic each is usually fine in memory. A large fleet in one data centre, where the issuer's capacity is the constraint, earns the shared store — as long as you have accepted that the store's availability is now part of your verification path. ## Two details that are easy to skip - **Single-flight.** Under a burst, several concurrent requests carrying the same credential should collapse into one lookup, with the rest waiting on its result. Without it, a cold entry means as many calls as there are concurrent requests. - **Keep it out of caches you did not choose.** Anything carrying the credential or the answer travels with `Cache-Control: no-store`, and neither the token nor the digest belongs in an access log. A cache you designed is a trade-off; a copy sitting in an intermediary is a leak.

  • Why not cache the fact that a credential was rejected for as long as that credential could possibly live?
    Because one rejection reason reverses. Expired, unknown and revoked verdicts are permanent for that token string, but a credential rejected because its validity window has not opened yet becomes valid later, and a long negative entry would keep refusing it for the rest of its life. Either cap negative entries at seconds, or set the cap from the reason when the answer supplies one.
  • The issuer times out mid-surge and the cache misses. What does the verifier return?
    It has no decision, so it must not manufacture one. Refuse the request with a status that says the service could not decide rather than one that says the credential is bad — a bad-credential answer sends clients off to the issuer you just failed to reach. Never store the failure as a cached verdict, positive or negative.
  • Your fleet moves from per-instance caches to a shared store and issuer traffic drops sharply. What did you take on?
    A dependency in the verification path: an outage or a failover in that store now stalls verification everywhere at once, where previously each instance degraded alone. The store also now holds token digests, so it inherits the access controls and retention rules of credential-adjacent data, and its latency is added to every cache hit.

A door list reprinted every two minutes. The reprint interval is not a detail of the printing — it is exactly how long a name cancelled upstairs still gets someone in downstairs.

saying these in an interview costs you the question

  • Uses the raw token string as the cache key
  • Keys the cache on the subject rather than the credential
  • Lets a cached entry outlive the credential it describes
  • Caches a failed lookup as if it were a rejection
  • Claims revocation is immediate while caching answers for minutes
  • Caches a not-yet-valid verdict for the credential's full lifetime