skip to content

Sixty turnstile verifiers meet access tokens carrying a `kid` their cached key set lacks, and the issuer answers slowly — what should each verifier do?

level: seniorimportance: must knowfreq 48%

answer

  1. an unobtainable key is not a maybe
  2. reject now, refresh in the background
  3. cooldown, single-flight, jitter
  4. known key identifiers keep verifying
  5. 503 means undecided, 401 sends clients to the issuer

basics

~20 s

Reject those tokens now — an unobtainable key is not evidence of validity — and schedule one bounded refetch instead of one per request. Tokens whose key identifier is already held keep verifying normally, so the outage stays confined to the unknown ones.

solid answer

~50 s

Three rules, in order. **Reject immediately**: the request must not wait on the issuer, and an unknown `kid` is a decided rejection, so `401` with `invalid_token`. **Bound the refetch**: one key-set fetch per verifier per cooldown window, single-flighted in the process, with the unknown key identifier held in a small negative cache for that window and the fetch jittered across the fleet so sixty devices do not arrive together. **Keep serving what you can**: tokens naming a key you already hold verify exactly as before, so the incident covers newly rotated keys and cold starts only. A verifier with no usable key set at all is in a different position — it cannot decide anything, and `503` is the honest answer there, because answering `401` sends every client to the issuer that is already struggling. Refresh on a schedule well inside the issuer's rotation overlap so an unknown `kid` stays rare.

code

pseudocode · 21 lines
pseudocode
on verify(token):
    header = parse_header(token)
    if header.alg not in accepted_algorithms:
        return reject(401, "invalid_token")       # cheapest refusal; no lookup, no fetch

    if keyset.is_empty():
        return undecided(503)                     # cold start: nothing can be decided here

    key = keyset.get(header.kid)
    if key:
        return verify_signature(token, key)       # issuer not involved at all

    if unknown_kids.contains(header.kid):
        return reject(401, "invalid_token")       # already refused this value this window
    unknown_kids.put(header.kid, ttl = cooldown)

    if now >= refetch_allowed_at:
        refetch_allowed_at = now + cooldown + jitter()
        single_flight("keyset", () => background_refresh_keyset())

    return reject(401, "invalid_token")           # this request never waits on the issuer

go deeper

for a junior

Remember the default: if the verifier cannot obtain the key, it refuses. A key it cannot fetch is not a reason to trust a token.

for a middle

Explain the mechanics that bound the damage — a cooldown, single-flight, a small negative cache of unknown identifiers, jitter across the fleet, and a background refresh the request never waits on.

for a senior

Show that you have operated this: which tokens keep working during an issuer outage, why a cold start is the real failure, and why a 401 during an outage makes the outage worse.

for a principal

Own the trade explicitly — how stale a key set may become before you stop trusting it, who may authorise extending that, and what the fleet's refresh schedule costs the issuer at steady state.

## Reject now, refetch later An unknown key identifier means one of two things: the issuer has rotated to a key you have not fetched yet, or somebody is sending you a value that was never real. **At the moment of the request, the verifier cannot tell those apart**, which is why the control is a rate limit rather than a detection. What it can do is refuse to make the request wait. Blocking a swipe on a network round trip to a slow issuer converts one slow dependency into a queue of held connections on sixty devices, and the queue arrives faster than it drains. Reject the token, trigger the refresh in the background, and let the next swipe benefit. When the verifier holds a key set and the named key is not in it, that is a **decision**: `401` with `invalid_token`. It may be a false negative during an honest rotation, and the answer to that is not to weaken the check — it is to make unknown identifiers rare by refreshing on a schedule comfortably inside the issuer's published overlap. ## Bound the refetch, or your own fleet becomes the attack If every unknown `kid` triggers a fetch, anyone who can reach a turnstile can make sixty devices hammer the issuer by minting nonsense with a random identifier each time. Four bounds turn that into a local no-op: 1. **A cooldown per verifier** — at most one key-set fetch per window, say one every five minutes, regardless of how many misses arrive. 2. **Single-flight** — concurrent misses inside one process cause one fetch, not one each. 3. **A bounded negative cache of unknown identifiers**, so repeats of the same value are answered locally for the cooldown's duration, and an identifier the attacker keeps varying is capped by the cooldown anyway. 4. **Jitter across the fleet** — sixty devices on identical cooldowns expire together and rediscover the issuer in one spike; a spread of plus or minus twenty per cent turns a spike into a trickle. The cheap checks come first, too. A token naming an algorithm outside the accepted list is rejected on a string comparison, before any identifier is looked up — so the cheapest-to-produce garbage costs the least to refuse. ## Keep verifying everything you still can The most important operational fact is the one people forget under pressure: **an unreachable issuer does not stop signature verification.** Tokens whose key identifier is already in the cached set verify locally, exactly as they did an hour ago. The blast radius of an issuer outage is therefore: | Situation | What the verifier can do | Answer | |---|---|---| | `kid` is in the cached key set | Full verification, unchanged | Normal `200` or a normal denial | | `kid` unknown, key set present | Decide: this key is not one it trusts | `401` `invalid_token`, plus a bounded refetch | | No key set at all — cold start | Nothing; it cannot decide | `503`, and alarm | | Verification is a call to the issuer per request, and the issuer is down | Nothing beyond cached verdicts | `503` once the cached verdicts run out | ## Stale but still signed A cached key set that is older than its intended refresh is not a weaker check — the signature is verified against a real published key exactly as before. It is an **older** check, and what ages is the one thing a key set conveys beyond the keys: which keys the issuer still stands behind. A key withdrawn for cause stays trusted for as long as you keep serving the stale set. So serving stale is usually right and must still be bounded: an absolute maximum age, beyond which the verifier stops rather than trusting an arbitrarily old set, with the cap set against the issuer's rotation overlap and an alarm well before it is reached. ## Cold start is the real outage A verifier that boots with an empty cache while the issuer is unreachable can verify nothing at all. Two mitigations, both cheap: persist the last known good key set locally and load it at boot with its fetch time, and stagger restarts so a rolling deployment never leaves the whole fleet cold at once. A persisted set is still a signed set; the only thing it costs is age, which the absolute cap already governs. ## The status you return changes the load on the issuer `401` means *I have decided your credential is unacceptable*, and well-behaved clients respond by going to get a new one — from the issuer you have just failed to reach. Returning `401` for an outage therefore conscripts your own fleet's clients into amplifying it. `503` says *I could not decide*, leaves the credential's standing untouched, and asks for a retry rather than a re-issue. ## Fail closed, and the only degraded mode worth having The default is to refuse. "Accept unverified tokens until the issuer returns" means a signature nobody checked, and the window always outlives the incident. If a degraded mode exists at all it should be narrow enough to describe in one sentence: **extend the stale key-set cap**, pre-authorised, time-boxed, alarmed, logged, with a named owner and an automatic expiry. It never means skipping the signature check — every token is still verified against a real key, just an older list of them.

  • A verifier boots cold during the surge with an empty key cache and cannot reach the issuer. What can it legitimately do?
    Only what it prepared for earlier: load a persisted last-known-good key set, with its fetch time, and serve from that under the same absolute age cap. That is still a real signed key set, so nothing about the check is weakened — only its age. Without a persisted set it can decide nothing and should say so with a 503 and an alarm, not guess.
  • Why is rate-limiting the refetch a better control than trying to detect forged key identifiers?
    Because the two causes are indistinguishable at the verifier: a legitimately rotated key and an invented identifier both look like a value that is not in the cached set. There is nothing to detect. Capping how often a miss can reach the issuer makes the question moot — a stream of invented identifiers costs a local lookup each and produces at most one fetch per window.
  • What is wrong with accepting tokens without checking the signature until the issuer comes back?
    It removes the only thing that makes a token worth anything, and the window reliably outlives the incident because nobody is watching for the moment it could close. The defensible degraded mode extends how stale a key set may be, which keeps every signature check real while relaxing only the freshness of the trusted list — pre-authorised, time-boxed, alarmed and auto-expiring.
  • How do you stop unknown key identifiers being common in the first place?
    Refresh the key set on a schedule comfortably inside the issuer's published rotation overlap, so a newly signed key is already cached before the first token using it arrives. Jitter the schedule across the fleet. Then an unknown identifier is genuinely unusual, which makes it safe for it to be both rejected and rate-limited.

saying these in an interview costs you the question

  • Blocks the request while refetching the key set from the issuer
  • Refetches the key set on every unknown key identifier
  • Accepts unverified tokens for the duration of an issuer outage
  • Thinks an issuer outage stops verification of already-cached keys
  • Returns 401 during an outage, sending every client to the failing issuer
  • Serves an arbitrarily stale key set with no maximum age