How should a JWT verifier cache and refresh issuer keys so rotation causes no outage?
answer
- Keys live in memory, not per request
- A miss is the signal, not a timer
- One fetch, however many misses
- Old and new keys coexist for a window
- No key found means reject, never skip
basics
~20 sCache the issuer's key set in memory, select keys by the token's kid, and on an unknown kid trigger one rate-limited refresh shared across threads. Serve from cache while refreshing, keep old keys through the overlap window, and reject rather than skip verification when no key is found.
solid answer
~50 sTreat the issuer's key set as a cached, refreshable resource on the verification path. Keep it in memory with a TTL, look up the token's `kid` in the cached set, and only when the `kid` is unknown fetch again — that makes rotation self-healing without polling. Guard that fetch: coalesce concurrent refreshes into one in-flight request, enforce a minimum interval between fetches, and negatively cache unknown key identifiers, otherwise a stream of tokens with bogus `kid` values turns into a request flood against the issuer, effectively a denial-of-service amplifier you built yourself. Serve stale keys while a refresh is in flight rather than blocking requests. Rotation is safe when the issuer publishes the new key before signing with it and keeps the old one published until the last token signed with it expires. If no trusted key matches, reject the token — never fall back to skipping verification.
code
json · 20 lines{
"keys": [
{
"kty": "RSA",
"use": "sig",
"alg": "RS256",
"kid": "2024-11-a",
"n": "0vx7agoebGcQSuu...",
"e": "AQAB"
},
{
"kty": "RSA",
"use": "sig",
"alg": "RS256",
"kid": "2025-02-b",
"n": "sXchDaQebHnPiGv...",
"e": "AQAB"
}
]
}go deeper
Know that the verifier holds the issuer's public keys, that kid selects which one, and that keys are cached rather than fetched per request.
Explain rotation mechanics from the verifier's side: multiple keys in the set at once, selection by kid, refresh when an identifier is unknown, and why a startup-pinned key breaks.
Show the hardening — single-flight refresh, rate limiting, negative caching of unknown identifiers, stale-while-revalidate serving, and rejecting rather than failing open when no trusted key matches.
Own the cross-service contract: a published rotation window at least as long as the maximum token lifetime, one shared verifier implementation with these guards, and metrics that make a rotation visible before it becomes an incident.
## Where this sits in validation Before a signature can be checked, the verifier must answer "which key?". For asymmetric signing the answer comes from the issuer's published key set, and the token's `kid` header names which entry to use. Everything interesting here is operational: the correctness rule is trivial (use a key you trust), but naive implementations produce either outages during rotation or traffic storms against the issuer. ## The cache is not optional Fetching keys per request would put a network call in front of every authenticated request — latency, and a hard dependency on the issuer being up. So the key set is cached in memory. Two consequences follow immediately: the cache can be *stale* when the issuer rotates, and the cache can *stampede* when many threads discover staleness simultaneously. ## Refresh on unknown kid, not on a timer alone Polling on a fixed interval is a poor primary strategy: too slow and rotation breaks authentication for up to one interval; too fast and you hammer the issuer forever to detect a change that happens quarterly. The event-driven rule is better — when a token presents a `kid` that is not in the cached set, that is the signal to refresh. Combine it with a modest background TTL so a long-idle process does not hold ancient keys. ## Guarding the refresh — the part candidates miss "Refresh when the `kid` is unknown" is exploitable as written. An attacker sends unsigned garbage tokens with random `kid` values; each one is a cache miss; each miss triggers a fetch. You have built a request amplifier pointed at your own identity provider. The guards: - **Rate limit / minimum refresh interval.** At most one refresh per N seconds regardless of how many misses occur. - **Single-flight coalescing.** Concurrent misses join one in-flight fetch instead of issuing N fetches. - **Negative caching.** Remember recently-seen unknown key identifiers for a short period and reject immediately without a fetch. - **Bounded parsing.** Cap the size of the fetched document and the number of keys accepted, so a hostile or broken response cannot exhaust memory. ## Serving during refresh While a refresh is in flight, keep answering from the existing cache. Tokens signed with keys you already hold must not queue behind a network call. Only requests whose `kid` is genuinely unknown wait — or, better, fail fast and let the client retry once the refresh lands. ## What makes rotation non-disruptive Rotation is a two-sided contract, and the verifier can only hold up its half: - The issuer publishes the new public key **before** it starts signing with it, so verifiers can learn it in advance. - Both keys remain published during an overlap window at least as long as the maximum token lifetime, so tokens signed with the outgoing key still verify. - Only after the last token signed with the old key has expired is that key withdrawn. The verifier's half: never assume a single key, always select by `kid`, tolerate several keys of the same type in the set, and re-read the set rather than pinning one key at startup. A verifier that caches "the" public key at boot will break at every rotation. ## Fetch integrity The key set must be retrieved over TLS from a URL that is part of your configuration or derived from the issuer's own discovery document — never from anything the token supplies. Validate the certificate chain; TLS is the only thing that binds the fetched keys to the issuer. Pin the issuer's host, and apply a connect and read timeout so a hung fetch cannot pile up threads. ## Failure behaviour If the key set cannot be fetched and no cached key matches the token, reject. The tempting failure modes are both wrong: skipping verification because keys were unavailable is an authentication bypass, and clearing the cache on a failed fetch destroys your ability to keep serving during an issuer outage. Keep the last good key set and prefer stale-but-trusted over empty. Also design for the issuer being down at process start. A cold cache plus an unreachable issuer means no authentication at all, so retry with backoff, expose the cache state in health signals, and consider seeding from a configured fallback for critical services. ## Observability Emit metrics for cache hits and misses, refresh attempts and failures, key-set age, and rejections due to unknown `kid`. A rising unknown-`kid` rejection rate is the earliest signal that a rotation is happening — or that someone is probing you. Log the `kid` and issuer on failure; never log the token itself. ## Interview framing Give the mechanism in one line — cache, select by `kid`, refresh on miss — then immediately name the three hazards that separate a working implementation from a naive one: refresh storms, blocking on refresh, and failing open when keys are unavailable.
- Why is refreshing the key set on every unknown kid, with no guard, a vulnerability?Because the `kid` is attacker-controlled. A stream of tokens with random key identifiers becomes a stream of cache misses and therefore a stream of fetches against your identity provider — you have built a traffic amplifier. Rate limit refreshes, coalesce concurrent ones into a single in-flight request, and negatively cache unknown identifiers for a short period.
- What must the issuer do for rotation to be invisible to verifiers?Publish the new public key before signing anything with it, keep the outgoing key published for at least the maximum token lifetime, and only then withdraw it. That overlap window means every token in circulation always has its key available. Verifiers cooperate by selecting on `kid` and tolerating multiple keys rather than pinning one at startup.
- The key set endpoint is down and your cache has expired. What do you do?Keep serving from the last good key set rather than emptying the cache — expiry should trigger a refresh attempt, not deletion. Retry with backoff. Tokens whose key you still hold verify normally; tokens naming a key you have never seen are rejected. Never fall back to accepting tokens without verification.
- Is a fixed polling interval enough on its own?Rarely. Too long and a rotation breaks authentication for up to one interval; too short and you poll continuously for an event that happens rarely. Refresh-on-unknown-kid handles rotation the moment it matters, with a modest background TTL as a safety net for long-idle processes.
saying these in an interview costs you the question
- Fetches the key set on every incoming request
- Caches one public key at startup and never re-reads
- Refreshes unconditionally whenever a kid is unknown
- Empties the cache when a fetch fails
- Accepts the token unverified if keys cannot be fetched