How do you rotate a JWT signing key using kid and a JWKS without rejecting live tokens?
answer
- Two keys must be valid at once
- Publish first, switch later
- Two different clocks govern the wait
- kid tells the verifier which one
- Old key stays until tokens expire
basics
~20 sPublish the new public key in the JWKS alongside the old one, wait longer than every verifier's cache lifetime, then switch signing to the new kid. Keep the old public key published until the last token signed with it has expired, and only then remove it.
solid answer
~50 sRotation is an overlap procedure, not a switch. **Step one:** generate the new key pair and add its public key to the JWKS with a distinct `kid`, while the issuer keeps signing with the old one. **Step two:** wait longer than the maximum cache lifetime verifiers apply to the key set, plus propagation, so that every verifier already holds both keys. **Step three:** flip the issuer to sign with the new `kid`; verifiers select on `kid` and find the key without a fetch. **Step four:** keep the old public key in the set until every token it signed has expired — maximum token lifetime plus clock-skew tolerance — then remove it. Verifiers must select strictly by `kid`, refresh on an unknown one with rate limiting, and serve from cache when the endpoint is briefly unreachable. Emergency revocation is a different procedure: you drop the key immediately and accept that live tokens die.
code
json · 8 lines{
"keys": [
{ "kty": "EC", "kid": "auth-2026-04", "use": "sig", "alg": "ES256",
"crv": "P-256", "x": "...", "y": "..." },
{ "kty": "EC", "kid": "auth-2026-08", "use": "sig", "alg": "ES256",
"crv": "P-256", "x": "...", "y": "..." }
]
}go deeper
Recall that more than one key can be published at once and that the token's kid tells the verifier which one signed it.
Explain why the new key is published before it is used, and why the old one must outlive the switch by at least the token lifetime.
Walk the full procedure with both clocks named — verifier cache lifetime before the switch, maximum token lifetime after — and say what you monitor at each step.
Own the policy: rotation cadence, where private keys live, what the emergency path costs in forced re-authentication, and how token lifetime is chosen to make that path survivable.
## Why rotation needs an overlap At any moment there are tokens in flight signed with the current key and verifiers holding a cached copy of the key set that may be minutes or hours old. A rotation that changes the signing key and the published set at the same instant breaks both groups: verifiers with a stale cache have not seen the new key, and tokens already issued reference a key you just deleted. The fix is to make the two populations overlap in time, which is exactly what the `kid` header parameter exists to support — it lets more than one key be valid simultaneously and lets each token say which one it used. ## The four-step procedure **1. Publish before you use.** Generate the new key pair, ideally inside a KMS or HSM so the private half is never exported. Add the public key to the JWKS with a fresh, distinct `kid`. Do not reuse an identifier; a `kid` that changes meaning is worse than no `kid` at all. The issuer keeps signing with the old key throughout this step. Nothing in the system changes behaviour yet. **2. Wait out the caches.** Every verifier caches the key set for some period. The safe wait is the maximum of those cache lifetimes plus a margin for deploys, cold starts, and any CDN in front of the endpoint. Publishing the key set with explicit, modest cache directives makes this window a number you control rather than a guess. Skipping this step is the single most common way rotations cause an outage. **3. Switch signing.** Point the issuer at the new key. Every new token now carries the new `kid`, and verifiers already hold that key, so they resolve it from cache with no network call. Old tokens still verify because the old key is still published. This is the moment to watch verification-error rates: a spike here means some verifier never picked up the new key, and rolling back is cheap because the old key is still active. **4. Retire the old key.** A token signed just before the switch remains valid for its whole lifetime. So the old public key must stay in the JWKS for at least the maximum access-token lifetime plus the clock-skew tolerance verifiers allow. After that, remove it from the published set and destroy or disable the private half. Removing it earlier produces exactly the failure rotation was supposed to avoid. ## What verifiers must do to make it work The issuer's procedure only works against well-behaved verifiers: - **Select by `kid`.** Trying every key in the set until one verifies appears to work and quietly removes the ability to reason about which key signed what. - **Refresh on an unknown `kid`, with a rate limit.** This is what lets an emergency rotation propagate faster than the cache TTL. Without a limit, tokens carrying random identifiers turn into an outbound request flood at the issuer. - **Cache negative results briefly**, so a burst of unknown identifiers does not retry per request. - **Serve from the last good copy on fetch failure.** Public keys change rarely; availability of the issuer's endpoint should not gate every API request. - **Bind algorithm to key.** After resolving a `kid`, confirm the key type matches the algorithm the verifier is configured to accept. ## Routine rotation versus emergency revocation These are different operations and conflating them is a common weak answer. **Routine rotation** is scheduled, overlapped, and invisible: no token is ever invalidated. Its cadence is a policy choice — commonly months — and the point is to exercise the procedure so it is boring and to bound how long any one key is exposed. **Emergency revocation** happens when a private key may have leaked. Then the overlap is the problem: you remove the compromised key from the JWKS immediately, force verifiers to refresh, and accept that every token signed with it stops validating. Users re-authenticate. Because that is disruptive, short token lifetimes are what make the emergency path survivable — with a fifteen-minute access token, the damage window and the recovery window are both small. ## The signal an interviewer is listening for A weak answer says "change the key and update the JWKS". A strong one names the two clocks that govern the procedure — the verifier cache lifetime, which sets how long you wait *before* switching, and the token lifetime, which sets how long you wait *after* — and can say what breaks if either is ignored.
- How does a verifier learn about a key that was published after its cache was populated?Two ways. The cache expires and it refetches on schedule, which is the normal path. Or a token arrives with an unknown `kid` and the verifier refetches on demand, which is what makes an emergency rotation propagate quickly. The on-demand path must be rate limited and its failures cached briefly, or attacker-supplied identifiers become a request amplifier.
- What does short token lifetime buy you in the rotation story?It shortens the tail after the switch: the old key can be retired sooner because tokens signed with it expire sooner. More importantly it makes the emergency path affordable — if the worst case for dropping a compromised key is that users re-authenticate within minutes rather than hours, you will actually pull the trigger when you need to.
- Should the kid be derived from the key material or chosen freely?Either works as long as it is unique and stable for that key. A thumbprint of the public key is convenient because it is collision-resistant and reproducible; a dated label such as an issue month is more readable during an incident. What matters is that a `kid` is never reused for different key material, since verifiers cache the mapping.
saying these in an interview costs you the question
- Swaps the key and the key set simultaneously
- Removes the old public key at switchover
- Reuses an existing kid for a new key
- Treats routine rotation and compromise identically
- Refetches the key set unbounded on unknown identifiers