How do you roll the key that signs a scheduler's session cookies without signing out every presenter in the same minute?
answer
- which rotation — the signing key
- verify with both, sign with one
- expand verification first, contract last
- overlap the longest remaining lifetime
- retire on evidence, not the calendar
basics
~20 sOverlap in a strict order: every instance must accept the new key for verification before any instance signs with it, and the retiring key stays in the verification set for at least the longest remaining lifetime of anything it signed.
solid answer
~50 sThis is key rotation, not identifier re-issue — the value in the cookie is untouched, only the key protecting it changes. Run it in four moves. First, deploy the new key into every instance's **verification** set and let that reach every instance. Second, flip signing to the new key. Third, keep verifying with the retiring key for at least the longest lifetime a value it signed can still have, plus clock skew. Fourth, drop it from the verification set, then destroy it. Getting the first two out of order means instances that lack the new key reject values signed by instances that have it — a fraction of requests failing that looks like a flapping bug. Retiring too early makes every outstanding value unverifiable at once, which at a radio station is the duty engineer locked out mid-transmission.
code
pseudocode · 17 lines# the current key signs; retiring keys only verify
keyring = { "2026-09": current_key, "2026-06": retiring_key }
current = "2026-09"
function issue(claims):
claims.kid = current
body = encode(claims)
return body + "." + mac(keyring[current], body)
function verify(cookie):
body, tag = split_last(cookie, ".")
kid = decode(body).kid
key = keyring.get(kid) # selector into OUR keyring only
if key == null: return UNAUTHENTICATED # unknown or already retired
if not constant_time_equals(tag, mac(key, body)):
return UNAUTHENTICATED
return decode(body)go deeper
Know that more than one key can be valid at once: the server signs with the newest and still accepts values protected by the previous one. That overlap is the reason a key change does not sign everybody out.
Give the order and say why each step is where it is — verification capability expands first, signing moves once, verification contracts last — and explain what an out-of-order rollout looks like from the caller's side.
Size the overlap against the longest remaining lifetime plus skew, and describe retiring on evidence: count requests still arriving under the old key identifier and remove it when that count has been zero longer than that lifetime.
Make two keys the steady state rather than an event, so the path is exercised continuously, and decide in advance the one case that inverts the rule — a suspected leak, where you cut over with no overlap and accept the mass sign-out.
Four different things get called rotation in an interview. This one is the key that signs or encrypts cookie-borne state, and it has an order of operations that is easy to state and routinely got wrong in the opposite direction. ## Which rotation this is Rolling the signing key changes the **key**; the identifier or payload in the caller's cookie is untouched. Re-issuing the identifier at sign-in or at a privilege change is a different rotation with different reasons, and rotating a stored machine credential is a third. A candidate who answers this question with "we call the framework's method that gives the caller a new identifier" has answered a different one. If the scheduler holds its sessions as server-side records under opaque references, there may be no cookie-signing key to roll at all — the question bites only where the cookie carries protected state. ## The order that keeps everyone signed in 1. **Add the new key to every instance's verification set** and let that deployment finish everywhere. Nothing signs with it yet. 2. **Flip signing to the new key**, once and only once every verifier holds it. Values signed with it are now accepted anywhere. 3. **Keep the retiring key in the verification set** for at least the longest remaining lifetime of anything it signed — the absolute lifetime of the value, plus clock skew, plus however long a caller may sit idle before returning with it. 4. **Remove it from the verification set, then destroy it**, in that order and not the reverse. The asymmetry is the whole lesson: **verification capability expands first and contracts last; signing capability moves in one step in between.** ## The two ways it goes wrong | mistake | what presenters see | how long it lasts | |---|---|---| | signing with the new key before every instance verifies it | intermittent sign-outs on some requests and not others | until the slowest instance updates | | retiring the old key before outstanding values expire | everyone signed out within the same minute | until each person signs in again | The first is the more confusing in practice, because it is proportional: if a fifth of instances lack the key, roughly a fifth of requests fail, and refreshing "fixes" it. Behind a load balancer this reads as flakiness rather than as a rollout ordering defect. The second is the more expensive. Every outstanding value becomes unverifiable simultaneously, which is why the overlap window is measured against the longest lifetime rather than the typical one. ## Knowing when it is safe to retire A key identifier inside the payload turns the retirement decision into evidence instead of arithmetic: - Verification becomes a lookup of one key rather than a trial of every key you hold, so the cost does not grow with the size of the keyring. - You can **count requests still arriving under the retiring identifier**. When that rate reaches zero and stays there past the longest lifetime, retirement is a measurement rather than a guess. - A value naming an identifier you no longer hold is simply unauthenticated. Treat the identifier as a selector into your own keyring and never as an instruction to obtain a key. ## The one time you skip the overlap If the key is believed to have leaked, the overlap is dropped **deliberately**. Every outstanding value protected by that key is forgeable by whoever holds it, so continuing to accept it is the harm you are trying to stop. You cut over immediately, accept that everyone is signed out in the same minute, and tell people why — for a station, that means warning the duty engineer before it happens rather than after. The general rule is "always overlap"; this is the case where the rule is wrong, and being able to name it is what separates a recited answer from an operated one. ## Practical notes that save the rollout - **Keep the keyring in configuration, not in code**, so adding a verification key is not a rebuild. - **Make two keys the normal state**, not an exceptional one. A path exercised twice a year fails the third time; a path always carrying a retiring slot does not. - **Alert when the retiring key is still verifying traffic** past the lifetime you expected, because that is either a clock problem or a lifetime longer than you documented. - **Roll on a schedule anyway.** A key that has never been rolled is a key nobody knows how to roll, and discovering that during a suspected leak is the worst possible moment.
- How long does the retiring key have to stay in the verification set?At least the longest remaining lifetime of anything it signed: the absolute lifetime of the value, plus clock skew between instances, plus the time a caller can plausibly sit idle before returning. Retiring on a fixed calendar date instead invalidates every outstanding value that was issued late in the old key's service.
- What does it look like when signing is flipped before every instance can verify the new key?Intermittent sign-outs, roughly proportional to the share of instances that have not updated. Reloading appears to fix it because the next request may land on an updated instance, so it reads as flakiness rather than as a rollout ordering defect — which is why the verification set is expanded first and separately.
- Does rolling the signing key give every presenter a new identifier?No. The value in the cookie is unchanged; only the key protecting it moves. Re-issuing the identifier itself happens for different reasons and on different triggers, and conflating the two leads teams to expect a rollover to end existing sessions — or worse, to rely on it for that.
- The keyring has been holding two keys for a year because nobody retired the old one. What is the harm?Any value ever signed with it is still accepted, so the rollover bought nothing: a copy taken before the rollover verifies today. Keep the retiring slot as a normal state, but retire the key itself on evidence — when no request has arrived under its identifier for longer than the longest lifetime.
Re-keying a building. You fit locks that accept both the old and new key before handing anybody a new key, then hand them out, and only when the last old key has been returned do you change the cylinders again. Doing it in the other order leaves the night engineer outside.
saying these in an interview costs you the question
- Switch the key at midnight and everyone just signs in again
- Add the new key and delete the old one together
- Rolling the signing key gives every caller a new identifier
- Keep the old key forever so nothing ever breaks
- Overlap always, even when the key has leaked