Your hub ships verification keys in each consumer's configuration file — how do you rotate the signing key without a flag day?
answer
- the clock is not yours here
- accept-both ships before sign-with-new
- measure adoption, do not assume it
- retire only after the last exp
basics
~20 sShip consumers a version that accepts two verification keys, confirm by measurement that they are running it, only then start signing with the new key, and withdraw the old one after the last token it signed has expired.
solid answer
~50 sThe overlap window has to live in the consumers' configuration, because there is nothing for them to poll. So the sequence is: first ship a consumer release that can hold and accept two verification keys at once and get it adopted; then measure adoption rather than assume it, using the client version each subscriber already reports, an explicit acknowledgement of the key it holds, or a canary where you sign a small share of tokens with the new key and watch pull failures; then cut over signing; then retire the old key once the longest-lived token signed with it has passed its `exp`, which is minutes, not months. The window is dominated by consumer adoption time, and you cannot shorten it by deploying anything of your own. If you cannot measure adoption, you are choosing between a flag day and never rotating, and whoever owns the subscriber contract should be told that in those words.
go deeper
Know that two verification keys can be trusted at the same time, and that a consumer which has not yet received the new key will reject every token signed with it. That is why the new key goes out before it is used.
Explain the ordered sequence — distribute, accept both, confirm, sign, retire — and be able to say what each step protects against and why retiring early rejects live tokens that are still perfectly valid.
Show that you size the window from consumer adoption time rather than token lifetime, that you measure adoption instead of assuming it, and that you can say out loud when the only honest options are a coordinated flag day or no rotation at all.
Own the structural fix and its cost: funding a fetchable key distribution, putting a notice period into the subscriber contract, and rehearsing the rollover on a schedule so the compromise case is a cutover rather than a project.
Several hundred subscriber newsrooms hold the hub's public verification key in a configuration file that was shipped to them. Nothing they run asks the hub for a newer one. That single fact turns a key rollover from a change the issuer schedules into a change several hundred other organisations have to make, which is why "we will just rotate at the weekend" is a flag day with better marketing. ## What makes this rotation different When consumers fetch verification keys from somewhere the issuer controls, the issuer owns the clock: make the new key available, wait out the consumers' refresh interval, then start signing. When the key arrives out of band, the issuer owns none of the clock. The window is bounded by your slowest subscriber's release process, which may be quarterly, and no deployment of yours shortens it. Two things follow at once: - **You need an inventory.** Rotation without a list of who holds which key is not a plan, it is an announcement. If you cannot enumerate consumers, the first deliverable of the rotation project is the list, not the key. - **The overlap is a consumer-side capability, not a document you publish.** If a subscriber's software can hold only one verification key, no schedule saves you — the change is simultaneous by construction. The real work of the first rotation is shipping every consumer a version that can hold two. ## The sequence 1. **Ship accept-both.** A consumer release that treats verification keys as a set and accepts a token signed by any member. Nothing is rotated yet; you are only buying the ability to rotate later. 2. **Distribute the new public key, still unused.** It sits in the set alongside the current one. Because you are not signing with it, a consumer that has it early loses nothing and a consumer that has it late has not broken anything. 3. **Measure adoption.** See below. This is the step that takes months and the step people skip. 4. **Canary the signing cutover.** Sign a small, reversible share of tokens with the new key and watch pull-failure rates and support volume. Ramp only if both stay flat. 5. **Cut over fully**, then **retire the old key** once the longest-lived token signed with it has passed its `exp`. That last step is measured in minutes and is the only short part of the whole schedule, though it is the part people picture when they hear "overlap window". ## Measuring adoption when you cannot see inside a consumer - **Version telemetry you already have.** Subscribers identify their client on every pull. Map client version to the key set that version shipped with, and adoption becomes a query you can run daily. - **An explicit acknowledgement.** Have the consumer report a short fingerprint of the verification keys it holds as part of a routine request or an operational report. It is the only signal that reflects the configuration actually loaded rather than the version notionally installed. - **The canary as measurement, not just as caution.** Signing one per cent of tokens with the new key tells you what no inventory can: whether the keys are really loaded in production. - **The honest admission.** If none of these is available, say plainly to whoever owns the subscriber contract that the choice is between a coordinated flag day and never rotating. That sentence is the deliverable; quietly hoping is not. ## Sizing the window The overlap must be at least the consumer adoption time, plus the longest lifetime of any token signed with the retiring key, plus a margin for clock differences between the hub and its consumers. In an out-of-band world the first term is measured in months and the second in minutes, so adoption dominates by orders of magnitude — which is exactly the intuition that a published-key-set world inverts, and why people who have only rotated in that world underestimate this one badly. ## Making the next rotation yours Treat the flag-day risk as a defect to design out, not a fact of life: - **Publish verification keys somewhere consumers fetch them,** and require refreshing as a condition of the feed, so the next rollover is bounded by a refresh interval you set. - **Put a deprecation clause in the subscriber contract** naming the notice period for a key change, so the slowest consumer has a contractual ceiling rather than an open-ended veto. - **Rotate on a schedule when nothing is wrong.** A rollover path that has never been exercised is not a path, and the first time you need it will be the compromise, under a clock, with the inventory half-built. The verifier's half of all this — how a consumer stores the set, when it refreshes, and what it does when it cannot — is the consumer's design problem and is decided on the other side of this boundary. The issuer's job is the schedule, the measurement and the retirement date.
- When the issuer does publish verification keys for consumers to fetch, what ordering must its operators follow?Make the new public key available first and leave it unused, wait at least one full consumer refresh interval plus a safety margin so that a consumer which just refreshed still picks it up, only then start signing with it, and withdraw the old key after the last token signed with it has passed its expiry. Publishing and signing in the same deploy is the classic self-inflicted outage.
- One subscriber never upgrades. What do you do?The window is bounded by the slowest consumer only until somebody decides it is not. Set a deprecation date from the contract's notice period, keep signing with the old key until it, and after it accept that the subscriber's pulls fail. That is a commercial decision and it should be made by the contract owner in advance, not by an engineer at cutover.
- The current signing key is suspected compromised. Does this sequence still apply?No — a compromise is the case the schedule cannot absorb. You cut over immediately and accept broken consumers, because every hour the old key stays trusted is an hour an attacker can mint. What the rehearsed sequence buys you is that accept-both is already deployed and the new key is already distributed, so the emergency is a cutover rather than a rebuild.
It is changing the locks on a building where four hundred tenants each hold a key you posted to them. You fit a cylinder that both the old and the new key open, wait until every tenant confirms the new key works in their door, and only then have the old one stop working. Announcing a date and re-keying overnight is the same act without the waiting, and it is how a hundred tenants end up on the pavement.
saying these in an interview costs you the question
- Rotate at 3 a.m.; hardly anyone is pulling copy then.
- Consumers will pick up the new verification key automatically.
- Retire the old key the moment you start signing with the new one.
- The overlap only has to cover the token lifetime.
- A stronger new key means the changeover is safe to rush.
- We do not need an inventory; we will email everyone.