With SAML metadata as your only lever, how do you move a signing key across several hundred counterparties you cannot compel?
answer
- overlap, never a swap
- publish both signing keys first
- old duration sets the clock
- freshness changes bind one generation late
- one document or one per relationship
basics
~20 sPublish the incoming key as a second KeyDescriptor while the outgoing one is still live, wait out the cacheDuration in the copies counterparties already hold, then switch and later withdraw the old one. The clock is set by the old duration, not the new.
solid answer
~60 sThe mechanism is an overlap: publish the incoming key as a second `<KeyDescriptor use="signing">` alongside the live one, let every counterparty's cached copy turn over, only then begin signing with the new key, and withdraw the old `<KeyDescriptor>` a full cycle later. The lever is freshness — `cacheDuration` bounds how stale a held copy may be and `validUntil` bounds whether it may be used at all — and the trap is that a freshness change binds **one generation late**: shortening `cacheDuration` today does nothing to copies fetched yesterday under the old value. The judgment call is structural. One document for all counterparties is cheap to maintain and makes the change all-or-nothing on a single clock; publishing a distinct document per relationship costs several hundred times the bookkeeping but lets you move in cohorts, with a per-party clock and a blast radius of one. Neither the document nor the protocol lets you force a refetch, so an announced date and a named contact per counterparty sit outside the mechanism and carry the parties the mechanism cannot reach.
code
pseudocode · 15 lines# T0 = the moment both signing keys are published
# D_old = cacheDuration carried by the copies counterparties already hold
# D_new = cacheDuration carried by the document published at T0
earliest_everyone_holds_both = T0 + D_old # NOT D_new
start_signing_with_new_key = earliest_everyone_holds_both + margin_for_slow_parties
withdraw_old_key_descriptor = start_signing_with_new_key + D_new
# guard: the document must stay usable for the whole sequence
for each document published during the rollover:
if document.validUntil <= withdraw_old_key_descriptor:
reject # the estate would hard-stop mid-rollover
# the lever you do not have
assert no_way_to_force_refetch(counterparties)go deeper
Recall the shape only: you do not swap a published key, you publish the new one next to the old one and give counterparties time to notice before you start using it.
Explain the four steps and which attribute governs each: several signing keys are allowed, the caching ceiling sets the wait, and the hard expiry must outlast the whole sequence.
Show the operational traps: the clock comes from the duration already in the field, the overlap must be sized for the slowest consumer, and the cutover is where you discover who never refetches.
Own the structural bet: one shared document versus one per relationship, justified by whether the relationships already differ, and the standing caching policy that fixes your emergency response time long before the emergency.
## What the document actually gives you A signing key change is trivial to publish and hard to land, because the party you need to reach acts on a **copy** of your document that it fetched at a time you do not know. The document gives you exactly two instruments: - **`cacheDuration`** — how long a holder may keep a copy before refetching. This is the floor on how fast any change can arrive. - **`validUntil`** — after which the content must not be used at all. This is a hard stop, not a nudge, and it is the same instant for everybody. Everything else — an announced date, a contact list, a support queue — sits outside the document. Worth being honest about that in an interview: freshness is the only lever *the document* gives you, not the only lever you have. ## The overlap play 1. **Publish both.** Add the incoming key as a second `<KeyDescriptor use="signing">`, leaving the outgoing one in place. Metadata permits several and declares no precedence between them, so a counterparty on either copy is able to accept content signed by whichever key you use. 2. **Wait out the held copies.** The earliest moment at which every counterparty can be assumed to hold both keys is the last publication plus the `cacheDuration` that was in force in the copies they already hold. 3. **Cut over.** Begin signing with the incoming key. Parties still on the old copy would fail here — hence step 2's margin. 4. **Withdraw.** Remove the outgoing `<KeyDescriptor>` in a later revision, one more cache cycle on, once nothing is signed with it. 5. **Check `validUntil` throughout.** Every document published during the sequence must carry a `validUntil` comfortably past the withdrawal date, or the entire estate hard-stops mid-rollover for reasons that have nothing to do with the key. ## The trap: freshness binds one generation late The most common way this is botched is an emergency reaction: shorten `cacheDuration` from a week to an hour and expect to move within the hour. The shorter value applies to copies fetched **after** it was published. A counterparty that fetched under `P7D` is entitled to hold that copy for a week, and nothing in the copy tells it that you have changed your mind. So the sequence is two-stage: - Publish the shorter `cacheDuration`. - Wait out the **old** duration. - Only from then on can you plan against the new one. A standing policy choice follows: the caching ceiling you run in normal times is the emergency response time you have bought, and you cannot buy it back on the day. ## The structural decision a lead actually owns | approach | maintenance cost | rollover shape | blast radius of a mistake | |---|---|---|---| | one document for all counterparties | one artefact, one renewal | all-or-nothing, on one clock | every relationship at once | | a document per relationship | several hundred artefacts and renewals | staged cohorts, per-party clocks | one relationship | In a bilateral estate both are legitimate, and the answer depends on how much the pairs already differ. If every counterparty gets the same endpoints, keys and formats, a single document is the honest representation and the cohort-by-cohort option is bought at a maintenance cost that will itself cause outages through un-renewed documents. If the relationships already differ — different `<NameIDFormat>` sets, different signing expectations — per-relationship documents are a truer description and the staged rollover comes free with them. ## The parties the mechanism cannot reach Some counterparties fetched once, at integration, and never again. Freshness attributes bind a consumer that refetches; they say nothing to one that does not. You cannot detect these from the document side, and the pairwise model gives you no authority over them. What that implies for the plan: - Size the overlap for the **slowest** party you know of, not the average. - Keep the overlap window open longer than the arithmetic requires, because the arithmetic only covers well-behaved consumers. - Treat the cutover as the moment you learn who they are: the failures after step 3 are the inventory you could not take beforehand. - Hold a named contact per counterparty, because for those parties the change is a conversation, not a document. ## What good looks like when you answer this A strong answer states the mechanism (overlap, not swap), names the clock (`cacheDuration` in the copies already held, not the one just published), bounds it (`validUntil` must outlast the whole sequence), and then says which of the two structural options it would choose **and why that choice depends on whether the relationships already differ**. A weak answer describes rotating a key and assumes everybody is looking.
- Why is shortening `cacheDuration` a poor first move once a rollover is already urgent?Because the shorter value only reaches copies fetched after it is published, so you still have to wait out the old duration before it means anything. It is a good move made in advance and a wasted one made under pressure — the response time was fixed by the value you were already running.
- What does publishing a separate document per counterparty buy, and what does it cost?It buys per-party clocks: cohorts can be moved independently and a mistake reaches one relationship. It costs several hundred artefacts to keep current, each with its own `validUntil` to renew — and un-renewed documents are themselves a leading cause of outages, so the cure can exceed the disease in a uniform estate.
- How do you find counterparties that never refetch your document?Not from the document. Freshness attributes only bind consumers that refresh, and the pairwise model gives you no visibility into their side. In practice the cutover is the discovery mechanism, which is the argument for a long overlap window and a named contact per counterparty held beforehand.
- Why must the overlap be sized for the slowest party rather than the stated duration?Because the duration describes well-behaved consumers only. A party that fetches on a slower schedule than it is entitled to, or that refetches only on restart, is invisible to the arithmetic. Adding margin costs you a longer window with two acceptable signing keys; omitting it costs that party a total outage.
saying these in an interview costs you the question
- Swaps the key in place instead of publishing an overlap
- Counts the new cacheDuration rather than the one already cached
- Believes a shortened cacheDuration takes effect immediately
- Forgets validUntil can expire mid-rollover and stop everything
- Assumes every counterparty refetches the document at all
- Treats the rollover as reversible once the old key is withdrawn