A contractor firm rotates its assertion-signing certificate without telling you — how does your service provider survive that?
answer
- one key means a flag day
- an accepted set, not a certificate
- refresh on a schedule, keep the last good
- the metadata URL is trust configuration
- alarm per connection, not globally
basics
~20 sHold a set of accepted signing certificates per connection rather than one, refresh each counterparty's metadata on a schedule and honour its cacheDuration and validUntil, keep the last good copy when a fetch fails, and alarm per connection so one firm's rotation is visible before its people are.
solid answer
~50 sRotation is only an outage if the connection can hold exactly one key. I store an accepted-certificate **set** per counterparty connection, so a firm can publish a new one, run both for a while, and retire the old without a flag day I have to coordinate. I refresh each firm's metadata document from its published URL on a schedule, honour `cacheDuration` and `validUntil` rather than caching forever, and treat a failed fetch as keep-the-last-good-copy-and-alarm rather than drop-trust. That makes the URL part of my trust configuration, so I verify the host's transport certificate, alarm when the published key set changes, and refuse to apply a document whose `validUntil` has passed. If instead I pin a downloaded copy, I am safer against that URL changing but I am a manual step behind every rotation — and an unannounced one takes that firm's sign-ins down all at once.
code
json · 15 lines{
"connectionId": "firm-northbank-drainage",
"issuerEntityId": "https://idp.northbank-drainage.example/entity",
"metadataUrl": "https://idp.northbank-drainage.example/metadata",
"metadataRefresh": {
"honourCacheDuration": true,
"minIntervalMinutes": 60,
"lastSuccessfulFetch": "2026-09-19T02:14:07Z",
"onFetchFailure": "keep-last-good-and-alarm"
},
"acceptedSigningCerts": [
{ "fingerprint": "9f:1c:...:04", "firstSeen": "2024-11-02", "lastVerifiedAt": "2026-09-18T21:40:11Z", "retireAfter": "2026-10-15" },
{ "fingerprint": "b3:77:...:e9", "firstSeen": "2026-09-19", "lastVerifiedAt": "2026-09-19T02:19:55Z", "retireAfter": null }
]
}go deeper
Understand that the key a counterparty signs with will change over time, and that your side has to be told about it somehow — by a document you fetch or by a person sending you one.
Explain why holding more than one accepted certificate per connection removes the coordinated cutover, and what a scheduled metadata refresh has to honour when it runs.
Show the operational design: refresh policy, keep-last-good on failure, alarms on key-set change, per-connection signature-failure metrics, and deliberate retirement of old keys after the new one has been observed.
Decide the estate-wide posture — refresh for most counterparties, pin for the few that justify a human review — and treat your own published metadata as a versioned interface whose changes every counterparty must consume on their schedule.
## The failure this is about At three in the morning a contractor firm's identity provider team replaces the key it signs with. Nobody told the water authority, because from their side it was routine. Every permit sign-in from that firm now fails signature verification, and because that firm is one tenant among many, your global error rate barely moves. The engineer on call sees *SSO is broken* from one customer and nothing in the dashboards. Everything that determines whether that is a five-minute annoyance or a four-hour outage was decided months earlier, in how the connection stores keys and how it learns about changes. ## Hold a set, not a key The single most valuable design decision here is that a connection's signing material is a **set**. A verification attempt succeeds if the signature checks against any member. - An overlap becomes possible: the firm adds a new key, signs with either for a period, and retires the old one, with no coordinated cutover. - A rollback becomes possible: if the new key turns out to be wrong, the old one is still accepted. - The removal of a retired key becomes a separate, deliberate, low-risk action, taken after you have observed the new one in use. A connection that stores one certificate turns every rotation into a synchronised deploy between two organisations, one of which does not work for you. ## Consuming metadata on a schedule, or pinning a copy There are two honest positions and they trade the same risk in opposite directions. | | Scheduled refresh from the counterparty's URL | Pinned downloaded copy | |---|---|---| | Unannounced rotation | absorbed automatically | breaks that connection's sign-ins at once | | What you trust | whatever that URL serves at refresh time | what a human reviewed on the day | | Change visibility | needs an alarm on key-set change, or it is silent | every change is a reviewed action | | Operational cost | a fetch job, its failures and its monitoring | a ticket per rotation, per counterparty | Scheduled refresh scales, and it is the right default once you have more than a handful of counterparties — but it makes that URL part of your trust configuration. So: verify the host's transport certificate strictly, refuse a document whose `validUntil` has already passed, honour `cacheDuration` instead of hammering or caching forever, and **alarm when the published key set changes** so a rotation is a notification rather than a silence. Keep the last good copy on a fetch failure; a firm's metadata endpoint being down must not log its people out. Pinning is defensible for a small number of high-value counterparties where a human review per change is affordable and the extra assurance is worth the manual step. ## Your own metadata is a thing you publish and version The other half of this is the document **you** publish: an `EntityDescriptor` for your service carrying an `SPSSODescriptor`, your `entityID`, one or more `AssertionConsumerServiceURL` endpoints with their `index` and `isDefault`, a `KeyDescriptor` for any key you sign or decrypt with, and the logout endpoint you accept messages on. Counterparties consume it to configure their side. Two rules that only hurt when they are broken: 1. **A different `entityID` per environment.** Test and production must not share one. If they do, a firm configured against your test entity can have its messages accepted by production, and the mistake is invisible until it matters. 2. **Version it and publish it as an artefact.** Adding an endpoint, changing a URL or adding a key is a change every counterparty has to consume, on their schedule, not yours. Treat it as a published interface with a change log, and give yourself the same overlap you asked them for — add the new endpoint before retiring the old one, and add a new key of your own alongside the current one rather than replacing it. ## Detection and the per-connection view The operational requirement that falls out of all of this is per-connection observability: signature-failure counts, last successful metadata refresh, current accepted key fingerprints and their first-seen dates, and the age of the oldest key still in the set. That last one is the metric that stops the accepted set from silently becoming a permanent collection of everything a firm has ever used. Put a review date on each entry and remove retired keys deliberately, once the new one has been observed in real traffic. ## What you cannot do You cannot make a counterparty announce a rotation, and you should not build a process that assumes they will. You can make the unannounced case cheap — a set, a refresh, an alarm — and you can make the announced case pleasant. Everything else is a phone call you do not control.
- Why not simply trust any certificate the refreshed metadata document contains, and drop the accepted set?That is effectively what a refresh does, but the set is what carries history and intent: which keys you have actually seen verify, when each appeared, and when a retired one may be removed. Without it you cannot answer whether a key in the document is new, whether the old one is still needed, or what changed last night.
- What do you publish in your own service-provider metadata, and why does the environment matter?Your `entityID`, an `SPSSODescriptor` with your `AssertionConsumerServiceURL` endpoints and their `index` and `isDefault`, a `KeyDescriptor` for keys you sign or decrypt with, and your logout endpoint. Environments need distinct `entityID` values so a counterparty configured against test can never have its messages accepted in production.
- A firm's metadata endpoint is unreachable for two hours. What should your refresh job do?Keep the last good copy, keep verifying against the accepted set, and raise an alarm naming that connection. Dropping trust because a fetch failed converts someone else's outage into your own, and the previously fetched material is not less valid because the host is down.
saying these in an interview costs you the question
- Store one signing certificate per counterparty and swap it on rotation day.
- If the metadata fetch fails, treat the connection as untrusted.
- Fetch the metadata once at onboarding and cache it forever.
- Share one entityID between the test and production deployments.
- A global signature-failure rate is enough to notice a rotation.
- Delete the old certificate as soon as the new one appears in the document.