skip to content

Every workload authenticating through an outside issuer began failing verification at 09:00 and recovered an hour later with no change on your side — what happened?

level: seniorimportance: must knowfreq 55%

answer

  1. nothing changed on your side
  2. it healed without anyone acting
  3. one issuer, all callers, at once
  4. how old is the copy you verify against
  5. refetch when the signing key is unknown

basics

~20 s

The issuer rolled the key it signs with, and the store was verifying against an older copy of that issuer's published keys. Every document signed with the new key failed until the store's scheduled refill of that copy caught up.

solid answer

~40 s

The shape of the symptom is the diagnosis: all callers of one issuer, at once, with nothing changed locally, and a recovery nobody performed. That points at the one thing shared by every caller of that issuer and owned by neither side's code — the store's copy of the keys the issuer publishes. The issuer rolled over; the new signing key was not in the copy; every document signed with it was rejected as unknown. An hour later the store's periodic refill fetched the current set and the rejections stopped. Two decisions made the window that long: the store refreshed on a fixed interval and did not refetch when an unrecognised signing key arrived, and the issuer did not keep the previous key published long enough to cover that interval.

code

pseudocode · 13 lines
pseudocode
on verify(document):
    keyId = document.signedWithKeyId
    key   = keyCopy.lookup(keyId)

    if key == null:
        return reject("unknown signing key")   # the whole outage lives here

    if not signatureValid(document, key):
        return reject("bad signature")

    return accept(document.issuer, document.subject)

# keyCopy is only ever refilled by refillOnTimer(everyHour)

go deeper

for a junior

Recall that a store verifying outside proof keeps a copy of the issuer's signing keys, and that a copy can be out of date.

for a middle

Explain how a key rollover produces an all-at-once verification failure, and what ends it when nobody intervenes.

for a senior

Diagnose from the symptom's shape, then separate the publisher's duty from the verifier's and harden the half you control.

for a principal

Decide how much of the estate may depend on an outside issuer's rollover discipline, and what notice and ownership you require before that dependency is accepted.

## What the symptom already tells you Before looking at anything, four facts narrow this almost to a single cause. - **All callers of one issuer, none of the others.** Whatever broke is shared by those callers and specific to that issuer. - **Simultaneous.** It is not a per-caller lifetime running out, which would stagger. - **No change on your side.** So the change was on theirs, or in something of theirs that you copy. - **It healed with no action.** Something on your side re-read state from theirs on a schedule. The only object that fits all four is the store's local copy of the keys that issuer publishes. ## What actually happened A store verifying documents it did not mint keeps a copy of the issuer's current signing keys, fetched from wherever that issuer publishes them, and consults it per request. At 09:00 the issuer began signing with a new key. The store looked up the key identifier carried with the signature, found nothing in its copy, and rejected — correctly, from its point of view: an unrecognised signing key is exactly what a forgery looks like. Every caller of that issuer hit the same wall in the same second, because they all present documents signed by the same key. At 10:00 the store's periodic refill fetched the published set again, the new key landed in the copy, and verification resumed. Nothing was fixed; the copy caught up. ## Why the window was exactly that long The outage length was set by the gap between the rollover and the next scheduled refill — which means the refresh interval is the worst-case outage for every future rollover, not a tuning detail. Two separate decisions produced it: 1. **The store refreshed only on a timer.** A verifier that refetches when it meets a signing key it has never seen would have recovered in one request instead of an hour. 2. **The issuer published the new key without keeping the old one available.** A publisher that carries both keys across a window longer than every consumer's refresh interval produces no outage at all, because consumers that have not refreshed yet still hold a key that validates. Either one alone would have prevented it. That is worth saying out loud in an answer, because it identifies which half you can actually change. ## The two sides of a rollover | side | what it owes | what it costs when skipped | |---|---|---| | the issuer, publishing | both keys available across a window wider than consumers' refresh intervals, and notice before the change | every consumer fails for as long as its own refresh interval | | the store, verifying | a refetch when an unknown signing key appears, bounded; alerting on the fetch itself failing | an outage of the full refresh interval, repeated at every rollover | You usually control only the second row. Design as if the first row will not happen, and treat it as a courtesy when it does. ## Making the failure structural rather than shorter - Refetch on an unknown key identifier, but **rate-limit it hard** — a flood of documents carrying invented key identifiers must not become a flood of requests at the issuer, and unbounded refetching converts a verification problem into an availability problem for both sides. - Treat the fetch of the published keys as a dependency with its own monitoring: alert when it last succeeded too long ago, not only when it errors, because a fetch that quietly keeps returning a cached response looks healthy. - Keep the previously seen keys until they stop appearing, so a rollover in the other direction does not reject documents already in flight. - Alert on the sharp shape of this failure — a sudden all-or-nothing rejection rate for one issuer — which is distinguishable from the diffuse background of bad callers. - Agree notice with the issuer's owners where you can, and record who they are; a trust relationship with no contact is one that fails like this again. ## What to check while it is still happening Ask, in order: is it one issuer or all of them; what does the rejection reason actually say (unknown signing key is different from bad signature, which is different from stale document); when did the store last successfully fetch that issuer's published keys; and does the copy contain the key identifier the current documents carry. Those four answers land on this cause or rule it out in a few minutes. ## The mistake to avoid afterwards Recovery is not a fix. The same rollover will happen again, and the next one may land at a worse hour or coincide with a fleet restart, when every workload is authenticating at once rather than renewing calmly. Treating the self-healing as closure is how this recurs annually.

  • Why did it recover without anyone doing anything?
    Because the store refills its copy of the issuer's published keys on a timer. The outage lasted the gap between the rollover and the next refill; nothing was repaired, the copy simply caught up. That interval is therefore your worst-case outage for every rollover to come.
  • What should a store do when a document arrives signed by a key identifier it has never seen?
    Refetch the issuer's published keys once and re-decide — with a hard rate limit, so documents carrying invented key identifiers cannot turn into a request flood at the issuer. Bounded that way, an hour of failures becomes one slightly slower request.
  • Whose side of a rollover is at fault here?
    Both, differently. A publisher should keep the old and new keys available across a window wider than any consumer's refresh interval. A consumer should not rely on that courtesy. Say which of the two you can actually change — usually only your own.

A gate desk that checks signatures against a specimen card photocopied last month. When head office issues new cards, every genuine signature looks like a forgery until the desk's photocopy is replaced.

saying these in an interview costs you the question

  • Blames clock skew when every caller of one issuer fails at once.
  • Assumes the issuer must have been unreachable.
  • Thinks a cached copy of published keys is harmless because keys rarely change.
  • Refetches the published keys on every failed verification, unbounded.
  • Says the fix is to pin one signing key permanently.
  • Treats the self-recovery as proof the problem is resolved.