Making every outward call depend on a platform token issuer adds a failure domain — how do you design for it?
answer
- minting is now on the critical path
- held tokens keep working; new starts fail
- the outage lands one lifetime in
- publish verification keys before signing with them
- break-glass is a permanent guarding cost
basics
~20 sDesign around the one property that makes it survivable: held tokens stay valid to their own expiry, so an issuer outage degrades over one lifetime and hits newly started workloads first. Token lifetime is the lever, traded directly against a stolen token's window.
solid answer
~50 sStart by naming the shape of the failure, because it is not an outage in the usual sense. When minting stops, every workload holding a valid token keeps working until that token expires, while workloads starting during the outage fail immediately because they never received one. So the blast radius is roughly *one token lifetime*, and lifetime becomes an explicit trade: shorter means a smaller window for a stolen token and less tolerance for an issuer blip; longer is the reverse. Then decide three things you now own: refresh ahead of expiry rather than on rejection, so a short blip is absorbed; how verification keys are cached and rotated with overlap, since verifiers usually survive longer than minting does; and whether any path keeps a static credential for break-glass, and what guarding it permanently costs.
go deeper
Recall that an issued token keeps working until it expires, so the component that mints tokens being unavailable does not stop calls immediately — it stops new workloads getting their first token.
Explain the ordering: held tokens stay valid, starting workloads fail first, and verifiers survive longer because they cache keys. Then say why token lifetime decides how long that grace lasts.
Show the operational design: refresh ahead of expiry rather than on rejection, hold rollouts while the issuer is unhealthy, rotate verification keys with overlap, and tolerate a little clock skew deliberately.
Own the trade itself. Pick lifetimes per class of workload with the stolen-token window on one side and outage tolerance on the other, decide whether any break-glass credential exists, and engineer the issuer's availability to match what now depends on it.
## What moved onto the critical path Static keys have one genuine virtue: they keep working while everything else is on fire, because nothing has to be alive for a workload to present one. Replacing them with issued identity trades that away. Minting is now a live dependency of starting a workload that calls outward, and verification is a live dependency of the far side accepting the call. That is a good trade, but it is a trade, and a lead is expected to have costed it rather than to have discovered it during the first incident. ## The shape of the outage This failure does not arrive all at once, and understanding the order is most of the answer. - **Workloads already holding a valid token keep working** until that token expires. Nothing checks with the issuer per call. - **Workloads starting during the outage fail immediately**, because they never received a first token. A rollout, an autoscaling event or a node replacement in the middle of an issuer outage is where the pain actually lands. - **Verification usually survives longer than minting**, because verifiers cache the issuer's published verification keys rather than fetching them per call. - **Full outward failure arrives about one token lifetime in**, not at minute zero. So with a ten-minute lifetime, a three-minute issuer outage is mostly invisible to steady-state traffic and highly visible to anything trying to start. That asymmetry is worth designing around directly: hold rollouts while the issuer is unhealthy rather than discovering it replica by replica. ## The lifetime lever | token lifetime | tolerance for an issuer outage | window for a stolen token | mint and refresh load | |---|---|---|---| | very short (a minute or two) | almost none; steady traffic breaks quickly | minimal | high, and refresh failures become routine | | moderate (five to fifteen minutes) | absorbs a short blip | small | modest | | long (hours) | survives a real outage | large, approaching a static key | low | There is no correct row. The judgment is which risk your estate actually carries: a mature issuer with its own redundancy argues for shorter, a single-region issuer in front of a latency-sensitive estate argues for longer, and the answer may legitimately differ for the workloads reaching your crown-jewel store and the ones reaching an internal report. ## The two failure modes people meet second 1. **Key rotation without overlap.** Verifiers cache verification keys. If the issuer signs with a new key before that key has been published long enough for every cache to pick it up, valid tokens are rejected everywhere at once — an outage with no failing component. Publish first, sign later, and retire the old key only after the longest cache lifetime has passed. 2. **Clock skew.** Expiry is a comparison between two clocks that belong to different machines. Skew across an estate turns into rejections nobody can reproduce, usually at the most heavily loaded hour. Allow a small tolerance, and monitor drift as an estate-level signal rather than trusting it. ## Break-glass, and what it costs The uncomfortable question is whether any path keeps a static credential for the case where the issuer cannot mint at all — typically restore tooling or the incident path itself. Both answers are defensible, and the cost is what decides it: - keeping one means a long-lived credential that must be stored offline, audited, rotated on a schedule and tested, forever, and it will be the single most attractive target in the estate; - keeping none means accepting that if the issuer is unrecoverable, the recovery path cannot authenticate either. If you keep one, keep exactly one, scope it to the recovery path only, and make using it loud enough that nobody does so casually. If you do not, then the issuer's own availability has to be engineered to the level of the things that depend on it, which is the honest consequence. ## What to measure 1. Mint failure rate and mint latency, per boundary — the leading indicator, because it fires before any call fails. 2. Verification rejections split by reason: expired, wrong audience, unknown key. Each points at a different fault, and the third is the rotation trap. 3. The distribution of remaining lifetime at the moment of use, which tells you whether refresh is genuinely happening ahead of expiry or at rejection. 4. Start failures attributable to identity, kept separate from ordinary start failures, so an issuer problem is not diagnosed as a bad release.
- Which part of this fails first when the issuer is unreachable, and why does that matter?Newly starting workloads, because they have no token at all, while steady-state traffic runs on tokens already held. It matters because rollouts, autoscaling and node replacement are exactly what you might otherwise attempt during the incident, and each one converts a survivable outage into an outage.
- Why does key rotation cause an outage with no failing component?Verifiers cache the issuer's verification keys. Sign with a new key before the caches hold it and every verifier rejects valid tokens at once, while the issuer, the network and the workloads all look healthy. Publish the new key, wait out the longest cache lifetime, then start signing with it.
- How would you argue for keeping a static credential on one path only?By scoping it to recovery, where the alternative is being unable to authenticate to fix the issuer itself. The case holds only if it is stored offline, granted the narrowest possible access, rotated and tested on a schedule, and its use raises an alert. Without those, it is the thing you spent this migration removing.
A pass desk that reissues everyone's pass every ten minutes: close the desk and the building keeps working until passes run out, but nobody arriving can get in at all.
saying these in an interview costs you the question
- Assumes every outward call fails the moment the issuer goes down
- Thinks verifiers contact the issuer on every request
- Rotates signing keys without an overlap window
- Refreshes tokens only after a call has been rejected
- Treats break-glass credentials as free once written down
- Ignores clock skew as a source of unexplained rejections