skip to content

Twelve services share one HS256 secret to verify access tokens — what is the risk, and how do you migrate?

level: principalimportance: should knowfreq 45%

answer

  1. Verification and issuance are the same capability
  2. Count the copies, then count the trust
  3. No way to attribute a forged token
  4. Accept before you issue, retire after expiry
  5. Bind each algorithm to its own key

basics

~20 s

Every one of those twelve services can forge tokens for the whole system, and no one can tell which copy of the secret leaked. Migrate to asymmetric signing in phases: publish a JWKS, teach verifiers to accept both algorithms by kid, switch issuance, then retire the secret.

solid answer

~50 s

The risk is that verification capability and issuance capability are the same thing under `HS256`. The secret sits in twelve deployment configurations, the CI system, the secret store, and probably a few laptops — and compromise of the weakest copy yields the power to mint a token asserting any subject or role. There is also no attribution: a forged token cannot be traced to which holder produced it, so an incident has no starting point. The migration is a rotation with an extra dimension. Stand up an asymmetric key pair at the issuer and publish the public key in a JWKS with its own `kid`. Roll every verifier to a version that resolves keys by `kid` and accepts an explicit allowlist of `{HS256, RS256}` during the window, binding each algorithm to its own key rather than letting the header choose. Then flip issuance to `RS256`. After the maximum token lifetime, drop `HS256` from every allowlist and destroy the secret.

code

text · 7 lines
text
phase 1  issuer publishes RS256 public key in JWKS (new kid)
         issuance unchanged: still HS256
phase 2  every verifier accepts {HS256 -> shared secret,
                                 RS256 -> key resolved by kid}
phase 3  issuer switches signing to RS256
phase 4  after max token lifetime + skew:
         drop HS256 from every allowlist, destroy the secret

go deeper

for a junior

Recall that a shared HMAC secret lets every verifier create tokens too, which is why systems with many verifiers use a private signing key and a published public key.

for a middle

Explain the phased sequence — publish, dual-accept, switch, retire — and why verifiers must be able to accept the new algorithm before any token uses it.

for a senior

Drive the rollout: inventory every secret holder, keep the algorithm bound to a specific key during the dual-accept window, and use per-service verification telemetry to prove each phase landed.

for a principal

Own the trust architecture and the risk narrative — blast radius, absent attribution, key custody in a KMS, how token lifetime bounds the transition window, and what this migration explicitly does not fix.

## Naming the risk precisely "Sharing a secret is bad" is not an answer. Three concrete properties are lost. **Verification implies issuance.** HMAC uses one key for both directions. A service that can check a token can create one. So twelve services each hold the authority of the identity provider, and any of them — or anyone who compromises one, or reads it out of a CI log, or finds it in a stale environment file — can mint a token claiming to be any user with any role. The authorization server's carefully guarded issuance path becomes irrelevant, because there are twelve unguarded ones beside it. **Blast radius is the union of every copy.** Security is set by the weakest holder: the least-maintained service, the noisiest log, the widest-scoped CI variable. This is the classic argument for a key hierarchy — capability should decrease as you move outward from the issuer, and here it does not decrease at all. **No attribution.** If a forged token surfaces, nothing distinguishes tokens minted by the issuer from tokens minted by service seven. You cannot scope the incident, so you must assume the whole secret is burned and rotate everything. A fourth, operational, property: **rotation is a flag day**. Changing the secret means twelve deployments must land close together, which is why in practice these secrets are never rotated and end up years old. ## The target state One asymmetric key pair at the issuer, private half in a KMS or HSM so it is never exported and every signing operation is auditable. Public half published in a JWKS at a stable HTTPS URL. Verifiers hold no secret at all — they fetch, cache, and verify. Onboarding a thirteenth service becomes a configuration entry rather than a secret hand-off, and compromising any verifier yields no forging power. ## The migration **Phase 0 — inventory.** Find every holder of the secret, including the ones not on the architecture diagram: batch jobs, admin tools, test fixtures, a partner integration. The migration is only finished when the last one is converted, so an incomplete inventory means the secret survives. **Phase 1 — publish.** Create the key pair, publish the public key in the JWKS with a distinct `kid`. Issuance is unchanged; nothing breaks. **Phase 2 — dual-accept.** Roll every verifier to a build that resolves the key from the token's `kid` and accepts an explicit allowlist of algorithms, `{HS256, RS256}`, with each algorithm bound to a specific key: the shared secret for `HS256` and the fetched public key for `RS256`. This is the delicate phase. The rule that keeps it safe is that the token never selects the key material — the verifier resolves a key first and then checks that the algorithm the token declares is the one that key is permitted for. A verifier that says "take the header's algorithm and use whatever key I have" during a dual-accept window is exactly the misconfiguration that key-substitution attacks depend on. Verify this phase is complete everywhere, with telemetry, before proceeding. **Phase 3 — switch issuance.** The issuer signs with the private key and stamps the new `kid`. Existing `HS256` tokens continue to verify until they expire; new ones are `RS256`. Watch verification failure rates per service — a spike identifies a verifier that never got phase 2. **Phase 4 — retire.** After the maximum token lifetime plus skew, remove `HS256` from every verifier's allowlist, delete the secret from every configuration and secret store, and revoke it wherever it is referenced. Confirm by deployment, not by intention: a secret that is merely unused is still a secret that leaked. ## The judgment an interviewer is testing Several things beyond the mechanics. **Shortening the window.** The dual-accept phase is when the system is at its most permissive, so its length is set by the token lifetime. If tokens live fifteen minutes, phase 4 follows phase 3 almost immediately; if they live a day, you carry the widened configuration for a day. That is a concrete reason short lifetimes are worth having before you start. **Sequencing.** Verifiers must accept the new algorithm before the issuer emits it, exactly as in an ordinary key rotation — publish and accept first, switch second, retire third. Reversing any pair of steps causes an outage. **Scope discipline.** The temptation is to bundle this with adopting a new identity provider, changing claim shapes, or introducing audience restrictions. Each of those is defensible on its own; together they make a failed rollout impossible to diagnose. Migrate the signing topology alone, then take the next thing. **Knowing what it does not fix.** Asymmetric signing changes who can *create* tokens. It does nothing about a stolen token, nothing about revocation, and nothing about over-broad claims. If the underlying worry is that tokens are too powerful or live too long, that is a separate piece of work.

  • How long should the dual-accept window last?
    Just longer than the maximum token lifetime plus clock-skew tolerance after issuance switches — that is the point at which no unexpired token is still signed with the old secret. It is the most permissive configuration the system will ever run, so it should be measured in token lifetimes, not in sprints.
  • What tells you the migration is actually finished?
    The secret is absent from every deployment configuration, secret store, CI variable, and test fixture, verified by inspection rather than by assumption, and `HS256` has been removed from every verifier's accepted-algorithm list and deployed. A secret that is merely unused still exists, and an allowlist entry that is merely unexercised still accepts tokens.
  • What problem does this migration not solve?
    Anything about tokens that already exist. A stolen token is equally usable under either algorithm, revocation is still unavailable for a self-contained token, and over-broad claims are just as over-broad. The migration changes who can create tokens; if the real worry is what a valid token can do, that needs scope, audience, and lifetime work instead.

saying these in an interview costs you the question

  • Switches the issuer before verifiers accept the new algorithm
  • Accepts whatever algorithm the token header declares
  • Retires the secret before existing tokens expire
  • Misses secret holders outside the main service list
  • Bundles the migration with unrelated claim changes

context