Your machine caller's client secret must be replaced at the authorization server — how do you sequence that without losing acquisition?
answer
- two live at once, then remove
- the registration change is not atomic
- held tokens hide the break for an hour
- verify by evidence, not by deploy
- overlap outlasts the slowest instance
basics
~20 sRegister the new credential alongside the old one so both authenticate the same client, roll the deployed value instance by instance as an ordinary deploy, confirm from acquisition evidence that every instance now uses the new credential, then remove the old one. Never change the registration and the fleet in one step.
solid answer
~50 sThe problem is that replacing a credential is atomic at the authorization server and never atomic across a fleet, so any single-step swap creates a window where the registered value and the deployed value disagree — and in that window acquisition fails with `invalid_client`. Worse, it fails **late**: every instance keeps polling happily on the token it already holds, so the break surfaces staggered, one instance at a time, as each entry expires. The fix is an overlap: two credentials live on the same registration at once, a per-instance roll with no coordination, verification from real acquisition evidence, then removal. Two things constrain the plan. The authorization server must actually support more than one live credential per client — check before you promise zero downtime, because if it does not, the overlap has to be carried by a second registration the callee also accepts. And the overlap window must outlast the slowest instance to roll plus the longest token lifetime plus your rollback budget.
code
json · 10 lines{
"client_id": "bms-site-poller",
"token_endpoint_auth_method": "client_secret_basic",
"credentials": [
{ "id": "cred-2026-03", "created": "2026-03-01", "status": "active",
"last_used": "2026-09-19T05:30:02Z" },
{ "id": "cred-2026-09", "created": "2026-09-18", "status": "active",
"last_used": null }
]
}go deeper
Recall the shape: two credentials are valid at the same time for a while, the deployed value moves during that window, and the old one is removed only afterwards.
Explain why the failure arrives late — acquisition breaks immediately, but each instance keeps working on the token it already holds until that token expires.
Show what you verify before removal: acquisition evidence per credential across every instance, an overlap window sized from the real roll time, and a watch period after removal.
Own the constraints and the costs: whether the authorization server supports an overlap at all, what each credential placement does to the roll, and what rotation leaves uncovered when the trigger was a suspected disclosure.
## Why a straight swap is an outage with a delay fuse Replacing one value at the authorization server is instantaneous. Replacing it in a fleet of site controllers is not: instances restart at different times, one is cordoned, one is mid-deploy, one was scaled up from an older release. Between those two events every instance still holding the old value fails to acquire — the token endpoint refuses the client — while the registration has already moved on. The delay is what makes it nasty. A cached machine token is unaffected by the credential that obtained it, so polling continues normally on every instance until its held token reaches expiry. The result is a failure that appears minutes or an hour after the change, scattered across the fleet in the order their tokens happen to run out, with no obvious correlation to the deploy that caused it. Anyone diagnosing it is looking at the wrong hour. ## The overlap 1. **Confirm the capability first.** Can the registration hold two credentials at once? If yes, everything below is routine. If not, this is a flag day, and the honest options are a maintenance window or carrying the overlap on a second registration that the callee is willing to accept for the duration — which is a bigger change and needs its own agreement. 2. **Add the new credential.** Both are now live. Nothing has been removed, so nothing can break. 3. **Roll the deployed value.** An ordinary release, instance by instance, at whatever rate you normally deploy. No coordination with the authorization server is needed, which is the entire point. 4. **Verify from acquisition evidence.** Not from "the release went out". You need to see acquisitions authenticating with the new credential from every instance, and no acquisitions still using the old one. 5. **Remove the old credential**, and watch for acquisition failures for at least one full token lifetime afterwards. The overlap window must be longer than the slowest instance to roll, plus the longest token lifetime, plus however long you would need to roll back. Anything shorter and step 5 is a guess. ## Verification is the step teams skip A deploy proves a value shipped. It does not prove that every process re-read it, that no instance is running an older release, and that nothing is pinned. Two signals are usable, and you should have arranged at least one *before* step 2: - **The authorization server's own view** — a last-used timestamp or an issuance log attributable to each credential. Where it exists, this is the authoritative answer and step 5 waits for it. - **The caller's own metric** — an acquisition counter tagged with which credential identifier was presented. It costs one label and it is the difference between a planned removal and a hopeful one. ## Where the credential rests, and what that does to step 3 | placement | how the new value reaches the process | what step 3 costs | how you verify | |---|---|---|---| | injected as an environment value at deploy | only by starting a new process | a full restart of every instance, on your deploy cadence | rollout completion plus per-credential acquisition metric | | mounted file, re-read on change | the file changes underneath a running process | no restart, if the process genuinely watches or re-reads on failure | acquisition metric, since no restart marks the moment | | platform-issued machine identity | the platform renews short-lived material for the workload | no human rotation at all: the material is short-lived by construction | the platform's own issuance record | The middle row has a trap: a process that reads the file once at startup is environment-equivalent no matter where the value is stored. If you are counting on a restart-free roll, the re-read path has to be exercised, not assumed. The third row moves the trust root to the platform the workload runs on, and it is only available where that platform issues such an identity in the first place. ## What the rotation does and does not reach Removing the old credential stops it obtaining **new** tokens. It does not shorten the life of any token already issued under it: each of those keeps working until its own `expires_in` elapses, which can be an hour. If the rotation is a response to a suspected disclosure rather than routine hygiene, that gap is the exposure you have left, and closing it is the issuer's revocation path, not the rotation. ## Changing the authentication method, not just the value Migrating a registration from a shared secret to an assertion signed with a private key (`private_key_jwt`) or to a certificate-based method is the same sequencing problem with one extra step: the registration's `token_endpoint_auth_method` changes, and many authorization servers treat that as a single-valued field. Register the new key or certificate alongside the current credential, verify acquisitions succeed with the new method in a non-production registration first, flip the method, and keep the ability to flip back until the roll is observed complete. The methods themselves differ in what they send and what they are worth — that is the client-authentication subject — but the deployment problem is identical: two things must be acceptable at once, or somebody is holding an outage.
- The authorization server allows only one credential per registration. What now?Zero downtime is off the table for that registration, so stop promising it. Either take a short window when the fleet is idle between polls and accept the failures it causes, or carry the overlap on a second registration that the meter service is willing to accept for the duration, then retire the first. Both are bigger decisions than a value change, which is why the capability check comes before the plan.
- How long should the overlap last?Longer than the slowest instance takes to roll, plus the longest token lifetime, plus the time you would need to roll back — and then long enough to have observed the old credential go unused across a full polling cycle. Setting it by calendar convenience is how a cordoned instance that never rolled gets discovered by its own failure.
- You are rotating because the credential may have leaked. Does the same sequence apply?The sequence is the same, but the tolerance is not: the overlap is now a window in which the possibly-disclosed credential still works, so you shorten it and accept some failures. Note also what rotation does not reach — tokens already obtained with that credential keep working until they expire, so if the exposure matters, the issuer's revocation path has to close that gap separately.
- Migrating from a shared secret to a key-signed assertion: what changes in the plan?One extra step and one extra risk. The registration's authentication method is often single-valued, so you cannot always have both methods live the way you can have two secrets live. Register the key alongside, prove acquisition with the new method somewhere non-production, flip the method, and keep the old credential registered and the flip reversible until the roll is observed complete.
saying these in an interview costs you the question
- Replaces the secret at the server and redeploys, expecting both to land together
- Calls the rotation done because the release finished, without checking acquisitions
- Believes removing the old credential immediately kills tokens already issued from it
- Plans a zero-downtime rotation without checking two credentials can be live at once
- Sets the overlap window by the calendar rather than by the slowest instance
- Assumes a mounted secret file is re-read by a process that only reads it at startup