A long-lived platform access key is found in a public repository — why is issuing a replacement key not the first move?
answer
- order of operations matters here
- stopping it beats replacing it
- creation is not cancellation
- count from publication, not discovery
- read back what it called
basics
~20 sCreating a new key does not disable the old one. Revoke the exposed credential first so it stops authenticating immediately, then issue replacements, then read back what the leaked credential actually called before it was stopped.
solid answer
~40 sA static platform key has no end date, so it stops working for one reason only: someone revoked it. Platforms normally let one identity hold several active keys at a time, which is exactly why rotating first is dangerous — you add a working credential without withdrawing the exposed one, and while you edit config and redeploy, the attacker still has a credential that authenticates. So the order is contain, confirm, replace, reconstruct: deactivate the exposed key, prove a call presenting it is refused, issue new credentials for the legitimate consumers, and only then read back what that credential did, starting from the moment it first became reachable rather than the moment someone noticed. Published credentials are harvested by automation and first use is routinely minutes later, so treat it as already used.
go deeper
Remember the order: stop the exposed credential first, replace it second. A static credential works until it is revoked, and creating another one changes nothing about the first.
Explain why platforms allow two active keys at once, and how that convenience turns rotate-first into a window where the attacker still has a working credential while you redeploy.
Show the investigation: an exposure window that starts at publication, a search of recorded use for calls your systems cannot account for, and an honest statement of what an empty result does and does not prove.
Frame the incident as evidence about the estate, not about one key. The question a lead answers is which classes of caller still require a value a human can paste, and what it would take to retire that requirement.
## The credential in question A **long-lived platform access key** is a pair of values — a public identifier and a secret — that a cloud platform accepts as proof of a principal's identity on every call it is presented with, with no end date. It is a **bearer** credential: the platform authenticates the string, not the person or the machine holding it, so a copy is as good as the original. Nothing in an incoming call distinguishes your scheduler from someone who read the value out of a public repository ten minutes ago. That single property drives the whole response. A credential that never expires stops working for one reason: somebody revoked it. ## Why a replacement is not a response Platforms commonly allow one identity to hold more than one active key, precisely so a changeover can overlap — the new credential starts working before the old one is withdrawn. That convenience is what makes *rotate first* dangerous during an exposure: - creating a key **adds** a working credential; it does not withdraw one; - while you edit configuration, rebuild and redeploy, the exposed key authenticates every call it is presented with; - publicly posted credentials are harvested by automation, and the gap between publication and first use is routinely measured in minutes; - the person redeploying is usually the person who should be containing, so rotating first also costs you the attention the containment step needed. **Revocation is the containment action; rotation is the recovery action.** They are not interchangeable, and the order is not a matter of taste. ## The drill 1. **Contain.** Deactivate or delete the exposed credential at the platform. Do this before you know the blast radius — a published credential is already somebody else's. 2. **Confirm.** Make one call with the old credential and watch it be refused. Revoking something you believed was the credential is a common and silent mistake. 3. **Replace.** Issue credentials for each legitimate consumer, preferably not another static key. 4. **Reconstruct.** Read back what that credential called, across the whole window it could have been in someone else's hands. 5. **Remove the cause.** A job that needed a pasted static key should be taking on a session that ends by itself. ## Deactivate or delete | action | effect on calls | reversible | effect on investigation | |---|---|---|---| | deactivate / disable | refused immediately | yes — which is itself a risk | the identifier still resolves, so recorded calls stay attributable | | delete | refused immediately | no | the identifier may be harder to attribute afterwards | Deactivating first and deleting once the investigation is finished is the usual sequence. The reason to keep the identifier is **attribution**, not the option of switching a publicly known credential back on. ## Sizing the window The exposure window opens when the value first became reachable by someone outside the team, not when someone noticed it: - the commit timestamp — including a commit later removed, because history, clones and forks keep it; - the publication time of any build image the value was baked into; - the timestamp of the job log, ticket or message that carried it. Then read the recorded use across that whole window and look for calls your own systems cannot account for: actions the job never performs, calls from places you do not operate in, and attempts that were refused for lack of permission — a refused call is still evidence that someone was probing. Finding nothing unexplained is weak evidence rather than proof: your records may not cover every service, and a caller who only read data leaves a thin trail. ## Why this is an argument about design Every step above exists because the credential had no natural end. A credential issued for one task and expiring on its own inverts the economics: the same exposure stops mattering when the window closes, with nobody noticing, revoking or redeploying. **Shortening a credential's life is the cheapest lever on blast radius** because it costs nothing at rest and needs no detection to work. It is not a substitute for narrowing what the credential may do — a short-lived credential with broad permissions does broad damage inside its window — and it does not retire this drill, because the identity that issues the short-lived credential is itself long-lived and can be exposed in its own way. What it does change is how often you run the drill, and whether a missed exposure becomes an incident or a non-event.
- Which timestamp should the investigation start from?The earliest moment the value was reachable by anyone outside the team — the commit that introduced it, the publication of a build image containing it, or the log line that printed it. Discovery time is when you started looking, not when the exposure began, and using it routinely understates the window by days.
- What if revoking the exposed key will break production jobs?Revoke anyway when the value is known to be public; it is already available to anyone who wants it, so the outage is a cost you have already incurred. Prepare replacements in parallel rather than in sequence, and treat the fact that nobody could enumerate the affected consumers as the finding to fix afterwards.
- What stops the same incident happening again?Removing the need for a value that a human can paste. Where the caller takes on a session for one task and that session expires by itself, an exposed credential ages out without anyone acting. Static keys remain only where a consumer genuinely cannot obtain a session, and those become a short, owned list.
Handing out a spare front-door key is nothing like issuing a visitor badge that stops working at six. Cutting a second spare does not make the first one stop opening the door — you have to change the lock.
saying these in an interview costs you the question
- Says issuing a new key automatically invalidates the leaked one
- Believes rewriting history to drop the commit contains the leak
- Counts the exposure window from the moment someone reported it
- Assumes nobody found the key because nothing looks broken
- Thinks narrowing the key's permissions removes the need to revoke