Which credentials do you rotate first after a malicious build plugin ran on your CI runners?
answer
- scope by what was readable
- re-entry and shipping power first
- signing keys before you rebuild
- rank data-plane creds by asset
- rotation does not undo past use
basics
~20 sAnything that lets an attacker back in or ship code: deploy roles, publish and registry tokens, signing keys. Then data-plane credentials ranked by what they reach. Rotation alone is not remediation — also audit what the stolen credentials already did.
solid answer
~50 sScope first, then order. Scope is everything the build process could read: environment variables injected into the job, files in the workspace, credential files in home directories on developer machines, and any ambient cloud role or token the runner could assume. Order by power to re-enter or to ship code — a stolen deploy role or publish token can re-poison everything you have just cleaned, so it goes first; signing credentials next, because they let an attacker produce artifacts your own verification would accept; then registry read tokens and data-plane credentials ranked by asset, so on a payments platform the ledger-writing credential outranks a metrics token. Short-lived per-job tokens shorten the rotation list but not the investigation: within their lifetime they had full authority, so go and read what they did. Rotating a credential never undoes an action taken with it.
go deeper
Know that anything the build could read — injected environment variables, files in the workspace, credential files on the machine — is treated as exposed, and that rotating means issuing new credentials and invalidating the old ones.
Be ready to explain scope and ordering: enumerate what was readable, then rotate deploy, publish and signing credentials before lower-value ones, and say why the ambient role a runner assumes counts even though no pipeline file names it.
Show the two judgments that matter: re-entry power first so you are not cleaning behind an attacker who still has access, and a post-rotation sweep for what the stolen credentials already did — published artifacts, new keys, added automation.
Own how completeness is proven and how long it takes. If nobody can enumerate what a job could read, or if a single credential cannot be rotated inside a day, those are design commitments to fund before the next incident rather than facts to discover during one.
## The setting A payments platform pulls a malicious version of a build plugin. Its code ran wherever a build ran — every developer laptop and every CI runner. The response opens from the worst credible assumption: the CI deploy role, the artifact-registry token and the signing credential are already gone. The question is what you rotate, in what order, and what rotation does not buy you. ## Step 1: scope by readability, not by memory The list of credentials to rotate is not the list you can recall. It is an enumeration of what was *readable from those environments*: - Secrets injected into the job as environment variables, including ones injected for steps that did not run. - Files in the workspace: config files, service-account material, anything a previous step wrote to disk. - The ambient identity of the host — a cloud role the runner assumes, or a token available to any process on the machine. - On laptops, the far broader set: personal access tokens with organisation scope, SSH keys, cloud CLI profiles, credentials for unrelated projects that happen to live in the same home directory. Anything you cannot demonstrate was unreadable is rotated. "Probably not exposed" is not a category. ## Step 2: order by what the credential can do to you next 1. **Re-entry and shipping power first.** Deploy roles, publish tokens, registry write credentials, CI administration tokens. If any of these remain valid, everything you clean can be re-poisoned, and you will chase your own tail. This is the one ordering mistake that actually costs you the incident. 2. **Signing credentials next.** A stolen signing key is worse than a stolen deploy token in one respect: it produces artifacts that your own verification, and possibly your consumers', will accept as genuine. Revoke before you rebuild, or you will re-sign into a compromised trust chain. 3. **Then data-plane credentials, ranked by asset.** Here the asset drives the order rather than the technology. On a payments platform, a credential that writes to the ledger — money and audit truth — outranks a read-only credential for a metrics store, even though both were equally readable. 4. **Then the long tail**, including third-party integration tokens that nobody owns and which are, for that reason, the ones most likely to still be valid in six months. ## Step 3: short-lived tokens do not exempt you "Our runners use per-job tokens that expired hours ago" is a good property and a bad answer. It shrinks the rotation list because there is nothing left to revoke, and it shrinks the *window*, which matters. It does not shrink the investigation: during its lifetime that token carried the job's full authority, and a payload that ran in the same job used it while it was valid. The right follow-through is to read the audit trail for what that identity did during the window rather than to declare the problem expired. ## Step 4: rotation is not remediation The single deepest point in this answer: **rotating a credential invalidates future use; it does not undo past use.** After rotation you still sweep for what was done: - Artifacts published or images pushed during the window. - New identities: keys, users, tokens, service accounts, added collaborators. - Automation added: webhooks, scheduled jobs, altered pipeline definitions. - Configuration changed: relaxed protections, added trusted publishers, modified deployment targets. Any of these outlives the credential that created it, and a response that rotated everything and swept for none of it has left the door propped open. ## Step 5: prove it, do not assert it Completeness comes from an enumeration you can show — the secret bindings each affected job had, plus an inventory of credential files on affected machines — not from a team's collective memory. Track it as a list with a state per credential, and let the incident close only when every entry is rotated, expired, or explicitly justified as unreadable. ## What a strong answer sounds like "Enumerate what those environments could read, rotate re-entry and shipping credentials first so the attacker cannot follow me around, then signing keys, then data-plane credentials by asset. Short-lived tokens change the rotation list, not the investigation. Then I go look at what those credentials already did, because rotation is forward-looking only."
- The runners mint short-lived tokens per job. Does that reduce the work?It reduces the rotation list, not the investigation. While it lived, the token had the job's full authority, and the payload ran inside that job. Read the audit trail for what that identity did during the window instead of treating expiry as containment.
- Developer laptops ran it too. What is on them that CI does not have?Usually more, not less: personal tokens with broad organisation scope, SSH keys, cloud CLI profiles, browser sessions and credentials for unrelated projects sharing the same home directory. A runner's identity is typically job-scoped; a laptop's identity is a person's.
- How do you know rotation is complete?From an enumeration, not from memory: the secret bindings each affected job had, plus an inventory of credential files on affected machines. Every entry ends as rotated, already expired, or justified as unreadable. Anything left ambiguous counts as outstanding.
- Why revoke signing credentials before rebuilding rather than after?Because rebuilding while a stolen signing key is still valid means you are re-establishing a trust chain an attacker can also produce artifacts in. Revoke, stand up the replacement, then rebuild, so everything you sign afterwards is distinguishable from anything they signed during the window.
Changing the locks stops the next visit. It does not recover what was carried out last night, and it does not close the window they left unlatched on the way.
saying these in an interview costs you the question
- Rotates only the secrets named in the pipeline config
- Treats rotation as the end of the response
- Says short-lived tokens mean no exposure
- Skips laptops because CI is the build system
- Rotates noisy credentials before deploy and publish ones
- Never checks what the stolen credential already did