skip to content

A CI helper script that thousands of pipelines download at run time was tampered with - how far does that reach?

level: seniorimportance: should knowfreq 57%

answer

  1. two hops: the jobs, then their outputs
  2. credentials and the artifact, not just secrets
  3. fetched at run time, so nothing recorded it
  4. reach exceeds what you can enumerate
  5. rotate by scope, rebuild from clean inputs

basics

~20 s

Every job that fetched it ran attacker code with that job's credentials and could alter what the job published, so the reach extends to those artifacts' consumers. Worse, a run-time fetch leaves no dependency record to query.

solid answer

~50 s

Reason in two hops. First hop: every pipeline run that fetched the helper during the window executed attacker-controlled code inside a process holding registry tokens, cloud roles, signing material and the source being built - so treat every credential reachable from those jobs as exposed, and every artifact they published as suspect, since the code could alter the output after tests passed. Second hop: consumers of those artifacts inherit the doubt. The uncomfortable part is enumeration: a helper fetched at run time appears in no lockfile and no component inventory, so you reconstruct scope from hosting access logs and CI logs, both bounded by retention, and you cannot see your consumers' pipelines at all. Reach exceeds enumerability, which is why you rotate and rebuild by scope rather than by evidence, and notify downstream with the specific artifact versions published in the window.

go deeper

for a junior

Understand that a CI job holds real credentials and produces the artifact users install, so code that runs inside it can both steal secrets and change what gets published.

for a middle

Be able to trace the two hops - affected jobs, then the artifacts they published and the consumers of those - and explain why a run-time fetch leaves no dependency record to query.

for a senior

Demonstrate the operational sequence: bound the window from hosting evidence, enumerate from CI logs while naming the retention gap, rotate by scope, rebuild from clean inputs, and notify downstream with specific versions.

for a principal

Own the systemic call: what the organisation accepts as an input to a build, whether every such input must be declared and recorded so blast radius becomes queryable, and who authorises a rebuild-everything decision and its cost.

## The two hops **Hop one: the jobs themselves.** A CI job is not a sandbox; it is a machine assembled to hold exactly the things an attacker wants. During the window, every run that fetched the helper executed attacker-controlled code with access to: - credentials injected into the environment - registry and package publishing tokens, cloud roles assumed by the runner, database or deploy credentials for the environments the pipeline touches; - the checked-out source, including anything private in it; - signing material, if the pipeline signs what it publishes; - and the *output* - the code can modify the artifact after tests pass and before it is published, which is what makes this more than credential theft. Outbound network calls from a build are normal behaviour, so exfiltration blends in. **Hop two: what those jobs produced.** Every artifact published during the window carries doubt, whether or not it was actually altered. Its consumers - other teams, deployed environments, external customers - are in scope by inheritance. This is the transitive part, and it is why one small shared helper can matter more than a large application dependency: shared build machinery sits upstream of many artifacts at once. ## Reach versus enumerability The defining property of this class is that **reach is large and enumeration is weak**, and a good answer says so out loud. A declared dependency leaves a trail: it appears in a manifest and a lockfile, it lands in a component inventory for the artifact, and you can query which builds included which version. A run-time fetch leaves none of that. There is no manifest entry, nothing in the artifact's component list, and often no local copy afterwards. So the investigation moves off the dependency graph and onto operational evidence: - **Hosting side** - object version history, modification timestamps and access logs for the location the helper was served from. This is what bounds the window, and it is the single most valuable artefact of the whole investigation. - **CI side** - job logs and network egress records that show which runs fetched it. Retention is the limiting factor; if logs roll at fourteen days and the window is longer, part of your answer is permanently unknowable. - **Downstream** - you cannot see your consumers' pipelines at all. Anything they built using your published artifact is their enumeration to run, on information only you can give them. ## How to run it 1. **Stop the bleeding.** Freeze the hosting location, remove or replace the object, and block the fetch so no further job runs the tampered code. 2. **Bound the window from the hosting side**, not from when someone noticed. Use object version history and write access logs; if you cannot establish a start time, choose the earliest defensible one and say that you did. 3. **Enumerate the affected jobs** from CI logs, and record explicitly how far back your logs go - the gap is part of the finding. 4. **Rotate by scope, not by evidence.** Do not rotate only the credentials you can prove were read. Rotate everything reachable from the affected jobs, because absence of evidence in a log that the attacker could influence is not evidence of absence. Prioritise long-lived and broadly scoped credentials, and above all publishing and signing keys. 5. **Re-derive the artifacts.** Rebuild everything published in the window from clean inputs on clean runners, and compare - a rebuild that does not match is a finding; a rebuild that matches raises confidence without proving innocence. 6. **Notify downstream with specifics.** Names, versions and digests of the artifacts published during the window, the window itself, and what a consumer should do. A vague advisory transfers your uncertainty without transferring the ability to act on it. 7. **Close the class.** The remediation that matters is not deleting one bad file; it is making run-time fetched code a declared, versioned, verified, recorded input so that next time the first question has a queryable answer. ## The judgment an interviewer is listening for Three things separate a strong answer. First, recognising that the *artifact*, not only the credentials, is at risk - candidates who stop at token theft miss the reason this class matters. Second, being explicit about the limits of enumeration and choosing a conservative response because of them, rather than claiming a precise blast radius the evidence cannot support. Third, treating the downstream notification as part of the job rather than a communications afterthought: you are the upstream in someone else's supply chain, and only you can tell them which versions to distrust.

  • Your CI logs only retain fourteen days and the exposure window is longer. What do you do?
    Bound the window from the hosting side instead, and treat the unlogged period as fully affected rather than unknown-therefore-clean. Rotate credentials for every pipeline that plausibly used the helper in that period, rebuild artifacts published then, and record the retention gap as a finding with an action to extend retention.
  • Why rotate credentials by scope rather than only those you can prove were accessed?
    Because the evidence lives in systems the attacker code could observe and, in some cases, influence, and because absence of a log line is not proof of non-use. Rotation is cheap relative to a second compromise. Scope-based rotation also forces the useful conversation about which credentials were long-lived and over-broad in the first place.
  • How would the investigation differ if the helper had been a declared, version-pinned dependency?
    The first question becomes a query rather than a forensic exercise: manifests, lockfiles and per-artifact component inventories tell you which builds included which version, independent of log retention. You would still rotate and rebuild, but you could scope precisely and give downstream consumers an exact affected list instead of a conservative one.
  • What do you tell downstream consumers of the artifacts published in the window?
    The window, the specific artifact names, versions and digests published inside it, what the tampering could have done, and a clear instruction - upgrade to a named rebuilt version, and treat credentials their own pipelines exposed to the affected artifact as suspect. Specifics let them run their own enumeration; a general advisory does not.

saying these in an interview costs you the question

  • Scopes the incident to stolen tokens only
  • Claims a precise blast radius the logs cannot support
  • Assumes clean job logs mean a job was unaffected
  • Rebuilds nothing because tests passed at the time
  • Treats downstream notification as optional
  • Deletes the bad file and calls it remediated

context