skip to content

Your self-hosted runners run inside the production VPC. Why does that turn every pull-request build into a network attack position?

level: seniorimportance: must knowfreq 52%

answer

  1. a build is code the proposer chose
  2. it inherits reach, not a password
  3. network location is not authorisation
  4. machine identity endpoint is the fast path
  5. fix placement, not hygiene

basics

~20 s

A build runs contributor-authored code with the runner's network identity. Anything the runner can reach - internal admin APIs trusting the network, the cluster control plane, the machine identity endpoint - becomes reachable by whoever can open a pull request.

solid answer

~50 s

A build job is arbitrary code chosen by whoever proposed the change: the build script lives in the repository, tests run, and dependency install hooks execute. Put that on a machine inside a trusted network zone and you have handed an outsider an execution position on your internal network. What they inherit is not a credential but **reach** - internal services that authorise by network location, the cluster control plane, another team's build infrastructure, and the platform's machine identity endpoint with whatever role the host carries. In a bank whose runner pool shares a segment with the payments admin API, anyone who can trigger a build has an unauthenticated position in front of it. The fixes are placement and identity, not hygiene: untrusted builds run in an isolated segment, internal services require real authentication, the machine identity endpoint is blocked, and pools are separated by trust tier.

go deeper

for a junior

Know that a build executes code proposed by the change, and that a runner placed on an internal network gives that code the same reach the machine has.

for a middle

Explain what is reachable in practice - services that trust the network, the control plane, the machine identity endpoint - and why the job inherits reach rather than a stolen credential.

for a senior

Demonstrate the fix in production terms: separate pools by trust tier with different network placement, expose single dependencies instead of relocating jobs, and scope host identity to the minimum.

for a principal

Own the boundary decision across the estate: where untrusted execution is allowed to sit, what internal services must stop trusting location, and how you fund that change when the services predate the pipeline.

## What a build job really is A CI job executes code that the change proposes. That includes the build configuration itself if it lives in the repository, the test suite, and any lifecycle or install hook a dependency defines. So the honest model is: **a pull request is a remote code execution request against your build fleet**, granted by design. The only question is what that execution can reach. ## Network position is a capability Self-hosting a runner usually means putting it where the builds need to be - inside the network with the artifact repository, the package proxy, the internal services builds integrate against. That placement is a capability handed to the job. It cannot be revoked by the job's own permissions, because most of what is reachable does not check permissions at all; it checks whether the caller is inside. What is typically reachable from a runner inside a trusted zone: - **Internal admin and operations APIs** that were built assuming only staff on the internal network could call them. - **The cluster control plane** and other orchestration endpoints. - **Databases and caches** that restrict by network reachability rather than by credential. - **Other build infrastructure** - the artifact repository (often with a write credential on the host), the package proxy, other teams' runners. - **The platform's machine identity endpoint**, from which a job can request credentials for whatever role the host was assigned. This is frequently the fastest path from build execution to cloud API access. - **The rest of the runner pool**, because runners in one pool usually sit in one segment. ## The scenario that makes it concrete A bank runs its own runners inside the VPC so builds can reach the internal package proxy. The same segment holds a payments administration API that authenticates nothing beyond being called from inside, and the cluster control plane. One repository in that organisation accepts external contributions and is wired to that pool. From that moment, opening a pull request is enough to run code with a line of sight to an interface that moves money and to the control plane that runs the services. No credential was stolen and no vulnerability was exploited; the trust boundary was drawn in the wrong place. The assets at risk here are not customer records - they are **money and the availability of internal services**, which is why this question survives the answer "but that repository has no secrets". ## Why this is worse than a compromised laptop A runner has no human attached, so there is no second factor and no one to notice odd behaviour. It usually carries less endpoint monitoring than a workstation. It sits in a more trusted network zone than a laptop ever would. And it re-executes on demand - the attacker does not need persistence, only the ability to open another pull request. ## Controls that actually change the picture 1. **Separate pools by trust tier, strictly.** Builds triggered by code that has not been reviewed by someone you trust run on a pool with no inward reach. Post-merge and release builds may run somewhere more privileged. Labels alone are not separation - the network placement has to differ. 2. **Put untrusted execution outside the trusted zone.** If a build genuinely needs an internal dependency, expose that one dependency to the untrusted segment through a controlled path rather than moving the whole job inside. 3. **Stop treating network location as authorisation.** Internal services should require real authentication and authorisation, so a foothold on the network is not automatically a foothold in the application. 4. **Block or proxy the machine identity endpoint** from build jobs, and give the host the smallest role it can do its work with. 5. **Scope credentials to the job.** Short-lived, narrowly scoped tokens issued for the job beat a long-lived token sitting on the host for anyone who lands there. 6. **Watch lateral connections.** Outbound filtering matters, but note the shape of this risk: the valuable targets are **inside**, so east-west visibility and segmentation are what apply here. ## A common wrong answer "We only let approved contributors trigger builds on that pool" is better than nothing but is not a boundary. Approval processes get bypassed by a compromised contributor account, by a repository that forgets to require approval, and by a dependency whose install hook runs in a build nobody thought of as untrusted. Design so that a build landing on an inward-facing host is impossible, not merely disallowed by policy. ## Saying it in an interview Lead with the framing - a build is remote code execution you granted on purpose - then name reach rather than credentials as the thing inherited, then give placement and real service authentication as the fix. Runner hygiene is a different problem and does not touch this one.

  • The team says the repository holds no secrets, so a foothold there is harmless. What do you say?
    That secrets in the repository are not the asset. The asset is the runner's network position and machine identity: the internal services that authorise by location, the control plane, and whatever cloud role the host carries. A repository with nothing valuable in it can still be the cheapest door into a segment full of valuable things.
  • How do you keep builds that legitimately need an internal package proxy from requiring a runner inside the trusted zone?
    Expose the one dependency rather than relocating the job. Publish the proxy through a controlled path reachable from the untrusted build segment, authenticated and read-only, with its own rate limits and logging. The build gets what it needs; the job never acquires line of sight to the control plane or the internal admin APIs.
  • Is a compromised runner worse than a compromised developer laptop?
    Usually yes. The runner sits in a more trusted network zone, has no human or second factor attached, typically carries lighter endpoint monitoring, and re-executes attacker code on demand every time a build is triggered. A laptop at least belongs to someone who may notice, and rarely has a direct line to the control plane.

saying these in an interview costs you the question

  • Treats network location as if it were authorisation
  • Says the repository has no secrets so the risk is low
  • Relies on contributor approval as the only boundary
  • Forgets the machine identity endpoint reachable from the job
  • Answers with runner cleanup instead of network placement

context