skip to content

Execution Environment Integrity

Where a job runs and what it is allowed to reuse decides whether one bad build stays contained: reused runners, shared caches, network access mid-build, and who maintains the pipeline everyone copies.

on this pageshow

explore

questions

21

What does a default-deny outbound network policy on a build job actually prevent?

level: juniorimportance: must knowfreq 60%

answer

  1. assume hostile code already runs here
  2. cut the channel, not the code
  3. loot out, second stage in
  4. artifact integrity is a different control
  5. DNS still counts as outbound

basics

~20 s

Default-deny outbound stops build-time code reaching unapproved hosts: no exfiltrating secrets or source, no pulling a second-stage payload. It does not make a malicious dependency safe - the code still runs and can still corrupt the artifact.

solid answer

~50 s

A build job executes a lot of code nobody reviewed - install hooks, plugins, generated test fixtures - while holding source, publish tokens and often a cloud identity. Default-deny egress means the job can only reach destinations you explicitly allowed, so that code has no channel to send what it stole and no channel to fetch a second stage, which is how most build-time payloads actually arrive. Routing through a proxy rather than dropping silently also gives you a log, and a denied connection from a build is a very high-signal alert because legitimate builds talk to a small, stable set of hosts. What it does not do: it does not stop the malicious code running, it does not clean the artifact you publish, and it leaks through anything you had to allow - including DNS resolution, which is itself an outbound channel.

go deeper

for a junior

Be ready to say in one breath what the control blocks: unapproved outbound connections from a build job, which means no sending secrets out and no fetching a payload in.

for a middle

Explain the mechanics - the job runs with no route by default, an allowed set is named explicitly, and traffic goes through a proxy so attempts are logged rather than silently dropped.

for a senior

Show where you would apply it first and what it misses: the install stage needs almost nothing, the artifact can still be backdoored, and any allowed host plus DNS remain viable channels.

for a principal

Own the framing that this is containment rather than prevention, and be able to argue what it is worth relative to the ongoing cost of maintaining allow-lists across an estate.

## What the control actually is A build job runs a great deal of code that nobody on your team wrote or reviewed: dependency install hooks, build plugins, code generators, test fixtures, container entrypoints. Default-deny outbound means the job starts with **no route off the machine**. Every destination it is allowed to reach is named explicitly - normally by forcing all traffic through an egress proxy or by attaching a network policy to the job's namespace - and everything else is refused. The mental model that makes this click: **assume the build is already running hostile code and ask what that code can do next.** ## The threat it answers Once arbitrary code executes on a runner, it can read whatever the job can read. In a typical pipeline that is a lot: the checked-out source, the environment (publish tokens, registry credentials, a short-lived cloud identity), the runner's disk, and the link-local instance-metadata endpoint that will hand out the host's cloud role to any local process. Having read all that, the attacker needs a channel. Two directions matter: - **Outbound** - get the loot off the box. Credentials, source, signing material, customer data present in test fixtures. - **Inbound over an outbound connection** - fetch a second stage. A very common pattern is that the published package is boring and small, and the real payload is downloaded at build time from an attacker-controlled host. This keeps the malicious code out of the artifact anyone can inspect. Default-deny closes both. It also closes long-lived command-and-control, and it closes the sloppier category of unpinned installers that fetch a script from a random host mid-build. ## What it does not prevent - this is the part interviews probe - **The dependency is still malicious and still runs.** Egress control is containment, not prevention of compromise. The hostile code executes on your runner with the job's privileges. - **The artifact can still be corrupted.** An implant compiled into the binary leaves your network through your own approved channel - the registry push you obviously allow. Nothing about egress policy detects that; that is what provenance, reproducible builds and review are for. - **Allowed destinations remain open.** If the internal package index is allow-listed and the job can publish to it, a secret can be smuggled out inside a package. If a public source-code host is on the list, a repository there is a perfectly good drop point. - **DNS is an outbound channel.** If the job can resolve arbitrary names through a resolver that talks to the internet, data can be encoded into query names and the lookup alone carries it out, even when the connection that follows is blocked. So the honest framing is: default-deny egress **reduces the blast radius and buys you evidence**. It does not make untrusted code safe to run. ## Preventive plus detective Route the traffic through a proxy rather than silently dropping packets, because the log is half the value. A legitimate build talks to a small, stable set of destinations, so a **denied connection attempt from a build job is unusually high-signal** - far better than most alerts a security team gets. The deny is the preventive half; the log line is the detective half, and it is what lets you answer later which build, which stage, and how much data moved. ## Where to apply it first Egress policy does not have to be uniform. The stage that runs the most untrusted code is dependency resolution and install, and it is usually also the stage that has no legitimate reason to talk to anything except the package proxy. Tightening that one stage, and leaving the deploy stage with its broader list, gets most of the value without a month of breakage. ## What a strong answer sounds like Something close to: the build is a place where third-party code executes with production-adjacent credentials; default-deny egress means that code cannot phone home with what it stole or pull down what it still needs, and every attempt it makes is logged. It does not stop the code running, it does not clean the artifact, and it leaks through any destination you had to allow - including DNS.

  • If the malicious dependency runs anyway, what have you actually gained?
    Blast radius and evidence. The code executes but cannot ship credentials or source anywhere, and the common loader pattern - a benign-looking package that downloads its real payload at build time - simply fails. You also get a log line for every attempt, which turns an invisible compromise into an alert and gives you something to scope from later.
  • Does a strict egress allow-list stop a compromised build from shipping a backdoored artifact?
    No. The implant travels out through a destination you obviously allow - the registry you publish to. Egress control is about unapproved channels; artifact integrity is answered by reproducible builds, provenance and review of what the build produced, not by network policy.
  • Which outbound path do teams most often leave open by accident?
    DNS. Jobs usually keep name resolution working because nothing builds without it, and a resolver that reaches the internet will happily carry data encoded in query names even when every TCP connection is refused. Logging DNS queries, and resolving only through a controlled resolver, is the usual answer.

It is the difference between stopping a thief getting into the building and making sure that if one does get in, there is no door, no window and no phone line to get the loot back out.

saying these in an interview costs you the question

  • Claims deny-all egress makes a malicious dependency harmless
  • Thinks blocking outbound also blocks attacks against the runner
  • Forgets DNS resolution is itself an outbound channel
  • Treats egress policy as a replacement for dependency scanning
  • Assumes allowed destinations cannot be abused for exfiltration

context

open as a page

What makes a build hermetic, and why is a hermetic build not automatically deterministic?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A hermetic build declares and pins every input before it starts and fetches nothing while it runs. It can still emit different bytes on each run, because embedded timestamps, absolute paths and file ordering vary independently of the inputs.

open as a page

Why is a build runner that keeps its state between jobs treated as a supply-chain risk, not just a flakiness risk?

level: juniorimportance: must knowfreq 60%

basics

~20 s

A reused runner lets one job write what the next job reads. Whatever an earlier job left on disk - tools on PATH, caches, unlocked keys - silently shapes the next artifact and exposes that job's secrets.

open as a page

A Go build cache keyed only on the go.sum hash is shared by every branch's CI jobs - what is the risk?

level: middleimportance: must knowfreq 60%

basics

~20 s

Any engineer who can push a branch computes the same key, writes that entry, and the release build restores it. A key describes content inputs, never who produced them, so one namespace across all branches erases the trust boundary.

open as a page

Your hardened base image is rebuilt monthly, but services pin it by digest. How do you make the rebuild reach them?

level: middleimportance: must knowfreq 60%

basics

~20 s

A digest names one immutable build, so a rebuild produces a new digest nobody consumes. Rebuilding is the supply side; you need a push mechanism: automated bump changes per repository, a build-time staleness failure, and retirement of the old digest.

open as a page

Your self-hosted runners run inside the production VPC. Why does that turn every pull-request build into a network attack position?

level: seniorimportance: must knowfreq 52%

basics

~20 s

A build runs contributor-authored code with the runner's network identity. Anything the runner can reach - internal admin APIs trusting the network, the cluster control plane, the machine identity endpoint - becomes reachable by whoever can open a pull request.

open as a page

In a CI pipeline, what is cache poisoning and why does the build not notice it?

level: juniorimportance: should knowfreq 48%

basics

~20 s

Cache poisoning means getting tampered content into a cache entry a later build restores. The build treats restored files as if it had produced them: nothing is signed, nothing is compared, and the log shows only a cache hit.

open as a page

What does a team lose by copy-pasting a shared CI pipeline template instead of referencing a versioned one?

level: juniorimportance: should knowfreq 52%

basics

~20 s

A copy stops receiving updates: security fixes made centrally never reach it, and the local edits made to it are invisible to the platform team. A versioned reference keeps one source of truth and makes adoption measurable.

open as a page

Your build runners now deny all outbound traffic - what belongs on the egress allow-list?

level: middleimportance: should knowfreq 47%

basics

~20 s

The package proxy, the image registry, source hosts, the CI control plane's log and artifact endpoints, identity and secret endpoints, telemetry, plus plumbing like DNS and time sync. Scope the list per stage rather than keeping one global union.

open as a page

The same commit produces byte-different OS packages on two build machines - what non-determinism do you hunt for?

level: middleimportance: should knowfreq 47%

basics

~20 s

Hunt the values that differ between machines but are not build inputs: embedded build timestamps, absolute workspace paths, filesystem file ordering, locale and timezone, hostname and username, and the uid, gid and permission bits recorded in the archive.

open as a page

Each CI job runs in its own container on a shared host - what contamination does that not prevent?

level: middleimportance: should knowfreq 47%

basics

~20 s

A fresh container is not a fresh machine. Whatever is mounted in from the host - workspaces, tool and dependency caches, a container daemon socket - plus the host's network position, carries state and access between jobs.

open as a page

A build job uploads a jar and the deploy job ships whatever it downloads - how do you close that gap?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Bind the handoff to content, not to a name. Have the build record the jar's digest, have the deploy recompute it and refuse on a mismatch, and carry the expected digest where the run's other jobs cannot rewrite it.

open as a page

How does running dependency install in a credential-less job contain a malicious package?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Install runs in a job holding no cloud role, no publish token and no repo write, with egress limited to the package proxy. Hostile install code then executes with nothing worth stealing and nowhere to send it.

open as a page

One step in an otherwise hermetic build graph downloads a reference data file - how do you find it and what replaces it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Find it by running the build in a sandbox with outbound access denied and seeing which step fails. Replace the fetch with a declared input pinned by content digest, fetched and verified in a separate phase before the sealed build runs.

open as a page

Your shared release template has a flag that skips the signing stage, and twelve services set it. What do you do?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Find out why the hatch is used before removing it: it usually marks a gap in the paved road. Make every use attributed, dated and reported, close the gaps, then narrow the hatch to an approved exception and delete it.

open as a page

You cannot make every runner ephemeral. How do you decide which builds may keep a persistent host?

level: principalimportance: should knowfreq 28%

basics

~20 s

Decide by the trust level of the code the job runs and the value of what lives on the host, not by team convenience. Persistence is an exception with an owner, an expiry, and compensating controls you actually enforce.

open as a page

A training pipeline restores a cached preprocessed feature shard - what integrity risk does that add?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

The shard is an unverified input to the model with no record of who produced it, and tampering breaks nothing: training succeeds and only the model's behaviour changes. Caching derived data moves integrity risk into data nobody reviews.

open as a page

An egress log shows a nightly build reached an unexpected host - how do you find which stage?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Attribution needs a per-job egress identity - a source address or proxy credential that puts the job id in every log line - intersected with retained per-step timestamps. Bytes sent versus received then separates exfiltration from a downloaded payload.

open as a page

A build declares both an internal package index and a public index as sources - what does that cost hermeticity?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Two declared sources mean the origin of a component is decided at build time by whichever index answers, not by your declaration. Hermeticity wants exactly one resolvable source, no fallback on a miss, and a build record naming the source and hash for each resolved component.

open as a page

Someone plants a compiler shim in a shared runner's tool cache. How do you detect that builds were tampered with?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Rarely from the artifact alone. Detection comes from rebuilding on known-clean infrastructure and comparing, from build records kept off the host that say which host and toolchain produced each artifact, and from integrity monitoring of the tool-cache path.

open as a page

How do you get 40 teams onto a hardened shared pipeline when a written policy alone has not worked?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Make the hardened path the fastest route to production so complying costs less than not complying, absorb the migration work centrally, keep the off-road route available but expensive, and measure adoption by supported version rather than by mere presence.

open as a page