skip to content

You cannot make every runner ephemeral. How do you decide which builds may keep a persistent host?

level: principalimportance: should knowfreq 28%

answer

  1. two axes: input trust, host value
  2. one cell you refuse outright
  3. persistence is a priced exception
  4. single tenant, bounded lifetime, no resident keys
  5. owner and expiry, or it is permanent

basics

~20 s

Decide by the trust level of the code the job runs and the value of what lives on the host, not by team convenience. Persistence is an exception with an owner, an expiry, and compensating controls you actually enforce.

solid answer

~50 s

Two axes decide it. **How trusted is the input** - post-merge and release builds of reviewed code are a different proposition from anything an outside or low-privilege contributor can trigger, and untrusted input should never land on a persistent host. **How valuable is what the host holds** - a fleet with signing material or production credentials sits at the highest bar regardless of who triggers the build. Where the constraint is real, such as hardware-bound fleets or caches that take an hour to warm, grant the exception and charge for it: single-tenant pools per repository or trust tier, reimaging from a golden image on a bounded cadence, no long-lived secrets on the host with signing done by a service the host calls, an immutable toolchain the job cannot write, and job-to-host assignment records kept off the machine. Then treat it as an accepted risk with a named owner and an expiry date, reviewed rather than inherited forever.

go deeper

for a junior

Understand that single-use runners are the safer default and that persistent ones exist because some builds are tied to specific hardware or slow-to-rebuild state.

for a middle

Be able to list what you would demand in exchange for persistence: single tenancy, periodic reimaging, no long-lived secrets on the host and a toolchain the job cannot modify.

for a senior

Show you can apply the trust-of-input and value-of-host axes to a real estate, refuse untrusted builds on privileged persistent fleets, and design the credential and signing path so the host holds nothing reusable.

for a principal

Own the exception process itself: who approves persistence, who funds the alternative, when each exception expires, and what you report to leadership to show the exception surface is shrinking.

## Why this is a judgment call, not a rule "Every job gets a fresh machine" is the right default and an incomplete strategy, because some fleets cannot comply. Hardware-bound builds tied to a specific licensed operating system, specialised accelerators, builds whose warm state takes tens of minutes to rebuild, and appliances you do not control all produce persistent hosts. A security lead who answers only "make them all ephemeral" has not engaged with the estate they were hired to protect. The job is to decide **where** persistence is tolerable and to price it. ## The two axes that decide it ### Axis 1: how trusted is the code the job will run Rank the triggers: - Release and post-merge builds of code that passed review - **trusted input**. - Builds of branches inside the organisation by engineers with commit rights - **semi-trusted**; the code is not reviewed yet, but the actor is known and accountable. - Builds triggered by outside contributors, or by anyone whose access to one low-value repository should not imply access to anything else - **untrusted input**. Untrusted input on a persistent host is the combination to eliminate first, because that is where an attacker with no standing gets to leave something behind for a later, more valuable job. ### Axis 2: what the host holds and can reach A host with signing material, production deployment credentials or reach into a sensitive network segment is high-value regardless of who triggers builds. High-value plus persistent means the compensating controls have to be strong enough that you would be comfortable explaining them to a customer. Combine them: the matrix has one cell you refuse outright (untrusted input on a high-value persistent host), one cell that is fine (trusted input, low-value host, single tenant), and two that need controls. ## What you charge for the exception Grant persistence only bundled with these: - **Single tenancy.** One repository, or one trust tier, per host. Contamination is only tolerable inside a blast radius you already accepted; a general pool serving everyone is not that. - **Bounded persistence.** Reimage from a golden image between jobs where the constraint allows it, and on a fixed cadence where it does not, so anything planted has a defined maximum lifetime rather than living until someone notices. - **No long-lived secrets resident on the host.** Issue short-lived, job-scoped credentials at job start. Move signing behind a service the host authenticates to and that logs every operation, so compromising the host does not mean walking away with the key. - **Immutable toolchain.** The compiler, SDK and language runtimes come from the build image and are not writable by the job account. This removes the most valuable thing an attacker could leave behind. - **Records off the host.** Job-to-host assignment, build metadata and integrity events, shipped somewhere the host cannot rewrite. Without them you can neither detect nor scope an incident on that fleet. - **A tested rebuild path.** You should be able to destroy any persistent host and have a replacement taking work quickly. If destroying a host is a project, you will clean it instead, and cleaning does not work. ## The organisational half This is where the question stops being technical. - **Every persistent fleet is an entry in an exception register**, with a named owner, the reason, the compensating controls in force, and an expiry date. Exceptions that renew silently become permanent architecture that nobody chose. - **Someone has to fund the alternative.** "Make it ephemeral" often means buying more hardware or accepting slower builds; if you do not identify who pays, the exception is permanent by default and you have written a policy that describes nothing. - **Say what you can actually prove.** When a customer or auditor asks about your build environment, the answer for these fleets is "single-tenant, reimaged on this cadence, no resident signing keys, with build records retained", not a claim of isolation you cannot demonstrate. - **Measure the exception surface.** How many fleets, what fraction of releases pass through one, and is the number going down. That trend is the real security metric here, not the existence of the policy. ## Anti-patterns to name - A **cleanup step** as the compensating control - it runs after the untrusted job, which may have subverted it. - **Labels as separation** - one pool addressed by different names is one pool. - A **written standard with no enforcement**, so the only teams complying are the ones who did not need to be told. - **Permanent exceptions with no owner**, which is how a fleet everyone has forgotten ends up signing your releases. ## Saying it in an interview Show that you can refuse the wrong cell of the matrix, grant the rest with priced controls, and put the exception on a register with an owner and a date. The failure mode being probed is a lead who either bans persistence and gets ignored, or waves it through because the fleet is old and important.

  • A team says their fleet cannot be reimaged between jobs because warming the cache takes forty minutes. What do you require instead?
    Bound the persistence rather than removing it: reimage on a fixed cadence, keep the fleet single-tenant so contamination stays inside an accepted blast radius, make the toolchain immutable so the cache is data rather than executables, and keep the cache scoped and ideally read-only to the job. Then register the exception with an owner and a review date.
  • How do you stop a persistent fleet from being the one that holds the signing keys?
    Move signing off the build host entirely. The host calls a signing service that authenticates the caller, applies its own policy about what may be signed, and logs every operation. The host then holds a short-lived capability to request a signature rather than the key itself, so compromising it does not yield reusable signing power.
  • What would you report to leadership to show this is improving?
    The exception surface, not the policy. How many persistent fleets exist, what share of production releases are built on one, how many exceptions are past their review date, and the trend over quarters. A shrinking share of releases passing through persistent hosts is evidence; a published standard is not.

saying these in an interview costs you the question

  • Answers only make everything ephemeral and stops there
  • Grants exceptions with no owner or expiry date
  • Accepts a cleanup step as the compensating control
  • Treats pool labels as if they were separation
  • Leaves signing keys resident on a long-lived host

context