skip to content

Runners, Agents & Isolation

Where the job actually executes and how strongly it is isolated from the next job. This is asked because self-hosted runners are the usual cost lever and, when they are persistent and shared, the usual way a build steals another team's credentials.

on this pageshow

questions

5

In a CI/CD system, what is the difference between a provider-hosted runner and a self-hosted runner, and what does choosing self-hosted actually change?

level: juniorimportance: must knowfreq 72%

answer

  1. who owns the machine
  2. fresh per job versus yours forever
  3. cost lever and trust lever
  4. private network reach is the real driver
  5. your machine, your cleanup and patching

basics

~20 s

A provider-hosted runner is a fresh machine the CI vendor creates per job, bills per minute, then destroys. A self-hosted runner is a machine you own and register yourself: usually cheaper at volume and able to reach private networks, but you inherit patching, isolation and cleanup.

solid answer

~50 s

Every CI platform splits into a coordinator that schedules jobs and machines that execute them; that executing machine is the runner or agent. With **provider-hosted** runners the vendor supplies a standard image, creates a clean instance for the job, destroys it afterwards, and bills per minute — you get guaranteed clean state and zero maintenance, but fixed hardware, no route into your private network, and a cost that scales linearly with build volume. With **self-hosted** runners you install the agent on your own machine and it polls the coordinator for work. People do this for cost at high volume, for hardware the vendor does not sell, and above all for reachability of private registries and databases. What changes is ownership: you now own OS patching, disk cleanup, capacity when a queue forms, and — the part that gets missed — the trust boundary, because a job on your machine is arbitrary code executing inside your network.

go deeper

for a junior

Be able to say plainly that a hosted runner is a clean machine the vendor creates and destroys per job, while a self-hosted runner is your own machine running the agent, and name cost and private-network access as the usual reasons to switch.

for a middle

Explain the mechanics: what persists between jobs on each, why hosted runners need an explicit cache step, and what obligations — patching, disk cleanup, capacity — move to you the moment you self-host.

for a senior

Show that you treat self-hosting as a trust decision. Talk about single-use instances, pool separation by trust level, and keeping untrusted contributions off machines that sit inside your network.

for a principal

Own the fleet strategy: model build-minute cost against operating a pool, decide which workloads justify a private footprint at all, and keep the high-trust pool small enough that one team can genuinely patch and audit it.

## The runner is where everything actually happens A CI/CD platform has two halves. The coordinator (the server, the control plane) decides which jobs should run, in what order, and records the results. The other half is a fleet of machines that actually run the work. The process on that machine — called a runner, an agent, or an executor depending on the vendor — receives a job, checks out the code, runs the steps, streams logs back, and reports success or failure. This matters because everything a pipeline "has" is a property of that machine, not of the pipeline file. How much disk is free, which compilers are installed, whether the internal package repository is reachable, what credentials happen to be sitting in the runner user's home directory — none of that is written in the YAML. When a build behaves differently than expected, the runner is usually the reason. ## The hosted deal Provider-hosted runners are machines the CI vendor operates. They boot from a standard image the vendor maintains, run exactly one job, and are discarded. You are billed for the minutes consumed. What you get: clean state on every run with no effort, an image kept patched and stocked with common toolchains by someone else, and elastic capacity — twenty pull requests at once means twenty machines, with no capacity planning on your side. What you give up: hardware choice is limited to the sizes on the price list; there is no route from that machine into your private network, so a job that must reach an internal database or artifact repository needs a tunnel or a publicly reachable endpoint; nothing persists between jobs, so any cache must be explicitly saved and restored through the platform's cache service; and the per-minute price scales with your build volume, which is fine at ten builds a day and material at ten thousand. ## Why teams self-host Four reasons come up repeatedly, and cost is only one of them. - **Cost at volume.** Long builds on large machines eventually cost more per minute than an always-on instance you rent directly. - **Hardware the vendor does not sell.** GPUs, very large memory, a specific CPU architecture, physical test devices, or a licensed operating system. - **Network reachability.** The build needs the internal package mirror, a licence server, or a database that will never be exposed publicly. - **Compliance.** Source code or data must stay inside a particular network or jurisdiction. A large warm cache is a fifth, weaker reason — weaker because the same effect is obtainable with a shared cache service without giving up clean machines. ## What you inherit Choosing self-hosted moves three obligations onto your team. **Operations.** Someone patches the OS and keeps the toolchain from drifting away from what developers use locally. Someone notices when the disk fills with old workspaces and container images. Someone owns capacity: a hosted pool absorbs a burst, a fixed self-hosted pool grows a queue. **Isolation.** If one machine runs many jobs, those jobs share a filesystem, a user account, a process table, and whatever the previous job left behind. Without a deliberate boundary, one team's build can read another team's credentials. **The trust boundary.** This is the one candidates skip. A self-hosted runner is, functionally, a general-purpose code-execution service sitting on your internal network, and it will execute whatever the pipeline tells it to. "It is behind our firewall" is the reason it is more dangerous, not less: a hosted runner that gets compromised can reach nothing of yours, while yours can reach everything the machine can. That is why untrusted contributions — pull requests from forks of a public repository — must not be routed to self-hosted machines. ## Deciding Three questions settle most cases. What does a build-minute actually cost us at current and projected volume? Does any job need something hosted cannot provide — hardware, private-network reachability, a data-residency guarantee? And can we run the fleet properly: single job per machine, destroyed afterwards, patched, and separated from untrusted work? The common answer is a hybrid. Hosted runners take the bulk of ordinary builds and every job triggered by an outside contributor, and a small self-hosted pool handles the jobs that genuinely need to be inside the network — typically deployments and specialised builds. That keeps the expensive, high-trust footprint small enough to actually operate well.

  • Beyond raw cost, what usually justifies moving to self-hosted runners?
    Reachability and hardware. The build needs an internal package mirror, licence server or database that is not publicly exposed; or it needs a GPU, a specific CPU architecture, very large memory, a licensed OS, or physical test devices that the hosted price list does not offer. Data-residency or compliance rules are the third common driver.
  • What is the first control you add once you take on self-hosted runners?
    Make them single-use: one job per instance, then destroy and replace it, so nothing survives from one job to the next. After that, separate pools by trust level, run jobs unprivileged, and keep anything triggered by an outside contributor off them entirely. Patching and disk cleanup follow, but statelessness removes the largest class of problems.
  • If a job on a hosted runner must reach a private database, what are the options?
    Either bring the runner to the network — a self-hosted instance inside it — or bring the network to the runner via an outbound tunnel or a VPN client the job starts, with a narrowly scoped, short-lived credential. A third option is to remove the need: run that step from inside the network as a separate deployment job the pipeline triggers.

saying these in an interview costs you the question

  • Self-hosted is simply cheaper, there is no downside
  • Self-hosted runners are safer because they are behind the firewall
  • Hosted runners keep the workspace between jobs, so caching is free
  • Hosted runners can reach our internal database by default
  • Any CI job is isolated automatically because it runs in a container

context

open as a page

A CI job can execute in a container on a shared host, in a virtual machine created for it, or directly on a shared bare-metal machine. What does each boundary actually contain, and what leaks across it?

level: middleimportance: should knowfreq 46%

basics

~20 s

A container isolates the filesystem view, processes and network but shares the host kernel, so a kernel bug or a privileged mount escapes it. A virtual machine gives the job its own kernel and is the strongest boundary. A shared bare-metal machine contains nothing: jobs share a user, disk and credentials.

open as a page

A build passes on your team's long-lived shared CI machines but fails on a freshly provisioned one. What kinds of leftover state on a reused runner cause that, and what removes the whole class of problem?

level: middleimportance: should knowfreq 56%

basics

~20 s

A reused runner carries state forward: leftover workspace files, globally installed tools, package-manager config and credentials in the agent user's home directory, stray processes, and cached images. The build silently depends on it. Single-use runners — one job per instance, then destroyed — remove the class.

open as a page

Your CI must post a coverage comment on pull requests opened from forks, so a team proposes running the fork's build in a job that already holds the repository write token. Why is that dangerous, and what is the safe pattern?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A fork pull request is untrusted code, and any credential present in the job's environment is available to it — build scripts, test hooks and dependency install hooks all execute before anyone reviews the diff. Split it: run the untrusted build with no credentials, hand its output to a separate trusted job that never checks out the fork's code.

open as a page

You own an autoscaling pool of single-use CI runners. How do you decide between scaling to zero and keeping warm capacity, and what would you measure to know the choice was right?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Trade idle machine cost against developer wait time, and price both. Measure queue wait separately from job duration, at high percentiles, alongside provisioning latency. Fix slow provisioning before buying warm capacity, then size a minimum pool to absorb the usual arrival burst.

open as a page