skip to content

A CI job can execute in a container on a shared host, in a virtual machine created for it, or directly on a shared bare-metal machine. What does each boundary actually contain, and what leaks across it?

level: middleimportance: should knowfreq 46%

answer

  1. shared kernel versus own kernel
  2. namespaces and cgroups, not magic
  3. the socket mount undoes it all
  4. hypervisor is the strong line
  5. bare metal shares user, disk, credentials

basics

~20 s

A container isolates the filesystem view, processes and network but shares the host kernel, so a kernel bug or a privileged mount escapes it. A virtual machine gives the job its own kernel and is the strongest boundary. A shared bare-metal machine contains nothing: jobs share a user, disk and credentials.

solid answer

~50 s

Think of it as three strengths. Running the job as a process on a shared machine isolates almost nothing — same user, same disk, same leftover credentials, same process table — so any job can read what any other job left. A container adds kernel-enforced namespaces and cgroups: separate filesystem view, process table, network stack and resource limits. That is a genuine boundary, but the kernel is shared, so a kernel vulnerability crosses it, and it collapses entirely if the job is given the host's container socket or extra privileges. A virtual machine gives the job its own kernel behind the hypervisor, which is why multi-tenant CI providers use per-job VMs; microVM technology such as Firecracker makes that cheap enough to do per job. The practical rule is that the boundary is only as strong as what you hand the job — a container with the host socket mounted is not isolated at all.

go deeper

for a junior

Know the ordering: a plain process on a shared machine isolates least, a container isolates the filesystem and processes, and a virtual machine isolates most because it has its own kernel.

for a middle

Explain the mechanisms — namespaces, cgroups, seccomp, the shared kernel — and be ready to say exactly why mounting the host's container socket into a job removes the boundary.

for a senior

Match the boundary to the trust model in a real fleet: which workloads justify per-job VMs, how you keep privileged image builds off shared machines, and what you audit in a runner's configuration.

for a principal

Own the tradeoff across the organisation: per-job VM isolation costs startup latency and duplicated resources on every job, so decide where that spend buys real separation and where a container pool per trust domain is enough.

## Why the boundary matters in CI specifically A CI job runs code that people change every day, including code from contributors you have never met. The question "what can this job touch" therefore has two separate audiences: security (can this job reach another team's secrets, or the host) and reproducibility (can this job see leftovers from the last one). The isolation boundary answers both. ## Bare process on a shared machine The weakest arrangement is the agent running jobs directly on a host, as its own operating-system user, in a working directory it reuses. Nothing here is a boundary in the security sense. Every job runs as the same user, so every job can read every file that user can read: the previous job's checkout, its `~/.docker/config.json`, its `~/.npmrc`, its git credential store, its ssh-agent socket. Jobs can see and signal each other's processes, bind the same ports, and modify globally installed tools. If you run two teams' pipelines on one such machine, you have not separated those teams at all. Giving each job its own OS user helps a little — file permissions then apply — but it does not contain a job that can install packages, does not stop port collisions, and does not clean anything up. ## Container on a shared host A container is a normal process on the host kernel with a set of restrictions applied. Linux namespaces give it a private view of a subsystem — mount namespace for the filesystem, PID namespace for the process table, network namespace for interfaces and ports, and optionally a user namespace to map root inside the container to an unprivileged UID outside. Control groups (cgroups) cap CPU, memory and I/O. Seccomp filters restrict which system calls it may make. For CI this is usually the right default: each job gets a clean filesystem from an image, cannot see other jobs' processes, cannot bind a port another job is using, and cannot exhaust the whole host's memory. Reproducibility improves too, because the toolchain comes from the image rather than from whatever is installed on the host. What it does not contain: the kernel is shared, so a kernel vulnerability reachable from inside a container is a host compromise. More importantly in practice, most CI container escapes are not exploits at all — they are configuration: ```bash # Both of these hand the job the host, not a sandbox: docker run -v /var/run/docker.sock:/var/run/docker.sock build-image docker run --privileged build-image ``` Mounting the host's container socket lets the job ask the daemon to start a new container with the host root filesystem mounted — that is root on the host, obtained through an API call rather than an exploit. `--privileged` drops the capability and device restrictions that made the container a boundary in the first place. Both appear constantly in CI because image builds and Docker-in-Docker seem to need them; the answer is a rootless or daemonless image builder, not a privileged container. ## Virtual machine per job A VM runs its own kernel on virtual hardware provided by a hypervisor. The interface between guest and host is much narrower than the system-call surface a container shares, which is why every large multi-tenant CI provider isolates customers at the VM level rather than the container level. The historical objection was cost: a full VM took tens of seconds to boot. MicroVM technology such as Firecracker, and sandboxed container runtimes such as gVisor (which intercepts system calls in userspace) and Kata Containers (which runs each container inside a lightweight VM), have narrowed that gap enough that per-job VM-grade isolation is practical. The tradeoff is startup latency and resource overhead per job — you pay boot time and a duplicated kernel and page cache for every job, which matters when jobs are short and numerous. ## Choosing The decision follows the trust model, not the technology preference. Jobs from a single trusted team, on a machine only that team uses, are adequately served by containers. Jobs from multiple teams with different secrets, or from outside contributors, want a VM-grade boundary and a fresh instance per job. And whichever you choose, audit what you hand the job: privileged mode, the host container socket, host path mounts, and host networking each individually erase the boundary you thought you had.

  • CI image builds are the usual reason teams reach for privileged containers. What is the alternative?
    Use a builder that does not need the host daemon: a rootless or daemonless image builder that constructs layers as an unprivileged user, or a remote build service the job talks to over an authenticated API. Either keeps the boundary intact. If a privileged builder is unavoidable, confine it to a dedicated pool of single-use machines that hold no other team's credentials.
  • Where do gVisor and Kata Containers sit between containers and full VMs?
    Both keep the container interface while strengthening what backs it. gVisor interposes a userspace kernel that services the container's system calls, shrinking the host kernel surface it can reach. Kata runs each container inside a lightweight VM, so the boundary is the hypervisor. Both cost some performance and some syscall compatibility in exchange for a narrower escape surface.

saying these in an interview costs you the question

  • Containers are virtual machines, so the isolation is equivalent
  • A container cannot affect the host under any circumstances
  • Mounting the container socket only lets the job build images
  • Running as a different OS user is enough to separate two teams
  • Isolation is purely a security concern, not a build-correctness one

context