skip to content

Why does mounting the container runtime's control socket into a container hand that container administrative access to the host?

level: middleimportance: must knowfreq 54%

answer

  1. the socket is the runtime's full API
  2. create-with-any-settings is in the protocol
  3. the new container is a sibling
  4. read-only mount does not filter requests
  5. caller needs no privilege of its own

basics

~20 s

Anything that can reach the runtime's control socket can ask it to create a privileged container with the host's filesystem mounted in, and the runtime obeys as a host-level process. No privilege inside the calling container is needed.

solid answer

~50 s

The service behind that socket is what actually creates containers on the machine, so it runs with administrative privilege. Its request set includes 'create a container with these settings' — and those settings include privileged mode and a writable mount of the host's own filesystem. A workload that can open the socket and speak the protocol simply asks for that container and starts it. The created container is a **sibling produced by the runtime**, not a child of the caller: nothing the caller was denied applies to it, so running the caller unprivileged and as an ordinary account changes nothing. Even without creating anything, socket access exposes every other container's configuration and environment on that host and a way to run commands inside them. Treat the mount as equivalent to an administrative account on the machine.

code

pseudocode · 18 lines
pseudocode
# running inside a container that can open the runtime's control socket

request = {
  createContainer: {
    image:          "any image this host can fetch",
    privilegedMode: true,
    hostMounts:     [ { hostPath: "/", mountedAt: "/host", access: "read-write" } ],
    command:        "append an entry to the host's scheduled-task table"
  }
}

send request        over the control socket
send startContainer over the control socket

# the new container was created BY the runtime, as a sibling:
#   - it does not inherit our privilege set, our call allowlist or our profile
#   - the account we run under never applied to it
#   - we needed no privilege except the ability to open the socket

go deeper

for a junior

Remember the equivalence: reaching the runtime's control socket from inside a container is the same as holding an administrative account on that host, because the runtime will build whatever container it is asked for.

for a middle

Explain the chain end to end — create request with privileged mode and a writable host mount, runtime obliges, sibling container writes into something the host executes — and why the caller's own restrictions never apply to that sibling.

for a senior

Refuse the mount and offer the narrower path: read the identity the platform already injects, or front the socket with a proxy that permits only named read requests. Name the neighbour-secrets exposure even when nothing is created.

for a principal

Decide the estate rule for per-host agents that demand this grant — whether such a workload is accepted at all, what evidence its vendor must provide, and who is allowed to change its spec once it runs everywhere.

## What the control socket is The runtime that actually creates containers on a host exposes a local control interface — usually a socket file in the host's filesystem — speaking a request/response protocol. Its requests are roughly: create a container from this image with these settings, start it, stop it, run a process inside a running one, copy files in or out, list what is running, read a container's full configuration. The service answering those requests runs with administrative privilege on the machine, because creating containers requires it. That is the entire mechanism, and the escalation follows from it in one step. ## Why a mount of it is a privilege grant - The request set includes **create with arbitrary settings** — there is no subset of the protocol that creates only well-behaved containers. - The container that gets created is a **sibling**, made by the runtime on the host, not a child of the requester. - Therefore **nothing the requester was denied applies to the new container**. The caller's trimmed privilege set, its call allowlist and its access-control profile were attached to the caller, not to the runtime. - The requester needs **no privilege beyond the ability to open the socket and speak its protocol**. The worked chain, which is what an interviewer wants to hear: 1. Ask the runtime to create a container that is privileged and mounts the host's root filesystem read-write. 2. The runtime obliges — that is its job, and the request is well formed. 3. Have that container write into something the host executes: its table of scheduled tasks, a start-up service definition, the trusted-keys file of the administrative account. 4. The host runs it. Elapsed time: seconds, with no exploit and no kernel defect involved. ## The three misunderstandings that keep this alive | What the team believes | What is actually true | |---|---| | "We mounted it read-only, so it is safe" | The read-only flag applies to the filesystem entry, not to the conversation. A socket is written to in both directions by design; the requests still go through. | | "Our code only calls the list endpoint" | The socket does not know which calls the code intends to make. The grant is the whole protocol, not the subset in use today. | | "The process inside runs as an unprivileged account" | Irrelevant. The privileged act is performed by the runtime on the caller's behalf; the caller only has to be able to ask. | ## What it gives even if nothing new is created Creating a privileged sibling is the headline, but socket access alone already yields, for **every** container on that host: - its full configuration, including values passed through its environment, which is where credentials most often live; - the ability to run a process inside it, i.e. a shell in someone else's workload; - the ability to copy files out of it; - the ability to stop it, which is an availability problem on top of the confidentiality one. So even a genuinely read-only use of the socket is a read of every neighbouring workload's secrets. ## Narrower ways to get what the workload actually wanted Most requests for the socket come from a per-host collector that wants to label what it collects with the name or identity of the workload it came from. That does not need the socket: - the platform already injects the workload's identity into its own environment or writes it into the path where the data lands — read it from there; - where richer data genuinely is needed, put a filtering proxy in front of the socket that permits exactly the read requests required, and let nothing but the proxy reach the socket itself; - runtime designs differ here — some have no single long-lived privileged service at all, so this exact route does not exist in the same form, and the equivalent question on those becomes what the workload can ask the machine to run on its behalf. ## How to say it in a review State the equivalence rather than arguing about intent: *this mount is an administrative account on this host, issued to this workload and to anything that compromises it*. Then ask the two useful questions — what data does it actually need, and where else is that data already written? A per-host collector that gets the socket is also the highest-value target on the estate, because it runs everywhere and holds the same grant everywhere.

  • The team says they only need to list containers — can the socket be mounted read-only?
    No. A read-only mount marks the filesystem entry, not the conversation, and a socket is written to in both directions by design, so the create request still goes through. If only specific read requests are needed, put a filtering proxy in front of the socket and let only the proxy reach it — the filtering has to happen at the protocol, not at the mount.
  • Does running the calling container under an unprivileged account prevent this escalation?
    No. The privileged act is performed by the runtime, which already holds administrative privilege on the host. The caller contributes only a well-formed request. Account hardening inside the caller is worth doing for other reasons, but it does not touch this route at all.
  • What does a compromised workload with socket access get from its neighbours, without creating anything?
    The full configuration of every container on that host, including environment values, which is where credentials usually are; the ability to run a process inside any of them; the ability to copy files out; and the ability to stop them. A purely read-only use of the socket already reads every neighbour's secrets.

Being locked in one office but holding a direct line to the building manager, who will unlock any door for whoever calls. You never picked a lock; you asked someone with the keys.

saying these in an interview costs you the question

  • Says a read-only mount of the socket makes it safe
  • Thinks the container it creates inherits the caller's confinement
  • Believes an unprivileged account inside the caller prevents the escalation
  • Treats socket access as equivalent to read-only monitoring data
  • Assumes only the runtime's own tooling can speak that protocol