When a container starts, what does the container manager prepare, and what does the low-level runtime actually do with it?
answer
- two programs, not one
- one unpacks, one fences
- a directory plus a document
- manager merges image defaults with the request
- runtime execs the first process, then exits
basics
~20 sTwo components split the work. A container manager pulls and unpacks an image into a root filesystem and writes a configuration document beside it; a low-level runtime then applies the fences named in that document and executes the declared command as the first process inside the boundary.
solid answer
~50 sStarting a container is two jobs, normally done by two separate programs. The **manager** is what you or a scheduler talks to: it resolves the image reference, pulls and verifies the layers, unpacks them into a single root filesystem tree, and merges the image's own recorded defaults — command, environment, working directory, user — with what the caller asked for — mounts, ceilings, which views to create — into one configuration document. That directory plus that document is the entire hand-off. The **low-level runtime** is invoked on it: it creates the per-resource views, joins the accounting and limiting group, mounts what the document lists, pivots onto the unpacked root filesystem, drops the privilege set, and executes the declared command as the first process inside the boundary. In the common design it then exits, leaving the workload running.
code
pseudocode · 21 lines# what the manager assembles, and all the runtime is given
bundle/
rootFilesystem/ # image layers unpacked, in order, into one tree
runtimeConfig:
process:
command: ["telemetry-ingester", "--listen", "9000"]
workingDir: "/srv"
user: 10001 # from the image, unless the caller overrode it
env: ["REGION=eu-west", "BATCH_SIZE=512"]
views: [filesystem, processTable, network, ids, hostname]
accountingGroup: "/workloads/ingester-7"
ceilings:
cpuQuota: "2 cores"
memoryMax: "512Mi"
mounts:
- source: "/var/lib/telemetry", target: "/data", mode: "read-write"
privileges:
keep: ["bind-low-port"] # everything else dropped
# runtime.create(bundle) -> runtime.start(bundle) -> the command replaces it -> runtime exitsgo deeper
Recall that starting a container is two steps: something fetches and unpacks the image, and something else fences and runs the process. Being able to say which of the two talks to a registry is already most of the answer.
Explain the hand-off concretely — an unpacked root filesystem plus a configuration document — and say what goes into that document: the command, the views, the mounts, the ceilings, the user. Note that an image contributes defaults, not fencing.
Show why the seam matters in production: the manager can be upgraded under running workloads, a runtime can be swapped per workload class, and two containers from one image can be fenced differently on the same host. Name what the runtime does last and what it does afterwards.
Weigh the split as an interface decision: a narrow, documented hand-off is what lets a fleet change its container stack without touching a single artifact, and what keeps the privileged part of the stack small. The cost is more moving parts to version and observe on every host.
Starting a container looks like one action from the outside and is two jobs with a precise hand-off between them. Knowing where the seam is explains why the same artifact runs on hosts whose tooling has nothing in common, why the component you restart is usually not the one your process is attached to, and why the fencing a container gets is not a property of its image. ## What the container manager does The **container manager** is the component a human or a scheduler actually asks for a container. Before any of your code runs, it: - **Resolves the reference** you gave it to a concrete image — a manifest listing an ordered set of content-addressed layers plus a configuration object. - **Pulls and verifies** the layers it does not already hold locally, checking each one against the digest the manifest claims for it. - **Unpacks** those layers, in order, into a single directory tree: the container's **root filesystem**. This is a plain directory on the host, not a magic object — a later layer's contents land on top of an earlier layer's. - **Assembles a configuration document**: one description of the process to run, the mounts to attach, the views to create, the ceilings to apply, the user to run as, and the privilege set to keep. It builds this by taking the image's own recorded defaults and overlaying whatever the caller asked for, with the caller winning on conflicts. - **Prepares everything that lives outside the boundary**: the network attachment, the host paths and managed volumes to be mounted, the accounting and limiting group the process will be placed in. At the end of all that, the manager holds a directory and a document. That pair is the whole of what it can hand on. ## What the low-level runtime does The **low-level runtime** is a small program that is handed exactly that pair and asked to turn it into a running, fenced process. Reading only the document, it: 1. Creates the per-resource views the document asks for — a separate view of the filesystem, the process table, the network, the user ids and the hostname. 2. Places the new process in the accounting and limiting group so its consumption is counted and capped. 3. Attaches the listed mounts and pivots onto the unpacked root filesystem, so that tree becomes what the process sees as `/`. 4. Drops to the declared user, trims the privilege set, and applies any confinement profile the document names. 5. Executes the declared command. That command becomes the **first process inside the boundary** — it is what receives the termination signal later, and its exit status is the container's exit status. In the common design the runtime process then exits: its job was to build the boundary and hand it to your program, not to babysit it. Designs genuinely differ here — some invocations keep the runtime in the foreground as the container's parent instead — but where it exits, something else has to hold the workload's output stream and collect its exit status, which is why these stacks leave a small supervising process behind. ## The hand-off at a glance | Concern | Container manager | Low-level runtime | |---|---|---| | Talks to a registry | Yes — pulls and verifies layers | No — never sees a registry | | Sees the image format | Yes — manifest, layers, config | No — only an unpacked tree | | Writes the configuration document | Yes, merging image defaults with the request | No — it only reads it | | Creates the views and applies the ceilings | No | Yes | | Executes your command | No | Yes | | Lifetime | Long-lived, or a short command that leaves state behind | Typically exits once the process is running | ## Why the split is worth having - **Replaceability.** Because the interface between the two halves is a directory plus a document, either half can be swapped: a different manager can drive the same runtime, and a different runtime — including a sandboxed one — can be dropped into the same slot without touching the image. - **Restartability.** The heavy, frequently-upgraded component is the manager. Keeping it out of the workload's process lineage is what makes upgrading it a non-event for running containers. - **Scope of privilege.** The part that needs elevated privilege to build a boundary is small, short-lived and auditable, rather than being fused into the large component that also does networking and registry traffic. ## Where this is commonly misread The most frequent error is imagining a single daemon that does everything from pull to fencing; the second is thinking the fences are baked into the image. They are not — an image records defaults such as the command, the environment and the user, and nothing about which views to create or what the ceilings are. Those are decided per start, by whoever requested it, and written into the configuration document. Two containers from the same image on the same host can therefore be fenced completely differently, and that is the normal case rather than an exception.
- Where do the image's own defaults, such as the declared command and user, enter this?The image carries a configuration object recording them. The manager reads it and folds it into the runtime configuration document, letting the caller's overrides win. The low-level runtime never reads the image at all — it only ever sees the document and the unpacked tree, which is why an override and an image default are indistinguishable by the time it runs.
- Does the container manager have to be a long-running daemon?No, and designs differ. Some stacks run a persistent service that owns every container on the host; others assemble the bundle in a short-lived command that exits, leaving only a small supervising process behind. The division of work is the same either way — pull, unpack, write the document, then invoke a runtime. What differs is whether anything stays resident to be asked about the container afterwards.
saying these in an interview costs you the question
- Thinks one daemon does everything from pulling the image to applying the fences
- Says the low-level runtime downloads and unpacks the image itself
- Assumes every container must die when its manager restarts
- Calls the unpacked root filesystem the image, as if they were the same object
- Thinks the fences and ceilings are baked into the image at build time