skip to content

What exactly does an OCI-compliant runtime such as runc receive as its input, and what are the main sections of the config.json it reads?

level: middleimportance: should knowfreq 40%

answer

  1. bundle = config.json + rootfs/
  2. process · root · mounts · linux · hooks
  3. namespaces: no path = create, path = join
  4. maskedPaths / readonlyPaths guard /proc and /sys
  5. create → start → kill → delete

basics

~20 s

It receives a filesystem bundle: a directory containing an already-unpacked rootfs and a config.json. config.json declares the process (argv, env, user, capabilities), the root filesystem path, mounts, and a Linux section listing namespaces, cgroup resource limits, seccomp, LSM labels and masked paths. The runtime just applies it.

solid answer

~40 s

runc's input is a **filesystem bundle** — a directory with `rootfs/` (already unpacked from the image) and `config.json`. It never sees an image or a registry. `config.json` sections you should be able to name: - `ociVersion` - `process` — `args`, `env`, `cwd`, `user` (uid/gid/additionalGids), `capabilities` sets, `rlimits`, `noNewPrivileges`, `terminal` - `root` — `path` to the rootfs and `readonly` - `hostname`, `mounts` — `/proc`, `/sys`, `/dev`, tmpfs, bind mounts, each with source, destination, type and options - `linux` — `namespaces` to create or join, `resources` (cgroup memory/cpu/pids/io limits), `seccomp`, `apparmorProfile` / `selinuxLabel`, `maskedPaths` and `readonlyPaths`, `uidMappings`/`gidMappings` for user namespaces, `devices`, `cgroupsPath` - `hooks` — createRuntime, startContainer, poststart, poststop Lifecycle verbs are fixed too: `create`, `start`, `state`, `kill`, `delete` — with `create` doing the namespace/cgroup setup and `start` releasing the process to exec.

code

bash · 5 lines
bash
mkdir -p /tmp/bundle/rootfs && cd /tmp/bundle
runc spec
docker export $(docker create alpine:3.20) | tar -C rootfs -xf -
runc run demo
runc state demo

go deeper

for a junior

Know that the runtime is handed a directory containing a rootfs and a config.json and simply applies what that file says.

for a middle

Name the main sections and what each configures, and explain that namespaces without a path are created while namespaces with a path are joined.

for a senior

Use the generated config.json as the debugging source of truth for security posture — capabilities, seccomp, read-only root, user, limits — rather than trusting CLI flags or the Dockerfile.

for a principal

Treat config.json as the enforcement point for a platform-wide container baseline, and reason about who owns generating it and how hooks and annotations extend it.

## The bundle: the runtime's only input The OCI runtime-spec defines a **filesystem bundle** — an ordinary directory containing exactly two things: ``` bundle/ config.json rootfs/ ``` `rootfs/` is the container's root filesystem, *already assembled*. Someone else — containerd's snapshotter, CRI-O, Podman — pulled the image, unpacked the layers, and mounted or materialized the merged view. `config.json` is the complete declarative description of the container to create. That is the whole interface. A runtime does not know what a registry is, what a layer is, or what a tag is. You can see this concretely: `runc spec` writes a default `config.json` in the current directory, and `runc run <id>` with a populated `rootfs/` next to it starts a container with no Docker anywhere in the picture. ## The sections of config.json **`ociVersion`** — the spec version this document conforms to. Runtimes reject bundles they cannot support. **`process`** — everything about the initial process: `args` (argv, already resolved — the runtime does not merge ENTRYPOINT and CMD, the container manager did that), `env`, `cwd`, `terminal`, `user` with `uid`, `gid` and `additionalGids`, `rlimits`, `noNewPrivileges`, and the Linux **capabilities** sets (`bounding`, `effective`, `permitted`, `inheritable`, `ambient`). This is where dropped capabilities actually get expressed. **`root`** — `path` (usually `rootfs`) and `readonly`. A read-only root container is simply `root.readonly: true` plus writable mounts where the app needs them. **`mounts`** — an ordered list, each with `destination`, `type`, `source`, `options`. Every container gets `proc` on `/proc`, `sysfs` or a bind on `/sys`, a `tmpfs` on `/dev`, `devpts`, `mqueue`, `shm`. Volumes and bind mounts requested by the user appear here as additional entries. Mount order matters because later mounts can land on paths created by earlier ones. **`hostname`** and, in newer versions, `domainname` — applied inside the UTS namespace. **`linux`** — the platform-specific block: - `namespaces` — an array of `{type, path}`. Omitting `path` means "create a new one"; supplying a path means "join this existing namespace". Joining is exactly how sidecar containers share a network namespace with another container. - `resources` — the cgroup limits: `memory.limit`, `cpu.shares`/`quota`/`period`/`cpus`, `pids.limit`, `blockIO` weights, hugepages. `cgroupsPath` says where in the cgroup hierarchy to place them. - `seccomp` — the syscall filter: a `defaultAction` (typically `SCMP_ACT_ERRNO`) plus per-syscall allow rules. - `apparmorProfile`, `selinuxLabel` — the LSM confinement to apply. - `maskedPaths` / `readonlyPaths` — kernel interfaces such as `/proc/kcore` and `/sys/firmware` bind-mounted over `/dev/null` or remounted read-only so a container cannot read or poke host internals. - `uidMappings` / `gidMappings` — user-namespace ID mapping (rootless and userns-remapped containers). - `devices` — device nodes to create, plus the cgroup device allow/deny rules. - `rootfsPropagation`, `sysctl`, `intelRdt`. **`hooks`** — programs the runtime invokes at defined points: `createRuntime` and `createContainer` (after namespaces exist, before the user process runs — where network plugins traditionally attach an interface), `startContainer`, `poststart`, `poststop`. **`annotations`** — free-form key/value metadata passed through from the caller; how a manager such as CRI-O smuggles higher-level context (pod name, sandbox ID) to the runtime and to hooks. ## The lifecycle The spec also fixes the verbs: `create` (set up namespaces, cgroups, mounts, and block just before exec), `start` (release the process to `exec` the entrypoint), `state` (report status and PID as JSON), `kill` (deliver a signal), `delete` (tear down). The two-phase create/start split exists so the manager can attach networking, set up cgroups or run hooks in the fully-prepared-but-not-yet-running container. `runc run` is just create-then-start for convenience. ## Why the split from the image config matters The image's own config (Entrypoint, Cmd, Env, User) is *input* to producing this document, not the document itself. The container manager merges image defaults with the caller's overrides — `-e`, `-u`, `--memory`, `-v`, `--cap-drop`, `--read-only` — into one flat `config.json`. So when you debug "why did my container run as root" or "why wasn't the memory limit applied", the authoritative artifact is the generated `config.json`, not the Dockerfile.

  • Why does the runtime-spec separate `create` from `start` instead of having a single run operation?
    After `create` the container's namespaces, cgroups and mounts exist but the user process has not been exec'd yet. That window lets the container manager attach a network interface to the new network namespace, adjust cgroups, or run createRuntime hooks against a fully-prepared container. `start` then releases the blocked init to exec the entrypoint.
  • How would one container be made to share another container's network namespace through config.json?
    Its `linux.namespaces` entry for `type: network` carries a `path` pointing at the other container's namespace file (for example `/proc/<pid>/ns/net`) instead of being pathless. A pathless entry means "create a new namespace"; a path means "join this existing one". That is the mechanism behind sidecars and Kubernetes pod networking.

saying these in an interview costs you the question

  • Thinking runc pulls or unpacks the image — the rootfs is already assembled when runc is invoked.
  • Saying config.json is the image config; it is a separate document generated by merging image defaults with runtime options.
  • Believing runc stays alive as the container's parent — it exits after the container is created and started.
  • Assuming cgroup limits are enforced by runc itself rather than by the kernel via the cgroup settings runc writes.
  • Ignoring maskedPaths/readonlyPaths and claiming /proc is fully exposed in a default container.

context