Because the OCI runtime specification is a written contract, runc can be replaced with alternative runtimes such as crun, gVisor's runsc, or Kata Containers. What does that swap actually change, and what stays identical?
answer
- contract = bundle + create/start/kill/delete
- crun = same model, faster C
- runsc = user-space kernel, small host syscall surface
- Kata = per-container lightweight VM
- spec fixes interface, not behaviour or perf
basics
~20 sEverything above the runtime is unchanged: same images, same registry, same manager, same config.json. What changes is how the container is isolated and executed — crun is a faster C reimplementation using the same kernel primitives, runsc intercepts syscalls in a user-space kernel, Kata boots a lightweight VM. Isolation strength and compatibility/performance trade off.
solid answer
~50 sThe contract is narrow: a runtime takes a **bundle** (rootfs + `config.json`) and implements `create`/`start`/`state`/`kill`/`delete`. Anything honouring it drops in. - **runc** — the reference: Linux namespaces, cgroups, seccomp, capabilities. Shared kernel. - **crun** — same model, written in C; lower memory and faster start, better cgroup v2 support. A pure drop-in. - **runsc (gVisor)** — a user-space kernel intercepts syscalls, so the container rarely touches the host kernel directly. Much smaller kernel attack surface; syscall-heavy and I/O-heavy workloads slow down, and some syscalls are unimplemented. - **Kata Containers** — each container/pod runs in a lightweight VM with its own kernel. Hardware-level isolation, higher start latency and memory overhead, needs virtualization support. Identical across all of them: the image, the registry, the build, the manager (containerd/CRI-O), the `config.json`. What differs is isolation boundary, startup cost, and how faithfully host kernel features are exposed. You select per workload — `--runtime=` in Docker, a containerd runtime handler in Kubernetes.
code
json · 7 lines{
"default-runtime": "runc",
"runtimes": {
"runsc": { "path": "/usr/local/bin/runsc" },
"crun": { "path": "/usr/bin/crun" }
}
}go deeper
Know that runc is one implementation of a written spec and that other runtimes exist which run the same images.
Describe the isolation mechanism of each option and that images, registry, manager and config.json stay unchanged across the swap.
Own the trade-off: name the workloads that justify a sandboxed runtime, the compatibility and throughput costs, and how you'd select per workload and validate with benchmarks.
Frame it as a tenancy and risk decision across a fleet — where the isolation boundary must sit, what mixing runtimes costs in operations and observability, and how you'd stage adoption.
## What the contract actually is The OCI runtime-spec pins down two things: the **bundle** — a directory with `config.json` and an unpacked `rootfs/` — and the **lifecycle operations** a runtime must expose: `create`, `start`, `state`, `kill`, `delete`. That is the entire surface a container manager depends on. Everything else about a runtime is an implementation choice, and that is precisely what makes runtimes swappable. ## The runtimes people actually name **runc** — the reference implementation, written in Go, donated by Docker. It creates Linux namespaces, writes cgroup limits, applies seccomp filters, capabilities, and LSM labels, pivots into the rootfs and execs the entrypoint. The container's processes are ordinary host processes sharing the host kernel. **crun** — a C implementation from the Podman ecosystem. Same isolation model and the same kernel primitives, but a smaller binary, noticeably lower memory per container and faster start; it also tracked cgroup v2 and OCI features earlier. It is a genuine drop-in: the behaviour a workload sees is the same. **runsc (gVisor)** — implements the same OCI verbs but *not* the same isolation model. A user-space process called the Sentry implements a large subset of the Linux syscall surface itself; the container's syscalls are intercepted (via ptrace or KVM) and handled by the Sentry, which makes only a small, tightly-restricted set of real calls to the host kernel. The host kernel attack surface drops dramatically, which is the point: kernel LPE bugs that would break out of a runc container often cannot be reached. Costs: syscall-heavy and network/file-I/O-heavy workloads see meaningful slowdowns, and unimplemented or subtly different syscalls break some software (certain databases, anything using exotic ioctls or direct hardware access). **Kata Containers** — runs each container (or Kubernetes pod) inside a lightweight virtual machine with its own guest kernel, using a hardware virtualization boundary. Isolation approaches VM strength. Costs: hundreds of milliseconds of extra start latency, a per-VM memory floor, nested virtualization required if you are already on VMs, and awkwardness with host device access and some volume types. There are others in the same family — Firecracker-based sandboxes, `youki` (Rust), Windows and WASM runtimes fronted by shims. ## What is genuinely unchanged - **The image.** Runtimes do not consume images; the manager unpacks them. The same OCI image runs under all of these. - **The registry and the build.** Nothing changes. - **The manager.** containerd or CRI-O still handles pull, snapshotting, the shim and lifecycle bookkeeping. - **`config.json`.** The same declarative document is passed down; every runtime is expected to honour the process, mounts, resources and security fields it can. ## What is not identical, despite the spec This is where senior candidates separate themselves. Compliance guarantees the *interface*, not equivalent *behaviour*: - **Kernel fidelity.** Under gVisor the guest kernel is a reimplementation; `/proc` contents, unusual syscalls, and kernel-version-specific behaviour differ. Under Kata the guest kernel is a real but *different* kernel from the host's, so host module or sysctl assumptions break. - **Performance profile.** Start latency, syscall cost, network throughput and memory floor all shift, sometimes by an order of magnitude for the wrong workload. - **Feature support.** Privileged containers, host networking, host PID namespace, device passthrough, and some volume drivers are limited or unavailable in sandboxed runtimes — exactly the features that leak host access, so the limitation is often intentional. - **Observability.** Host-level tools (perf, eBPF probes, ps on the host) see sandboxed containers differently or not at all. ## How you select one Docker: `dockerd` is configured with named runtimes in `/etc/docker/daemon.json`, then `docker run --runtime=runsc …`. containerd exposes **runtime handlers**; Kubernetes surfaces them as `RuntimeClass` objects that a pod selects by name — the details of scheduling that belong to the orchestration side, but the mechanism is the same OCI swap underneath. ## The judgment answer Use the sandboxed runtime where the threat model demands it and the workload tolerates it: untrusted or multi-tenant user code, CI runners executing arbitrary builds, third-party plugins, functions-as-a-service. Keep runc or crun for trusted first-party services where a kernel-shared boundary plus a tight seccomp profile, dropped capabilities, read-only rootfs and a user namespace is proportionate. Mixing runtimes in one cluster is normal and is the usual answer — you pay the isolation tax only where it buys something. Always benchmark the specific workload before committing; the overhead is workload-shaped, not a single number.
- If gVisor is so much safer, why not run everything under it?Because the safety comes from interposing a user-space kernel on every syscall, which costs throughput on syscall-, network- and file-I/O-heavy workloads, and because its syscall surface is a subset — some software simply does not run, and host-level features like device passthrough or privileged mode are restricted. For trusted first-party services, runc with dropped capabilities, seccomp, a read-only rootfs and a user namespace is usually proportionate at far lower cost.
- What would make you reach for Kata Containers over gVisor?When you need a hardware-enforced boundary and full kernel compatibility at the same time: the workload requires syscalls or kernel behaviour gVisor does not faithfully implement, or a compliance requirement asks for VM-grade isolation between tenants. You pay for it with startup latency in the hundreds of milliseconds, a per-VM memory floor, and a dependency on hardware virtualization being available.
The spec is like a standard electrical socket: any appliance that fits will draw power, but a plug adapter does not make a hair dryer and a server rack interchangeable in what they cost or how safely they run.
saying these in an interview costs you the question
- Claiming OCI compliance means identical behaviour and performance — it only guarantees the interface.
- Thinking swapping the runtime requires rebuilding images or changing the registry.
- Describing gVisor as "a VM" or Kata as "just a stricter runc" — the isolation mechanisms are distinct.
- Assuming a sandboxed runtime is a free upgrade with no compatibility or throughput cost.
- Believing the runtime is what pulls the image or enforces the image's security scan policy.