skip to content

Kernel Isolation & Runtimes

What the kernel does when you start a container: the namespaces that isolate it, the cgroups that bound it, and the dockerd, containerd, shim and runc chain that launches it under the OCI specs. Asked to separate CLI users from people who can debug a runtime.

part ofDockeroverview, primer and where to startread it →
on this pageshow

questions

27

Kubernetes 1.24 removed dockershim - do images built with `docker build` still run?

level: juniorimportance: must knowfreq 68%

answer

  1. Only one component was actually deleted
  2. The image format is an open standard
  3. A translator inside the node agent
  4. Deprecated in 1.20, removed in 1.24

basics

~20 s

Yes. Dockershim was the Kubernetes node agent's built-in adapter for talking to Docker Engine, and only that adapter was removed. An image is a standard OCI artefact, so containerd and CRI-O pull and run exactly the same images.

solid answer

~40 s

Yes - nothing about the image changes. Dockershim was code inside the Kubernetes node agent that translated Container Runtime Interface (CRI) gRPC calls into Docker Engine API calls; Kubernetes deprecated it in 1.20 and deleted it in 1.24. A node loses that adapter, not the ability to run your images: containerd and CRI-O implement CRI directly and run the same standard image format that `docker build` and `docker push` produce. Dockerfiles, registries, tags and the container's runtime behaviour are untouched, because behaviour comes from the image config and the OCI runtime spec. What does change is node-level tooling: on a containerd node the `docker` CLI no longer lists the cluster's containers, so node debugging moves to `crictl`, and anything that bind-mounted `/var/run/docker.sock` or shelled out to `docker` on the node has to be replaced.

code

dockerfile · 8 lines
dockerfile
FROM golang:1.22 AS build
WORKDIR /src
COPY . .
RUN CGO_ENABLED=0 go build -o /out/fanout ./cmd/fanout

FROM scratch
COPY --from=build /out/fanout /fanout
ENTRYPOINT ["/fanout"]

go deeper

for a junior

Be ready to answer the panic version of this in one sentence: an adapter was removed, not your images. Recall that dockershim lived inside the Kubernetes node agent and that containerd and CRI-O run the same images your docker build produces.

for a middle

Explain what the adapter translated - CRI gRPC calls into Docker Engine API calls - and name the two CRI services. Expect a follow-up on what concretely changes on a node once it is gone, starting with docker ps no longer showing anything useful.

for a senior

Show you have run the migration: an inventory of workloads mounting the Docker socket, node agents that shell out to docker, log-shipping assumptions, and a node-by-node rollout with a way back. The interesting risk is never the image; it is the tooling around the node.

for a principal

Own the framing: an in-tree special case for one vendor's engine became a maintenance liability, and a published contract is what made the swap survivable. Argue how you keep your own platform's runtime layer replaceable so the next such change is a configuration change.

## The headline, and what it actually said In December 2020 the Kubernetes 1.20 release notes deprecated a component called *dockershim*. That was compressed online into 'Kubernetes is dropping Docker', and it has been an interview question ever since. The accurate statement is far narrower: one adapter inside one Kubernetes node component was deprecated, and in Kubernetes 1.24 (May 2022) its code was deleted. Nothing about how images are built, stored or shaped changed. ## What dockershim was The node agent that starts and stops containers does not link against any particular container engine. It speaks the **Container Runtime Interface (CRI)**, a gRPC contract with two services: - **RuntimeService** - pod sandbox and container lifecycle: `RunPodSandbox`, `CreateContainer`, `StartContainer`, `StopContainer`, `RemoveContainer`, `ListContainers`, plus `Exec`, `Attach`, `PortForward` and status calls. - **ImageService** - `PullImage`, `ListImages`, `ImageStatus`, `RemoveImage`, `ImageFsInfo`. Any runtime implementing those two services over a unix socket can be driven by the node agent. containerd does, through a built-in CRI plugin; CRI-O was written for nothing else. Docker Engine does not. It exposes its own HTTP API on `/var/run/docker.sock`, designed years before CRI existed. So Kubernetes shipped **dockershim**: a CRI implementation compiled into the node agent that accepted CRI gRPC calls and re-issued them as Docker Engine API calls. It also had to invent the concepts Docker Engine has no notion of - a CRI *pod sandbox* became a small `pause` container whose namespaces the pod's real containers joined. That made Docker Engine the one runtime with special, in-tree support maintained by the Kubernetes project itself, while every other runtime lived out of tree and was maintained by its own community. Add the fact that Docker Engine is itself a layer above containerd - so a Docker node meant node agent to dockershim to dockerd to containerd, where a containerd node meant node agent straight to containerd - and the shim was carrying maintenance cost for hops nobody needed. ## Why your images are untouched Building an image and running one are two different contracts, and only the running side changed. `docker build` produces an image - a manifest, a config and a set of layer blobs - in a format standardised under the Open Container Initiative and pushed over a standardised registry protocol. containerd and CRI-O consume exactly that. (Docker's own manifest media types predate the OCI ones; both are accepted.) A Dockerfile is an input to a *builder*, not a runtime instruction set, and once the image exists nothing in it records which engine built it. Concretely, all of this keeps working after the removal: - `docker build`, `docker buildx build`, your Dockerfiles and multi-stage builds - `docker push` / `docker pull`, Docker Hub, any private registry - Docker Desktop on developer laptops - the container's runtime behaviour - ENTRYPOINT/CMD, ENV, USER, WORKDIR, exposed ports, signal delivery - because that comes from the image config and the OCI runtime spec, and `runc` is the same `runc` on either path The one-line answer is: the image is a standard artefact, dockershim was a translator for one engine's API, and deleting the translator does not invalidate the artefact. ## What genuinely changed, on the node The disruption was operational, and it was real: 1. **A runtime migration.** Every node had to be provisioned with, or switched to, containerd or CRI-O, with the node agent pointed at that runtime's socket via `--container-runtime-endpoint`. 2. **`docker ps` stops being the node debugging tool.** On a containerd node there is no dockerd holding those containers. Even where Docker Engine is still installed for other reasons, its containers live in containerd's `moby` namespace while the cluster's live in `k8s.io`, so neither client lists the other's. Node debugging moves to `crictl`. 3. **Nothing can bind-mount `/var/run/docker.sock` any more.** Workloads that built images in-cluster, or agents that inspected containers through the Docker Engine API, lost the socket they mounted. Those builds move to a builder that does not need the engine, or out of the cluster entirely. 4. **Log and metric assumptions.** Tooling that read Docker's `json-file` log layout on the node, or scraped Docker's own stats endpoint, has to read what the CRI runtime and node agent produce instead. 5. **If you genuinely need Docker Engine as the runtime**, `cri-dockerd` is the out-of-tree adapter that puts the shim's logic back in front of it as a separate daemon. ## How to answer it Lead with 'no, the images are fine', give the reason in one clause - the image format is a standard the other runtimes implement - then show you know where the pain actually was: a node-by-node runtime migration, `docker ps` giving way to `crictl`, and every workload or agent that assumed a Docker socket on the node. A candidate who says 'we had to rebuild our images' or 'Dockerfiles are dead' has swallowed the headline instead of reading it.

  • If the images are unaffected, why was the removal disruptive at all?
    Because the disruption was on the node, not in the image. Every node needed a runtime migration to containerd or CRI-O; `docker ps` stopped showing the cluster's containers; workloads that bind-mounted `/var/run/docker.sock` for in-cluster builds lost their socket; and node agents that shelled out to `docker` or parsed Docker's json-file log layout had to be rewritten against CRI-native equivalents.
  • Does this mean Docker Engine cannot be installed on a Kubernetes node any more?
    No. You can install Docker Engine on a node for builds or legacy tooling; the node agent simply will not drive it unless cri-dockerd sits in front. The two coexist without seeing each other: Docker Engine's containers live in containerd's `moby` namespace, while the cluster's live in `k8s.io`, so neither `docker ps` nor `crictl ps` lists the other's containers.
  • Which services does the CRI contract define, and why does that matter here?
    Two gRPC services: RuntimeService, which covers pod sandbox and container lifecycle plus exec, attach and status, and ImageService, which covers pulling, listing, inspecting and removing images. It matters because dockershim's whole job was implementing those two services on top of Docker Engine's own API - and containerd and CRI-O implement them natively, so the adapter had nothing left to add.

Dockershim was a translator hired so one guest could speak to the host in his own language. Firing the translator does not change the letters everyone had already written - it just means the host now talks directly to guests who speak the standard tongue.

saying these in an interview costs you the question

  • Says Docker-built images must be rebuilt for containerd
  • Claims Dockerfiles are no longer supported by Kubernetes
  • Thinks Kubernetes uninstalled Docker Engine from nodes
  • Believes containerd and CRI-O use a different image format
  • Assumes `docker ps` still lists pod containers on a containerd node
  • Says the removal happened in 1.20 rather than the deprecation

context

open as a page

What are Linux namespaces, which kinds exist, and how do they make a container look like a separate machine?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Namespaces are a kernel feature that gives a process its own view of a global resource. Separate namespaces exist for process IDs, network stacks, mounts, hostname, IPC, users and cgroups. A container is just a process placed in a fresh set of them.

open as a page

The Open Container Initiative (OCI) publishes three specifications — image-spec, runtime-spec and distribution-spec. What does each one define, and where does each apply in the life of a container?

level: juniorimportance: must knowfreq 55%

basics

~20 s

image-spec defines the image format on disk and on the wire: layers, an image config, a manifest, and an index, all addressed by digest. distribution-spec defines the registry HTTP API used to push and pull those blobs. runtime-spec defines the unpacked bundle — a rootfs directory plus config.json — that a runtime such as runc executes.

open as a page

The `docker` command-line tool does not itself create containers. Describe the client/daemon architecture behind it and what actually happens when you run a docker command on a Linux host.

level: juniorimportance: must knowfreq 52%

basics

~20 s

The docker CLI is a thin HTTP client. It sends a request over the Unix socket /var/run/docker.sock to the long-running daemon dockerd, which does all the real work — pulling images, creating containers, managing networks — and streams results back. Containers are children of the daemon's stack, not of your shell.

open as a page

A colleague says "a Docker container is just a lightweight virtual machine." Explain what a container actually virtualizes compared with a hardware virtual machine running under a hypervisor, and why the distinction matters in practice.

level: juniorimportance: must knowfreq 88%

basics

~20 s

A virtual machine virtualizes hardware: the hypervisor gives each guest its own kernel and virtual devices. A container virtualizes the operating system: processes share the host kernel and are isolated by kernel features. A container is a process, not a machine.

open as a page

Explain the difference between the container options --cpus, --cpu-shares and --cpuset-cpus, and why an application can be slow while its CPU usage looks low.

level: middleimportance: must knowfreq 48%

basics

~20 s

--cpus is a hard quota per scheduling period (cpu.max), --cpu-shares is a relative weight that only matters under contention (cpu.weight), and --cpuset-cpus pins the container to specific cores. A container that burns its quota early is throttled until the next period, so latency spikes while average utilization looks modest.

open as a page

What are Linux control groups (cgroups), and what actually happens in the kernel when a container is started with `--memory=512m --cpus=1.5`?

level: middleimportance: must knowfreq 58%

basics

~20 s

cgroups are the kernel feature that limits and accounts resources for a group of processes. The runtime creates a cgroup per container and writes the flags into its control files: on cgroup v2, memory.max becomes 536870912 and cpu.max becomes '150000 100000' (150 ms of CPU per 100 ms period).

open as a page

Inside a container, your application runs as PID 1. What does the Linux kernel treat differently about PID 1, and what problems does that cause?

level: middleimportance: must knowfreq 62%

basics

~20 s

PID 1 is special: default signal handlers are ignored, so it will not die from SIGTERM unless it handles it, and it inherits orphaned children and must reap them. An app that does neither ignores graceful shutdown and accumulates zombie processes.

open as a page

Describe the objects that make up an OCI container image — what a manifest, an image config and a layer blob each contain, and how digests tie them together.

level: middleimportance: must knowfreq 50%

basics

~20 s

An image is a graph of digest-named blobs. The manifest lists a config descriptor and ordered layer descriptors. The config JSON holds runtime defaults (Entrypoint, Cmd, Env, User), build history and rootfs.diff_ids. Each layer blob is a tar of filesystem changes. The image's own ID is the digest of its config.

open as a page

On a Linux host running Docker Engine, walk through the chain of components involved in starting a container — from the daemon down to the process that becomes PID 1 inside the container — and say what each layer is responsible for.

level: middleimportance: must knowfreq 50%

basics

~20 s

dockerd takes the API request and handles images, networks and volumes, then asks containerd to run the container. containerd manages the snapshot and metadata and starts a containerd-shim per container. The shim invokes runc, which creates namespaces and cgroups and execs the entrypoint. runc then exits; the shim stays as the container's parent.

open as a page

Why does each running container get its own long-lived shim process sitting between the container manager and the container's init process, and what does that buy you when the container manager or Docker daemon is restarted?

level: middleimportance: must knowfreq 45%

basics

~20 s

The shim is the container's real parent. It keeps the stdio/TTY open, waits on the process so the exit code is never lost, and holds the container alive independently of the daemon. Because the daemon is not in the process tree, containerd or dockerd can restart or upgrade while containers keep running, then re-attach through the shims.

open as a page

Why does a typical container start in tens of milliseconds and consume roughly the memory of its own processes, while a comparable virtual machine takes tens of seconds and reserves hundreds of megabytes? Walk through where the time and the memory actually go.

level: middleimportance: must knowfreq 66%

basics

~20 s

A VM must emulate firmware, boot a kernel, probe devices and run an init system, and its guest kernel plus OS services permanently occupy RAM. A container skips all of that: the runtime sets up isolation and execs your process against the already-running host kernel.

open as a page

A container exits with code 137 in production. How do you confirm it was a cgroup memory kill rather than something else, and how do you find the cause?

level: seniorimportance: must knowfreq 54%

basics

~20 s

137 means the process died from SIGKILL (128+9), which can be a cgroup OOM kill or an external kill such as a stop timeout. Confirm with docker inspect .State.OOMKilled, the kernel log line 'Memory cgroup out of memory', and the cgroup's memory.events oom_kill counter, then compare peak usage against memory.max.

open as a page

If a process inside a Linux container exploits a kernel vulnerability, how does the blast radius compare with the same exploit fired inside a hardware virtual machine — and which container-level controls actually shrink it?

level: seniorimportance: must knowfreq 56%

basics

~20 s

In a container the kernel is the isolation boundary, so a kernel privilege-escalation bug means host compromise and every co-tenant container with it. In a VM the attacker owns only that guest and must also break the hypervisor. Shrink it by cutting syscall and capability reach.

open as a page

On a containerd Kubernetes node, what replaces `docker ps`, and how does crictl's view differ?

level: middleimportance: should knowfreq 44%

basics

~20 s

crictl replaces it. The docker CLI talks to dockerd, which is not the runtime on such a node. crictl speaks the CRI gRPC API instead and is sandbox-aware: crictl pods lists pod sandboxes, crictl ps lists the containers inside them.

open as a page

When you run a container with the default bridge network, what does the Linux kernel actually set up so it can reach the internet, and why can it still bind port 8080 while another container also uses 8080?

level: middleimportance: should knowfreq 55%

basics

~20 s

The container gets its own network namespace — a private stack with its own interfaces, routes, iptables rules and full port range. Docker creates a veth pair, puts one end inside as eth0 and attaches the other to the docker0 bridge, then NATs outbound traffic via the host.

open as a page

How does a single image tag serve both amd64 and arm64 hosts, and how does a client end up pulling the right variant?

level: middleimportance: should knowfreq 45%

basics

~20 s

The tag points at an OCI image index (manifest list) instead of a single manifest. The index lists one manifest per platform, each annotated with os and architecture. The client fetches the index, matches its own platform, then pulls only that manifest's config and layers.

open as a page

What exactly does an OCI-compliant runtime such as runc receive as its input, and what are the main sections of the config.json it reads?

level: middleimportance: should knowfreq 40%

basics

~20 s

It receives a filesystem bundle: a directory containing an already-unpacked rootfs and a config.json. config.json declares the process (argv, env, user, capabilities), the root filesystem path, mounts, and a Linux section listing namespaces, cgroup resource limits, seccomp, LSM labels and masked paths. The runtime just applies it.

open as a page

What changed between cgroup v1 and cgroup v2 on Linux, and how does that affect container resource limits?

level: seniorimportance: should knowfreq 38%

basics

~20 s

v1 had a separate hierarchy per controller with inconsistent interfaces; v2 has one unified hierarchy, one consistent API (memory.max, cpu.max, io.max), pressure metrics, and coordinated memory-plus-I/O accounting so buffered writes are attributed correctly. Some v1-only options, such as disabling the OOM killer, no longer exist.

open as a page

What is cri-dockerd, and what does keeping Docker Engine as the node runtime cost?

level: seniorimportance: should knowfreq 34%

basics

~20 s

cri-dockerd is the out-of-tree adapter that implements the CRI services and forwards them to the Docker Engine API - dockershim's code as a standalone daemon. It costs an extra hop, another component to patch and version-match, and cgroup-driver alignment.

open as a page

How can one container be made to share another container's network or process namespace, and which command lets you inspect a running container's namespaces from the host?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Namespace membership is per type, so containers can share some and not others: --network container:<name>, --pid container:<name>, --ipc container:<name>, or host for the host's own. From the host, nsenter -t <pid> -n -p -m joins a running container's namespaces via /proc/<pid>/ns/.

open as a page

containerd can be run on its own without Docker Engine. What is containerd responsible for by itself, and what is the Container Runtime Interface (CRI) that lets an orchestrator's node agent drive it directly?

level: seniorimportance: should knowfreq 42%

basics

~20 s

On its own, containerd handles image pull and storage, snapshots/rootfs preparation, container and task lifecycle via shims, and low-level networking hooks — exposed over a gRPC API on a Unix socket. CRI is a standard gRPC API (image and runtime services) that containerd implements as a plugin, so a node agent can drive it without Docker Engine.

open as a page

Explain what micro-VM and sandboxed runtimes such as Firecracker, Kata Containers and gVisor give you that a standard container started by runc does not, and what each one costs.

level: seniorimportance: should knowfreq 38%

basics

~20 s

They add a second isolation boundary under the container. Firecracker and Kata run each workload in a stripped-down VM with its own kernel and a tiny device model; gVisor keeps one host kernel but services most syscalls in a user-space kernel. Cost: some start-up, memory and feature loss.

open as a page

How would you choose memory and CPU limits for a latency-sensitive containerized service, given that exceeding the memory limit kills the process while exceeding the CPU quota only delays it?

level: principalimportance: should knowfreq 34%

basics

~20 s

Treat the two asymmetrically. Memory is a cliff: size it from observed peak plus headroom and cap the runtime below it so failures are diagnosable. CPU is a slope: quota only delays work, so set it above burst demand, size thread pools to it, and watch throttling counters instead of shaving it to average usage.

open as a page

You are designing a platform that will execute code submitted by untrusted customers. How would you decide between plain containers, micro-VMs, and full virtual machines as the isolation boundary, and what would you put around whichever you pick?

level: principalimportance: should knowfreq 40%

basics

~20 s

Decide from the threat model, not preference: untrusted code needs a boundary below the shared kernel, so micro-VMs are the default; plain containers only for trusted code; full VMs when you need coarse, long-lived, compliance-visible separation. Then add tenant-separated nodes, network egress control and hard resource caps.

open as a page

Because the OCI runtime specification is a written contract, runc can be replaced with alternative runtimes such as crun, gVisor's runsc, or Kata Containers. What does that swap actually change, and what stays identical?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Everything above the runtime is unchanged: same images, same registry, same manager, same config.json. What changes is how the container is isolated and executed — crun is a faster C reimplementation using the same kernel primitives, runsc intercepts syscalls in a user-space kernel, Kata boots a lightweight VM. Isolation strength and compatibility/performance trade off.

open as a page

The Docker daemon on a Linux host is unresponsive, yet the application containers on that host are still serving traffic. Explain why that is possible and how you would inspect and manage those containers while the daemon is down.

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Containers are parented by per-container shim processes, not by the daemon, so they keep running when it hangs. Inspect them one layer down with ctr -n moby containers ls / tasks ls, or from the OS with ps --forest, nsenter into the namespaces and the cgroup tree. Fix or restart the daemon; with live-restore it re-attaches without stopping anything.

open as a page