skip to content

What is cri-dockerd, and what does keeping Docker Engine as the node runtime cost?

level: seniorimportance: should knowfreq 34%

answer

  1. Dockershim's code, moved out of the tree
  2. An adapter daemon in front of the engine
  3. One more hop, one more thing to patch
  4. Point the runtime endpoint at a different socket

basics

~20 s

cri-dockerd is the out-of-tree adapter that implements the CRI services and forwards them to the Docker Engine API - dockershim's code as a standalone daemon. It costs an extra hop, another component to patch and version-match, and cgroup-driver alignment.

solid answer

~50 s

Since Kubernetes 1.24 the node agent speaks only CRI, and Docker Engine does not serve that API. `cri-dockerd`, maintained outside Kubernetes by Mirantis, is the old dockershim logic packaged as its own daemon: it listens on `unix:///var/run/cri-dockerd.sock`, implements RuntimeService and ImageService, and re-issues every call against the Docker Engine API. Point the node agent's `--container-runtime-endpoint` at it and Docker Engine is your runtime again. The price is real: the chain becomes node agent, cri-dockerd, dockerd, containerd, shim, runc instead of node agent, containerd, shim, runc; you patch and version-match one more daemon that now sits in the pod-start path; the cgroup driver must agree across all three components; the pod sandbox is emulated with a `pause` container; and new CRI features reach containerd and CRI-O first. Use it only for a real on-node dependency on the Docker Engine API, ideally with a removal date.

code

bash · 8 lines
bash
# Where the node agent looks for its CRI runtime
# containerd, the default:
#   --container-runtime-endpoint=unix:///run/containerd/containerd.sock
# Docker Engine, via the adapter:
#   --container-runtime-endpoint=unix:///var/run/cri-dockerd.sock

systemctl status cri-docker.service
crictl --runtime-endpoint unix:///var/run/cri-dockerd.sock ps

go deeper

for a junior

Recall the shape rather than the detail: after the dockershim removal the node agent speaks only CRI, and cri-dockerd is a separate piece of software you install if you still want Docker Engine to be the runtime.

for a middle

Explain the mechanics: the adapter implements the two CRI services on its own socket and forwards to the Docker Engine API, and the node agent's runtime endpoint has to be pointed at it. Be able to draw both chains and say which layers each one adds.

for a senior

Demonstrate the operational judgement: an inventory of what genuinely needs the Docker Engine API, the extra patching and cgroup-driver alignment you take on, and a migration that moves node pools rather than the whole fleet at once.

for a principal

Own the strategy: adapters are bridges with expiry dates, not architecture. Decide whether the fleet standardises on the ecosystem's default runtime, what a vendor support contract is genuinely buying, and how the platform avoids re-acquiring this kind of translation layer next time.

## Where cri-dockerd fits Since Kubernetes 1.24 the node agent speaks only CRI. It connects to whatever unix socket `--container-runtime-endpoint` names and expects to find RuntimeService and ImageService there. Docker Engine serves neither - it serves its own HTTP API on `/var/run/docker.sock`. `cri-dockerd` is the dockershim code lifted out of the Kubernetes tree and shipped as a standalone daemon, maintained outside Kubernetes by Mirantis. It listens on its own socket (by default `unix:///var/run/cri-dockerd.sock`), implements the two CRI services, and re-issues each call against the Docker Engine API. You install it as a system service, point the node agent's runtime endpoint at it, and the node once again runs Docker Engine as its container runtime. The chain becomes: ``` node agent -> cri-dockerd -> dockerd -> containerd -> shim -> runc ``` against the default: ``` node agent -> containerd -> shim -> runc ``` ## What it costs **Two extra processes in the pod-start path.** Every sandbox creation, container start, image pull and status poll crosses two more process boundaries and one more API translation. The latency is small; the failure surface is not - dockerd being wedged now stops pods from starting, where on a containerd node it would be irrelevant. **One more component to patch and version-match.** You now track kubelet, cri-dockerd and Docker Engine versions together, and a Docker Engine CVE sits directly in your cluster's container-start path rather than in an optional build tool. **Configuration that must agree.** The cgroup driver has to be the same choice across the node agent, cri-dockerd and Docker Engine; a mismatch gives you nodes that look healthy at first and misbehave under memory pressure. Adding a second place where that choice is expressed is a real source of node-level drift. **Emulated concepts.** CRI's pod sandbox has no native equivalent in Docker Engine, so cri-dockerd recreates the dockershim trick: a `pause` container the pod's containers share namespaces with. Container statuses and log paths are likewise translated rather than native, which matters to anything on the node that parses them. **Feature lag.** New CRI capabilities land in containerd and CRI-O first, because that is where the ecosystem's work happens; the adapter follows later, if at all. **Node footprint.** dockerd resident on every node, forever, for a translation layer. ## What it buys There is exactly one good reason: something on the node genuinely needs the Docker Engine API and cannot be replaced yet. A vendor-supported Docker Engine stack with a support contract behind it is a second. A third is scheduling: a phased migration where you cannot re-provision nodes in the same window as the control-plane upgrade, and want cri-dockerd as a temporary bridge with a removal date attached. Note what is *not* a reason: building images. Building has nothing to do with which runtime the node agent drives. You can install Docker Engine or buildx on a node purely to build, while the node agent drives containerd - and better still, build in CI and push to a registry. ## A concrete decision A 23-node cluster runs a chat-message fan-out service, a static Go binary in a scratch image, behind an upgrade past 1.24. The team's plan is cri-dockerd on all 23 nodes 'so nothing changes'. The right first move is an inventory of what actually calls the Docker Engine API on those nodes. In practice it is usually two things: a log shipper reading the `json-file` layout, and a node agent shelling out to `docker inspect`. Both have CRI-native equivalents, and replacing them is a day of work against a permanent second daemon to secure, patch and reason about. If one genuine dependency survives the inventory, run cri-dockerd on the small subset of nodes that need it and keep the rest on containerd, rather than degrading the whole fleet to the slowest common denominator. ## What does not change Termination semantics are identical. The node agent's grace period is carried into CRI's `StopContainer` as a timeout, so a service configured with a 4-second grace period still gets SIGTERM in PID 1 and SIGKILL when the timeout expires, on either path. The adapter adds hops, not different behaviour - and the images, as always, are the same images.

  • Do you need cri-dockerd if your team wants to keep building images with Docker on those machines?
    No, and this is the most common wasted install. Building is unrelated to which runtime the node agent drives: you can have Docker Engine or buildx present purely to build while the node agent talks to containerd. Better still, move builds out of the cluster into CI so nodes only ever pull images, which removes the argument entirely.
  • What has to line up between the node agent, cri-dockerd and Docker Engine on a node?
    The runtime endpoint the node agent is pointed at must be the adapter's socket rather than containerd's; the cgroup driver must be the same choice in all three; and the three versions must be patched together, since the adapter tracks both the CRI version above it and the Docker Engine API below it. Any drift shows up as nodes that register fine and then misbehave.
  • How would you decide to remove cri-dockerd from an existing fleet?
    Inventory what actually calls the Docker Engine API on those nodes - usually a log shipper reading the json-file layout and an agent shelling out to `docker inspect` - and check each for a CRI-native equivalent. Migrate those, then move nodes to containerd in a rolling fashion, one pool at a time, with the old runtime still installable as a rollback. Container behaviour and images do not change, so the blast radius is node-level, not workload-level.

saying these in an interview costs you the question

  • Says cri-dockerd is maintained in-tree by Kubernetes
  • Claims it adds no hops because Docker is the runtime anyway
  • Thinks containerd nodes cannot run Docker-built images
  • Assumes cri-dockerd is the default path after 1.24
  • Installs it just to keep `docker build` working on nodes
  • Ignores that dockerd is now in the pod-start failure path

context