Describe the Container Network Interface (CNI) contract used by Kubernetes: which component invokes the plugin, at what point in pod startup, what input it receives, and what it must return.
answer
- exec contract, not a daemon
- runtime → sandbox netns → CNI ADD
- /etc/cni/net.d config, /opt/cni/bin binaries
- env vars + stdin config → stdout Result JSON
- conflist chain; DEL must be idempotent
basics
~20 sWhen the CRI runtime creates a pod sandbox it runs a CNI plugin binary from /opt/cni/bin, using the JSON config in /etc/cni/net.d. It passes the command (ADD/DEL/CHECK) and the sandbox network namespace via environment variables plus config on stdin, and expects a JSON result with the assigned interface, IPs and routes on stdout.
solid answer
~50 sCNI is a thin exec-based contract, not a service. The kubelet asks the CRI runtime (containerd or CRI-O) to create a **pod sandbox**; the runtime creates an empty network namespace and then invokes CNI. It reads the lexically first config in `/etc/cni/net.d/*.conflist`, finds the named plugin binaries in `/opt/cni/bin`, and executes them. Input arrives two ways: environment variables (`CNI_COMMAND`, `CNI_CONTAINERID`, `CNI_NETNS`, `CNI_IFNAME`, `CNI_ARGS`, `CNI_PATH`) and the network config JSON on **stdin**. The plugin wires the namespace — typically a veth pair — calls its IPAM plugin for an address, and prints a JSON result on stdout listing interfaces, IPs, routes and DNS. A non-zero exit with a JSON error means failure. A `conflist` is a chain: each plugin runs in order (for example bridge, then portmap, then bandwidth). `DEL` must be idempotent, and `ADD` is called once per sandbox, so all containers inherit the result.
code
bash · 7 linesCNI_COMMAND=ADD \
CNI_CONTAINERID=8f3c1d... \
CNI_NETNS=/var/run/netns/cni-2a7f \
CNI_IFNAME=eth0 \
CNI_ARGS=K8S_POD_NAMESPACE=prod;K8S_POD_NAME=api-0 \
CNI_PATH=/opt/cni/bin \
/opt/cni/bin/bridge < /etc/cni/net.d/10-bridge.conflistgo deeper
Know that a plugin binary is executed when a pod is created and that it is what gives the pod its interface and IP; naming the two directories is already a good answer.
Give the full flow with the inputs and outputs — env vars plus stdin config, JSON Result on stdout — and explain the chain and the ADD/DEL commands.
Connect the contract to operational symptoms: ContainerCreating on one node, leaked IPAM entries from non-idempotent DEL, node NotReady until the plugin DaemonSet writes its config.
Discuss why an exec contract was chosen — swappable datapaths, no privileged API dependency — and the cost: per-pod fork latency, node-local state to reconcile, and upgrade coupling between the DaemonSet, binaries and spec version.
## Where CNI sits CNI is deliberately minimal: a specification for **executables**, not a daemon or an API server. That is why so many implementations exist — a plugin is any binary that reads a config and configures a namespace. The call chain during pod startup is: 1. The kubelet decides to run a pod and calls the CRI runtime (`containerd`, `CRI-O`) with `RunPodSandbox`. 2. The runtime creates the sandbox container and its **empty network namespace**. 3. The runtime — not the kubelet — invokes CNI against that namespace. 4. Only after CNI returns successfully does the runtime start the pod's application containers. That ordering explains the most common symptom: if CNI fails, the pod never leaves `ContainerCreating`, and the event text is a `failed to setup network for sandbox` message. No application container has started, so container logs are empty. ## Discovery: config and binaries Two directories matter, both on every node: - `/etc/cni/net.d/` — network configuration files. The runtime picks the lexically first valid `*.conflist` (or legacy `*.conf`). Plugins usually ship a DaemonSet that writes this file on startup, which is why a node can be `NotReady` with "network plugin not ready" until that pod lands. - `/opt/cni/bin/` — the plugin executables named by the config, plus the reference plugins (`bridge`, `host-local`, `loopback`, `portmap`, `bandwidth`). ## The invocation The runtime forks the binary and passes: - **Environment variables**: `CNI_COMMAND` (`ADD`, `DEL`, `CHECK`, `VERSION`, and in newer spec versions `GC` and `STATUS`), `CNI_CONTAINERID` (the sandbox ID), `CNI_NETNS` (path to the network namespace, e.g. `/var/run/netns/cni-…`), `CNI_IFNAME` (almost always `eth0` — the name *inside* the pod), `CNI_ARGS` (semicolon-separated key/values; Kubernetes passes `K8S_POD_NAMESPACE`, `K8S_POD_NAME`, `K8S_POD_INFRA_CONTAINER_ID`), and `CNI_PATH` (where to find other plugin binaries). - **stdin**: the network configuration JSON, including `cniVersion`, `name`, `type`, plugin-specific fields, an `ipam` block, and any `runtimeConfig`/`capabilities` the runtime injected (port mappings, bandwidth limits). The plugin returns on **stdout** a JSON *Result*: `cniVersion`, `interfaces`, `ips` (address, gateway, which interface), `routes`, and optionally `dns`. Failure is a non-zero exit code with a JSON error object (`code`, `msg`, `details`) — that message is what surfaces in the pod's events. ## What a plugin actually does on ADD A typical L3 plugin: create a veth pair; move one end into `CNI_NETNS` and rename it to `CNI_IFNAME`; call the configured **IPAM** plugin (a separate binary such as `host-local`, `dhcp`, or a vendor one like `calico-ipam`) to allocate an address from the node's slice of the cluster CIDR; set the address, default route and any extra routes inside the namespace; configure the host side (bridge port, host route, eBPF program); and return the Result. IPAM is a sub-contract with the same exec shape, so address management is swappable independently of the datapath. ## Chaining A `.conflist` holds a `plugins` array executed in order. The Result of each plugin is passed as `prevResult` to the next, so later plugins refine what earlier ones built: `bridge` creates connectivity, `portmap` adds host port mappings, `bandwidth` adds shaping, `tuning` sets sysctls. On `DEL` the chain runs in reverse. ## Delete and idempotency `DEL` is called when the sandbox is torn down and **must succeed even if it has already run or if the namespace is already gone** — the runtime retries. A plugin that errors on a missing namespace produces pods stuck in `Terminating` and, worse, leaked IP allocations: the IPAM store still shows an address in use for a container that no longer exists. That is the origin of the classic "node has free capacity but no more pod IPs" incident, cleared by garbage-collecting the IPAM store. ## Version notes worth stating Since the removal of dockershim, all supported Kubernetes versions go through CRI, so the runtime invokes CNI and the kubelet's old network-plugin flags are gone. Spec 1.0 dropped the old multi-version result plumbing; 1.1 added `GC` and `STATUS` so a runtime can ask a plugin to reconcile leaked resources rather than relying on `DEL` alone.
- A pod is stuck in Terminating and its node slowly runs out of pod IPs. How does the CNI contract explain that?Teardown calls the plugin with CNI_COMMAND=DEL, and DEL must be idempotent and tolerant of an already-missing namespace. If it errors, the runtime cannot complete sandbox removal, so the pod lingers in Terminating and the IPAM store keeps the address marked allocated. Over time the node's range fills with entries for containers that no longer exist, and new pods fail to get an IP until the store is garbage-collected.
- Why is IPAM a separate plugin rather than part of the network plugin?CNI splits the datapath from address management so the two can vary independently: the same bridge or routing plugin can use host-local ranges, DHCP, or a vendor IPAM that talks to a cloud API. The network plugin invokes the IPAM binary with the same exec contract and uses the returned addresses and routes. It also means IPAM state — and its failure modes, like exhaustion or leaks — lives in its own store you can inspect.
saying these in an interview costs you the question
- Saying the kubelet calls CNI directly — the CRI runtime does, on sandbox creation
- Thinking CNI is a long-running daemon or gRPC service rather than an exec'd binary
- Believing CNI runs once per container in the pod instead of once per sandbox
- Claiming CNI configures Services or kube-proxy rules — it only wires pod interfaces, addresses and routes
- Assuming DEL failures are harmless rather than the cause of leaked IPs and stuck Terminating pods