A Container Storage Interface (CSI) driver is usually deployed as both a Deployment and a DaemonSet. What does each part do, and which storage operations belong to which?
answer
- Controller = storage API, no filesystem
- Node = DaemonSet, privileged, mounts
- Stage once per node, publish once per pod
- Sidecars watch K8s, driver speaks CSI
- attachRequired: false for NFS-style
basics
~20 sThe controller plugin (a Deployment) does cluster-wide work against the storage API: create, delete, attach, detach, snapshot, expand. The node plugin (a DaemonSet on every node) does machine-local work: format, stage-mount and bind-mount the volume into the pod.
solid answer
~50 sCSI splits work by *where it must happen*. **Controller plugin** - runs anywhere, typically a Deployment with leader election and one or two replicas. It talks to the storage system's API and never touches a filesystem: `CreateVolume`, `DeleteVolume`, `ControllerPublishVolume`/`ControllerUnpublishVolume` (attach/detach a disk to a node), `CreateSnapshot`, `ControllerExpandVolume`. It runs alongside Kubernetes-maintained sidecars - `external-provisioner` watches PVCs, `external-attacher` watches VolumeAttachment objects, `external-snapshotter`, `external-resizer`. **Node plugin** - a DaemonSet, so a copy exists on every node that may run a pod using the driver. It is privileged, mounts the host's mount propagation paths, and implements `NodeStageVolume` (format if needed and mount once per node to a global staging directory), `NodePublishVolume` (bind-mount into the pod's volume directory), plus the unstage/unpublish teardown. The `node-driver-registrar` sidecar registers its socket with kubelet. Not every driver needs both halves. NFS-style drivers often have no attach step at all and set `attachRequired: false` in the CSIDriver object.
code
yaml · 9 linesapiVersion: storage.k8s.io/v1
kind: CSIDriver
metadata:
name: nfs.csi.k8s.io
spec:
attachRequired: false
podInfoOnMount: true
volumeLifecycleModes:
- Persistentgo deeper
Know there are two parts: one that asks the storage system for a disk, and one on each node that mounts it.
Name the RPCs on each side, the DaemonSet/Deployment shape, and at least two sidecars and what they watch.
Use the split as a triage map - provision, attach and mount failures each point to a different component and a different log - and mention blast radius and privilege.
Discuss it as capability negotiation and failure isolation: CSIDriver capability flags, controller HA and leader election, per-node attach limits, and the trust implications of privileged node plugins fleet-wide.
## Why there are two halves at all Storage work divides cleanly into two kinds: 1. Things done through the storage system's API, from anywhere with credentials and network access - "create me a 50 GiB disk", "attach disk vol-123 to instance i-abc". 2. Things that can only be done *on the machine where the container will run* - formatting the block device, mounting it, bind-mounting it into the container's directory. CSI names these the **Controller Service** and the **Node Service**, and Kubernetes deploys them accordingly. ## The controller plugin Usually a Deployment (or a small StatefulSet) with leader election, because these operations are cluster-wide and must not be duplicated. It holds the cloud or storage credentials - a good reason not to run it on every node. RPCs it implements: - `CreateVolume` / `DeleteVolume` - dynamic provisioning and reclaim. - `ControllerPublishVolume` / `ControllerUnpublishVolume` - attach and detach the volume to a node, meaningful for block storage such as EBS or a vSphere disk. - `CreateSnapshot` / `DeleteSnapshot` - snapshot support. - `ControllerExpandVolume` - grow the backing volume. - `ValidateVolumeCapabilities`, `GetCapacity`, `ListVolumes`. The controller plugin does not watch Kubernetes objects itself. That is the job of **sidecar containers** in the same pod, maintained by the Kubernetes storage community: - `external-provisioner` - watches unbound PVCs whose StorageClass names this driver, calls `CreateVolume`, and creates the PV object. - `external-attacher` - watches `VolumeAttachment` objects and calls `ControllerPublishVolume`. - `external-resizer` - watches PVC size edits and calls `ControllerExpandVolume`. - `external-snapshotter` - watches VolumeSnapshot objects and calls `CreateSnapshot`. - `livenessprobe` - exposes driver health. This sidecar pattern is why the vendor's binary can stay Kubernetes-agnostic: the sidecars speak Kubernetes, the driver speaks CSI. ## The node plugin A DaemonSet, privileged, with host paths mounted (`/var/lib/kubelet` with bidirectional mount propagation, and often `/dev`). Kubelet calls it directly over a UNIX socket; there is no API-server round trip in the mount path. RPCs: - `NodeStageVolume` - called **once per node per volume**. This is where a raw block device gets a filesystem (if it has none) and is mounted at a global staging path. Doing it once per node is what lets several pods share one volume on the same node. - `NodePublishVolume` - called **once per pod**. Bind-mounts (or for ephemeral inline drivers, directly mounts) the staged volume into that pod's directory, applying read-only and fsGroup settings. - `NodeUnpublishVolume`, `NodeUnstageVolume` - teardown in reverse. - `NodeGetInfo` - reports the node's ID in the storage system's terms and its topology labels, which the scheduler uses; also reports the maximum attachable volume count, surfaced in the `CSINode` object. - `NodeExpandVolume` - grows the filesystem after the backing device was expanded. The `node-driver-registrar` sidecar writes the driver's socket path into kubelet's plugin registration directory so kubelet discovers it. ## Drivers that skip a half The `CSIDriver` object advertises what is needed. A network filesystem driver (NFS, CephFS, EFS) usually sets `attachRequired: false`: there is nothing to attach, the node just mounts an export, so no VolumeAttachment objects are created and `external-attacher` is not needed. Conversely, a driver that only provides ephemeral inline volumes may have no controller service at all. ## Why the split matters operationally When storage misbehaves, the split tells you where to look. Provisioning stuck in `Pending`? Look at the controller pod and `external-provisioner` logs. Attach stuck (`Multi-Attach error`, `Volume not attached`)? Look at VolumeAttachment objects and `external-attacher`. Attached but the pod is stuck in `ContainerCreating` with a mount error? That is the node plugin on that specific node - check its DaemonSet pod, whether it is running at all, and the kubelet logs. Also note the blast radius asymmetry: a broken controller pauses new provisioning cluster-wide, while a broken node plugin only affects pods scheduled to that node.
- Why is NodeStageVolume separate from NodePublishVolume instead of one call?Staging is per node, publishing is per pod. Formatting and mounting the device once at a global staging path lets multiple pods on the same node bind-mount the same volume cheaply and consistently, and it makes the expensive, risky step (mkfs, device mount) happen exactly once. Teardown mirrors this: the last pod's unpublish is followed by a single unstage.
- Why does the node plugin run privileged, and what does that imply for security?It must mount filesystems and often access raw block devices on the host, which requires privileged mode plus host path mounts with bidirectional mount propagation. That makes every CSI node plugin a root-equivalent component on every node, so driver provenance, image signing and upgrade hygiene matter as much as for kubelet itself.
saying these in an interview costs you the question
- Saying the controller plugin mounts the volume - it never touches node filesystems
- Assuming every driver needs an attach step; NFS-style drivers set attachRequired: false
- Thinking the vendor driver watches Kubernetes objects - the sidecars do that
- Believing the node plugin runs only where the volume is used, rather than as a DaemonSet on all nodes
- Confusing NodeStageVolume (once per node) with NodePublishVolume (once per pod)