skip to content

Walk through everything that happens from the moment a user creates a PersistentVolumeClaim until the container has the volume mounted, when a Container Storage Interface (CSI) driver backs it. Where does that flow typically get stuck?

level: seniorimportance: should knowfreq 42%

answer

  1. Provision, bind, schedule, attach, stage, publish
  2. WaitForFirstConsumer fixes zone pinning
  3. VolumeAttachment = the attach receipt
  4. Multi-Attach = RWO stuck on a dead node
  5. One node failing = node plugin DaemonSet

basics

~20 s

PVC created; external-provisioner calls CreateVolume and creates the PV; the PVC binds; the scheduler places the pod; the attach controller creates a VolumeAttachment and external-attacher calls ControllerPublishVolume; kubelet calls NodeStageVolume then NodePublishVolume. It sticks most often at attach.

solid answer

~50 s

1. **Provision** - `external-provisioner` sees an unbound PVC whose StorageClass names the driver, calls `CreateVolume` on the controller plugin, and creates a matching PV. With `volumeBindingMode: WaitForFirstConsumer` this waits until the pod is scheduled, so the volume lands in the right zone. 2. **Bind** - the PV/PVC controller binds the pair; the PVC goes `Bound`. 3. **Schedule** - the scheduler filters nodes by the volume's topology and by each node's attach limit (from `CSINode`). 4. **Attach** - `kube-controller-manager`'s attach/detach controller creates a `VolumeAttachment`; `external-attacher` calls `ControllerPublishVolume`, and marks it attached. 5. **Mount** - kubelet calls `NodeStageVolume` (format, mount to a global staging path, once per node) then `NodePublishVolume` (bind-mount into the pod directory) on the node plugin. Common sticking points: `Pending` PVC (no matching StorageClass, quota, or `WaitForFirstConsumer`), zone mismatch, per-node attach limits, and `Multi-Attach error` when a ReadWriteOnce volume is still attached to an unreachable node.

code

bash · 5 lines
bash
kubectl describe pvc data-0
kubectl get pv -o custom-columns=NAME:.metadata.name,HANDLE:.spec.csi.volumeHandle,ZONE:.spec.nodeAffinity
kubectl get volumeattachments -o wide | grep <pv-name>
kubectl describe pod app-0 | tail -30
kubectl -n kube-system logs ds/ebs-csi-node -c ebs-plugin

go deeper

for a junior

Be able to say the flow is: claim, provision, bind, schedule, attach, mount, and that describe on the PVC and pod shows where it stopped.

for a middle

Name the component and object at each hop, and explain WaitForFirstConsumer.

for a senior

Diagnose from symptoms: node affinity conflict, attach limits, Multi-Attach on an unreachable node, single-node mount failures - and know why forcing detach unsafely risks data corruption.

for a principal

Talk about designing around the flow: topology-aware provisioning as a default, attach-limit-aware capacity planning, node-failure automation, and whether stateful workloads should rely on reattach at all versus application-level replication.

## The full path ### 1. Provisioning A PVC names a StorageClass; the StorageClass names a CSI driver as its `provisioner`. The `external-provisioner` sidecar in the controller pod watches PVCs. When it sees an unbound one for its driver it calls `CreateVolume` with the requested size, access mode, parameters and (if topology-aware) the allowed topology. On success it creates a PersistentVolume object whose `spec.csi.volumeHandle` is the real volume ID. If the StorageClass uses `volumeBindingMode: WaitForFirstConsumer`, nothing is provisioned until a pod referencing the PVC is scheduled. This is essential for zonal block storage: provisioning first would pin the disk to a zone before knowing where the pod can run. ### 2. Binding The control plane's PV/PVC controller binds claim to volume; the PVC status becomes `Bound`. Static provisioning skips step 1 and binds to a pre-created PV. ### 3. Scheduling The scheduler's volume-aware predicates check: does the node satisfy the PV's `nodeAffinity` (topology), and is the node already at its attachable-volume limit? The limit is reported by the node plugin via `NodeGetInfo` and stored in the `CSINode` object. ### 4. Attach For drivers with `attachRequired: true`, kube-controller-manager's attach/detach controller creates a `VolumeAttachment` object naming the volume and the node. `external-attacher` sees it, calls `ControllerPublishVolume`, and sets `status.attached: true` when the storage system reports the device attached to that machine. ### 5. Stage and publish Kubelet's volume manager waits for the attachment, then calls the node plugin over its UNIX socket: - `NodeStageVolume`: discover the device, `mkfs` if it is unformatted, mount at a global staging directory under `/var/lib/kubelet/plugins/kubernetes.io/csi/...`. Once per node per volume. - `NodePublishVolume`: bind-mount from staging into the pod's volume directory, applying read-only, `fsGroup` ownership and SELinux context. Only then does the container start. Teardown runs in reverse: `NodeUnpublishVolume`, `NodeUnstageVolume`, `ControllerUnpublishVolume`, and `DeleteVolume` if the reclaim policy is Delete and the PVC is gone. ## Where it gets stuck, and how to tell **PVC stays `Pending`.** `kubectl describe pvc` is the first stop. Either no StorageClass matched (typo, no default class), the provisioner is not running, the backend rejected the request (quota, unsupported size or access mode), or the class is `WaitForFirstConsumer` and no pod has been scheduled yet - which is normal, not a fault. **Pod `Pending` with a volume node-affinity conflict.** The PV was provisioned in zone A, the pod can only run in zone B (taints, capacity, node selectors). Symptom: `node(s) had volume node affinity conflict`. Zonal disks cannot move; fix is `WaitForFirstConsumer` for new volumes, or scheduling the pod back into the volume's zone. **Pod `ContainerCreating` with attach errors.** Check `kubectl get volumeattachments` and the `external-attacher` logs. Two classic cases: the node hit its attach limit (`CSINode` shows the maximum), or the storage API is throttling. **`Multi-Attach error for volume ...`.** A ReadWriteOnce volume is still attached to another node. This normally means a node went unreachable: its pods are stuck `Terminating`, so Kubernetes will not detach, because it cannot prove the old container stopped writing. Resolution is to make the old node's state authoritative - remove the failed Node object or use the non-graceful node shutdown taint - never to force-delete the pod and hope, since two writers on one filesystem corrupts it. **Mount errors on one node only.** The node plugin DaemonSet pod is missing, crash-looping, or the driver's socket is not registered. Check the DaemonSet on that node, `node-driver-registrar` logs, and kubelet logs for the driver name. ## The mental model Each step has a distinct owner, a distinct Kubernetes object you can inspect, and a distinct log. Provision -> PVC + provisioner. Bind -> PVC status. Schedule -> pod events. Attach -> VolumeAttachment + attacher. Mount -> node plugin + kubelet. Naming the object at each hop is what turns a vague "storage is broken" into a five-minute diagnosis.

  • A pod is stuck with 'Multi-Attach error for volume pvc-abc'. What is happening and what do you do?
    The volume is ReadWriteOnce and is still attached to a node that has not confirmed release - typically a node that went NotReady, leaving pods in Terminating so the detach never completes. The safe fix is to make the old node's state definitive: confirm it is really gone and delete the Node object, or rely on the non-graceful node shutdown taint so Kubernetes can force-detach. Force-deleting the pod alone is dangerous, because if the old kubelet is alive you get two writers on one filesystem.
  • Why does WaitForFirstConsumer exist, and when would you not use it?
    With immediate binding, the volume is created before the scheduler runs, so for zonal storage it may be provisioned where the pod cannot be placed, producing a node-affinity conflict. WaitForFirstConsumer defers CreateVolume until the pod is scheduled so topology is known. You would skip it for non-topological backends such as a network filesystem reachable from every node, or when you deliberately want the volume to exist before any workload.

saying these in an interview costs you the question

  • Skipping the attach stage entirely and jumping from PVC to mount
  • Claiming force-deleting the stuck pod is the correct fix for Multi-Attach
  • Believing the scheduler ignores volume topology and per-node attach limits
  • Saying kubelet calls the controller plugin - it only talks to the node plugin's socket
  • Thinking a Pending PVC under WaitForFirstConsumer is always an error

context