skip to content

Why is a Kubernetes volume bound with ReadWriteOnce described as node-scoped rather than pod-scoped, what problems does that cause, and which access mode fixes it?

level: middleimportance: must knowfreq 54%

answer

  1. attach is machine-level -> node scope
  2. same node = two writers, no warning
  3. different node = Multi-Attach error, rollout stalls
  4. RWOP: one pod cluster-wide, stable 1.29, CSI SINGLE_NODE_SINGLE_WRITER
  5. fixes: Recreate strategy, StatefulSet per-replica volumes

basics

~20 s

Storage attaches to a machine, so ReadWriteOnce limits the volume to one node - any number of pods on that node can mount and write to it. That silently allows concurrent writers. ReadWriteOncePod restricts it to exactly one pod cluster-wide and is enforced by Kubernetes.

solid answer

~60 s

Attachment happens at the **machine** level: a cloud disk or LUN is attached to a node, and after that any pod on that node can mount it. So `ReadWriteOnce` means "one node", and two pods co-scheduled on that node both get read-write access with no warning. That causes two distinct problems. **Silent concurrent writers.** A Deployment with `RollingUpdate` may briefly run old and new pods together. If both land on the same node, both write to the same volume, and a database or any store assuming exclusive access can corrupt itself. Kubernetes arbitrates nothing. **Wedged rollouts.** If instead the new pod is scheduled to a *different* node, the attach layer refuses - it will not attach a ReadWriteOnce volume to a second node while the first holds it - and you get `Multi-Attach error for volume ...`, with the rollout stalled until the old pod terminates. The fix for the first is **`ReadWriteOncePod`** (stable in 1.29, needs a CSI driver with `SINGLE_NODE_SINGLE_WRITER`), which Kubernetes enforces at admission and mount so only one pod cluster-wide can use the claim. The fix for the second is workload shape: `strategy: Recreate`, or a StatefulSet with per-replica volumes.

code

yaml · 37 lines
yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: sqlite-data
spec:
  accessModes:
    - ReadWriteOncePod
  resources:
    requests:
      storage: 20Gi
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: ledger
spec:
  replicas: 1
  strategy:
    type: Recreate
  selector:
    matchLabels:
      app: ledger
  template:
    metadata:
      labels:
        app: ledger
    spec:
      containers:
        - name: app
          image: example/ledger:3.2
          volumeMounts:
            - name: data
              mountPath: /var/lib/ledger
      volumes:
        - name: data
          persistentVolumeClaim:
            claimName: sqlite-data

go deeper

for a junior

Say that the volume attaches to a node, so pods sharing that node share the volume, and that ReadWriteOncePod restricts it to one pod.

for a middle

Explain the attach-then-mount split, both failure modes, and the version and CSI requirement for ReadWriteOncePod.

for a senior

Diagnose a stuck rollout, know why a dead node pins a volume for minutes, and choose the workload shape (Recreate versus StatefulSet) rather than only the access mode.

for a principal

Make it policy: stateful single-writer workloads default to StatefulSets with per-replica volumes and RWOP where drivers allow, plus a documented runbook for force-detach on unreachable nodes.

## Where the node scope comes from Persistent storage in Kubernetes is delivered in two stages. First **attach**: the volume is made visible to a node - a cloud API call attaches the disk, or an iSCSI/NVMe session is established. Then **mount**: the kubelet mounts that device (or network share) into the container's filesystem namespace. The first stage is inherently machine-level. Nothing about attaching a disk to a virtual machine knows what a pod is. Once the device is present on the node, the kubelet can mount it into as many pods as are scheduled there. Hence: `ReadWriteOnce` means *one node may hold it read-write*, and says nothing about how many pods on that node use it. ## Problem one: silent concurrent writers The dangerous case is quiet. Two pods on the same node open the same data directory. A relational database with a single-writer assumption, an embedded store like SQLite or RocksDB, or any application keeping a lock file or in-memory index will misbehave: torn writes, corrupted indexes, split state. Nothing in Kubernetes prevents it, warns about it, or repairs it. How do two pods end up together? A `RollingUpdate` Deployment referencing one PVC creates the replacement pod before terminating the old one, and the scheduler may well pick the same node (it is often the best fit - the volume's topology constraints point there). A manually scaled `replicas: 2` on a Deployment with a shared PVC does the same. A Job or CronJob overlapping its previous run does too. ## Problem two: the multi-attach stall The mirror-image failure. If the replacement pod is scheduled to a **different** node, the attach layer will not attach the ReadWriteOnce volume there while the original node still holds it. The new pod stays `ContainerCreating` and events show: `Multi-Attach error for volume "pvc-..." Volume is already exclusively attached to one node and can't be attached to another` Usually it clears once the old pod terminates and the detach completes. It becomes a real outage when the old node is unreachable: Kubernetes cannot confirm the pod is gone and will not force-detach, so the volume stays pinned for several minutes (the pod-eviction and force-detach timers), or indefinitely until an operator intervenes with the out-of-service taint on the dead node. ## The fix for exclusivity: ReadWriteOncePod `ReadWriteOncePod` (alpha 1.22, beta 1.27, **stable 1.29**) is the only pod-scoped access mode. When a PVC is bound with it, Kubernetes enforces that exactly one pod cluster-wide may use that claim: admission rejects a second pod referencing it, and the mount path enforces single-writer semantics. It requires a CSI driver that reports the `SINGLE_NODE_SINGLE_WRITER` capability; in-tree legacy plugins do not support it. Use it wherever a second writer would be a correctness bug rather than merely a surprise - single-instance databases, embedded stores, anything holding a private lock file. It costs nothing in throughput; it just closes a hole. Note what it does *not* fix: it makes the second pod fail fast instead of corrupting data, but the rollout still cannot overlap. Correctness improves, availability does not. ## The fix for rollouts: change the workload shape If a workload owns a single volume, a rolling update is the wrong strategy. Options: - **`strategy: Recreate`** on the Deployment: old pod terminates fully before the new one starts. Downtime, but no overlap and no multi-attach. - **StatefulSet with `volumeClaimTemplates`**: each replica gets its *own* RWO volume, and the ordered update policy never runs two pods with the same identity at once. This is the standard shape for stateful workloads. - **Node affinity pinning** so the replacement lands on the same node: this avoids the multi-attach error, but re-opens the concurrent-writer window, so it is only safe combined with RWOP or an application that tolerates two readers. - **RWX backend**: only if the application is genuinely designed for a shared filesystem. ## How to recognise it in an interview Say the mechanism (attach is machine-level, hence node scope), name both failure modes (silent co-writers on the same node, multi-attach stall on different nodes), name RWOP with its version and CSI requirement, and finish with the workload-shape fix. Candidates who only recite "RWO means one node" without the two consequences sound like they have read the docs but never broken a cluster.

  • A rollout is stuck with 'Multi-Attach error for volume ... already exclusively attached to one node'. What is happening and how do you resolve it?
    The replacement pod was scheduled to a different node while the old pod still holds the ReadWriteOnce volume, and the attach layer refuses a second node. Normally it clears as soon as the old pod terminates and the detach completes. To stop it recurring, switch the Deployment to strategy: Recreate, or move the workload to a StatefulSet with volumeClaimTemplates so each replica owns its own volume.
  • Does ReadWriteOncePod make rolling updates safe for a single-volume workload?
    No. It makes them safe from data corruption, not from downtime: the incoming pod is rejected rather than allowed to write concurrently, so the rollout simply cannot overlap. You still need Recreate or a StatefulSet update strategy, and you should expect a gap between the old pod terminating and the new one being ready.

Attaching a volume is like docking a filing cabinet in a room. ReadWriteOnce says only one room gets the cabinet - it says nothing about how many people are in that room. ReadWriteOncePod says only one person may be in there.

saying these in an interview costs you the question

  • Saying ReadWriteOnce prevents two pods from writing at once.
  • Assuming Kubernetes arbitrates or locks concurrent writes when two pods share a volume.
  • Treating the Multi-Attach error as a storage-driver bug rather than the mode working as designed.
  • Thinking ReadWriteOncePod eliminates the need to change the update strategy.
  • Pinning both pods to one node to dodge multi-attach without realising that re-enables concurrent writers.

context