skip to content

Storage and Volumes

Giving a pod storage that outlives it: PersistentVolumes and Claims, StorageClasses and dynamic provisioning, access modes, CSI drivers, and day-2 expansion and snapshots. Stateful workloads are where Kubernetes gets hard, so interviewers dig in.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

29

A PersistentVolumeClaim in Kubernetes declares accessModes such as ReadWriteOnce, ReadOnlyMany, ReadWriteMany or ReadWriteOncePod. What does each of those mean, and what is the unit they are scoped to?

level: juniorimportance: must knowfreq 62%

answer

  1. RWO/ROX/RWX are node-scoped; RWOP is pod-scoped
  2. RWO = many pods OK if same node
  3. RWX needs a shared filesystem (NFS/CephFS/EFS/Azure Files)
  4. RWOP stable in 1.29, needs CSI SINGLE_NODE_SINGLE_WRITER
  5. mode is a match request + attach constraint, not file permissions

basics

~10 s

ReadWriteOnce: mounted read-write by one node. ReadOnlyMany: mounted read-only by many nodes. ReadWriteMany: mounted read-write by many nodes. ReadWriteOncePod: exactly one pod cluster-wide. Except for ReadWriteOncePod, the scope is the node, not the pod.

solid answer

~60 s

Access modes describe how many **nodes** may mount a volume and in which direction. - **ReadWriteOnce (RWO)** - mountable read-write by a single **node**. Multiple pods on that same node can all use it; a pod on a second node cannot. - **ReadOnlyMany (ROX)** - mountable read-only by many nodes at once. - **ReadWriteMany (RWX)** - mountable read-write by many nodes at once. Requires a shared-filesystem backend such as NFS, CephFS, EFS or Azure Files; block devices cannot offer it safely. - **ReadWriteOncePod (RWOP)** - mountable by exactly **one pod** in the whole cluster; the only pod-scoped mode. Stable since Kubernetes 1.29 and requires a CSI driver that supports it. Two things to stress. First, a mode is a **request matched against what the backend supports** - listing RWX does not make a block volume shareable; if no PV or StorageClass can satisfy it, the claim stays Pending. Second, the mode is not a per-mount permission: whether a container writes is governed by `readOnly` on the volume mount, not by the access mode.

code

yaml · 23 lines
yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: db-data
spec:
  accessModes:
    - ReadWriteOncePod
  resources:
    requests:
      storage: 50Gi
  storageClassName: fast-block
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: shared-uploads
spec:
  accessModes:
    - ReadWriteMany
  resources:
    requests:
      storage: 200Gi
  storageClassName: nfs-shared

go deeper

for a junior

List the four modes and state clearly that RWO/ROX/RWX count nodes while RWOP counts pods.

for a middle

Add that the mode is matched against backend capability, that RWX needs a shared filesystem, and that readOnly on the mount is what governs writing.

for a senior

Connect modes to workload shape: RWO plus RollingUpdate causes multi-attach stalls, RWOP guarantees a single writer, and RWX costs POSIX semantics and throughput.

for a principal

Set the platform default (block RWO per replica), decide when shared filesystems are permitted at all, and weigh the operational cost of running an RWX backend versus pushing teams to object storage or a database.

## What access modes describe A `PersistentVolume` advertises the access modes its backend can support; a `PersistentVolumeClaim` requests one (a list is allowed, but binding uses one mode). The binding logic pairs a claim with a PV - or a StorageClass provisioner - that can satisfy the requested mode and size. If nothing can, the PVC sits in `Pending` and the pod that references it never starts. The key thing to internalise: **the unit is the node**, not the pod, for three of the four modes. That is a consequence of how storage is attached - a cloud disk or SAN LUN is attached to a *machine*, and once it is attached, everything on that machine can use it. ## The four modes **ReadWriteOnce (RWO)** - the volume may be attached read-write to a single node. Any number of pods scheduled on that node can mount it simultaneously and all can write. This surprises people who read "Once" as "one pod". It is the mode almost all block storage supports: AWS EBS, GCE Persistent Disk, Azure Disk, Ceph RBD, iSCSI LUNs, most local volumes. **ReadOnlyMany (ROX)** - the volume may be attached to many nodes, but only read-only. Typical for a pre-populated dataset, a model file or static content shared by many replicas. Support is backend-specific: some block backends allow multi-attach in read-only mode, and file backends generally do. **ReadWriteMany (RWX)** - many nodes, all read-write, at the same time. This requires a backend that implements a distributed or network filesystem with its own locking: NFS, CephFS, GlusterFS, AWS EFS, Azure Files, Google Filestore, Portworx, Longhorn (which serves RWX through an NFS layer). A plain block device cannot do it: two nodes writing to one block device with independent, non-cluster-aware filesystems corrupt it, because each node caches metadata it believes it owns exclusively. **ReadWriteOncePod (RWOP)** - read-write by exactly one **pod** cluster-wide. Alpha in 1.22, beta in 1.27, stable in 1.29. It exists precisely because RWO's node scope surprised people and offered no way to guarantee a single writer. Kubernetes enforces it at admission and mount time, and it requires a CSI driver reporting the `SINGLE_NODE_SINGLE_WRITER` capability. Use it for single-writer databases where two concurrent writers would corrupt data. ## Requested, matched - and only partly enforced An access mode is a **matchmaking attribute plus an attach-layer constraint**. Two frequent misunderstandings: - *Requesting a mode does not create the capability.* Ask for RWX from a StorageClass backed by cloud block storage and the claim never binds; you see a Pending PVC and a pod stuck in `ContainerCreating`. The volume type has to support it. - *It is not a file-permission setting.* Binding with `ReadOnlyMany` does not, by itself, make writes fail inside a container in every implementation; what reliably makes a mount read-only is `readOnly: true` on the pod's `volumeMounts` entry (or `persistentVolumeClaim.readOnly`). Treat the access mode as "how many machines may attach, and how", and the pod spec as "what this container may do". The attach layer *does* enforce the important part: with RWO, the external attach machinery will refuse to attach the same volume to a second node while the first still holds it - the origin of the `Multi-Attach error for volume ...` event. ## Choosing Start with RWO: it is universally supported, fastest, and correct for anything with a single writer per instance - which includes every StatefulSet replica that has its own volume. Reach for RWX only when several pods genuinely must write to the *same* filesystem at once, and know the price: a network filesystem in the data path, weaker POSIX semantics (especially locking), and lower throughput. Many apparent RWX needs are better served by object storage, a database, or giving each replica its own RWO volume. Use RWOP when a single writer is a correctness requirement rather than an expectation, and your CSI driver supports it. Use ROX for read-only shared datasets, remembering that some backends implement it as a genuine multi-attach and others by simply mounting read-only. ## Where you see it go wrong The standard incident: a Deployment with `strategy: RollingUpdate` and an RWO PVC. The new pod is scheduled to a different node, the volume cannot detach from the old node until the old pod terminates, and the rollout wedges with a `Multi-Attach` event. The mode was doing exactly what it promised; the workload shape was wrong for it.

  • Two pods on the same node both mount a PVC bound with ReadWriteOnce, and both write. Is that allowed?
    Yes. ReadWriteOnce is scoped to the node, so once the volume is attached there, any number of pods on that node may mount it read-write. Kubernetes will not stop them, and it will not arbitrate their writes - if the application cannot tolerate concurrent writers you must enforce that yourself, or use ReadWriteOncePod, which restricts the volume to a single pod cluster-wide.
  • A PVC requesting ReadWriteMany stays Pending forever against a cloud block StorageClass. What is happening?
    The provisioner cannot satisfy the requested mode, because block devices such as EBS or Azure Disk cannot be safely mounted read-write from several nodes at once. No PV is created, the claim never binds, and any pod referencing it stays in ContainerCreating. Either switch to a file-based StorageClass that supports RWX (NFS, EFS, Azure Files, CephFS), or redesign so each replica gets its own ReadWriteOnce volume.

Think of a volume as a workshop key. RWO hands one key to a building (a node) - everyone inside can use the workshop. ROX makes many read-only copies for many buildings. RWX means every building can work in it at once, which only works if the workshop was designed for crowds. RWOP hands the key to exactly one person.

saying these in an interview costs you the question

  • Reading ReadWriteOnce as 'one pod' rather than 'one node'.
  • Assuming requesting ReadWriteMany makes any backend support it.
  • Treating ReadOnlyMany as the way to make a mount read-only for a container (that is readOnly on the volumeMount).
  • Thinking ReadWriteOncePod works on any driver, without the CSI capability.
  • Believing Kubernetes coordinates concurrent writers for you once the mode allows multiple mounts.

context

open as a page

What is the Container Storage Interface (CSI) in Kubernetes, and why were storage drivers moved out of the Kubernetes core codebase onto it?

level: juniorimportance: must knowfreq 50%

basics

~20 s

CSI is a standard gRPC contract for storage drivers. A vendor ships the driver as ordinary pods in the cluster instead of code compiled into Kubernetes, so drivers install, upgrade and get fixed independently of the Kubernetes release.

open as a page

What does the `volumeClaimTemplates` field on a Kubernetes StatefulSet do, and how does it differ from listing a PersistentVolumeClaim under a Deployment's pod volumes?

level: juniorimportance: must knowfreq 58%

basics

~10 s

volumeClaimTemplates makes the StatefulSet controller create one PersistentVolumeClaim per replica, named <template>-<statefulset>-<ordinal>. A Deployment's pod spec references one existing PVC, so every replica shares the same volume.

open as a page

What is a Kubernetes StorageClass, and how does a PersistentVolumeClaim that references one end up with real storage mounted into a Pod?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A StorageClass is a named storage tier: a provisioner plus its parameters. When a PersistentVolumeClaim names that class, the provisioner creates real storage and a matching PersistentVolume, which is bound to the claim. No admin pre-creates the volume.

open as a page

In a Pod spec you can mount an emptyDir volume or mount a PersistentVolumeClaim. Explain the lifecycle difference between the two and when each is the right choice.

level: juniorimportance: must knowfreq 72%

basics

~20 s

An emptyDir is created when the Pod is placed on a node and deleted with the Pod, so it survives container restarts but not rescheduling. A PersistentVolumeClaim points at storage whose lifetime is independent of the Pod, so data survives deletion and rescheduling.

open as a page

How do two containers in the same Kubernetes pod share files through an emptyDir volume, and does that data survive a container crash?

level: juniorimportance: must knowfreq 72%

basics

~20 s

An emptyDir is an empty directory the kubelet creates when the pod starts on a node. Every container that mounts it sees the same files. It lasts as long as the pod, so a container crash and restart keeps the data.

open as a page

Why is a Kubernetes volume bound with ReadWriteOnce described as node-scoped rather than pod-scoped, what problems does that cause, and which access mode fixes it?

level: middleimportance: must knowfreq 54%

basics

~20 s

Storage attaches to a machine, so ReadWriteOnce limits the volume to one node - any number of pods on that node can mount and write to it. That silently allows concurrent writers. ReadWriteOncePod restricts it to exactly one pod cluster-wide and is enforced by Kubernetes.

open as a page

A Container Storage Interface (CSI) driver is usually deployed as both a Deployment and a DaemonSet. What does each part do, and which storage operations belong to which?

level: middleimportance: must knowfreq 45%

basics

~20 s

The controller plugin (a Deployment) does cluster-wide work against the storage API: create, delete, attach, detach, snapshot, expand. The node plugin (a DaemonSet on every node) does machine-local work: format, stage-mount and bind-mount the volume into the pod.

open as a page

After raising a Kubernetes PersistentVolumeClaim's spec.resources.requests.storage, why can kubectl get pvc still show the old capacity, and what completes the resize?

level: middleimportance: must knowfreq 62%

basics

~20 s

Expansion has two phases. The storage backend grows the volume first, then the kubelet grows the filesystem on the node. status.capacity, which kubectl shows, changes only after both finish. FileSystemResizePending means the node step is waiting for the volume to be mounted.

open as a page

When a pod in a StatefulSet is rescheduled to a different node, how does it end up with the same data it had before, and what can prevent that from happening?

level: middleimportance: must knowfreq 48%

basics

~20 s

The replacement pod keeps the same ordinal name, so it binds the same ordinal-suffixed PersistentVolumeClaim; the volume is detached from the old node and attached to the new one. Zone pinning, per-node attach limits, and a stuck detach from an unreachable node can block it.

open as a page

How does a Kubernetes PersistentVolumeClaim end up bound to a PersistentVolume, and what do the PersistentVolume phases Available, Bound, Released and Failed mean?

level: middleimportance: must knowfreq 58%

basics

~20 s

A controller matches an unbound claim to a volume with a compatible class, access mode, volume mode and at least the requested capacity, then binds them one-to-one. PV phases: Available (free), Bound (claimed), Released (its claim was deleted but the volume is not reclaimed yet), Failed (automatic reclamation errored).

open as a page

In a Kubernetes StorageClass, what is the difference between volumeBindingMode: Immediate and volumeBindingMode: WaitForFirstConsumer, and what production failure does the second one prevent?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Immediate provisions and binds a volume as soon as the claim exists, before any Pod is scheduled, so the disk can land in a zone or node the Pod cannot reach. WaitForFirstConsumer delays provisioning until a Pod using the claim is scheduled, letting the scheduler's decision drive volume placement.

open as a page

What does the persistentVolumeReclaimPolicy field on a Kubernetes PersistentVolume control, and what actually happens to the data when a PersistentVolumeClaim is deleted under Retain versus Delete?

level: seniorimportance: must knowfreq 48%

basics

~20 s

It decides the volume's fate once its claim is deleted. Delete removes the PersistentVolume object and destroys the backing storage and its data. Retain keeps both: the volume goes to the Released phase with data intact and needs manual cleanup or reuse by an administrator.

open as a page

A team takes a nightly Kubernetes VolumeSnapshot of a checkout database's PersistentVolumeClaim. Why is that, on its own, not a real database backup?

level: juniorimportance: should knowfreq 46%

basics

~20 s

A VolumeSnapshot normally lives on the same storage system as the volume, is only crash-consistent, and is deleted along with its Kubernetes object under the Delete policy. A backup needs application consistency, a separate failure domain, independent retention and a tested restore.

open as a page

A PersistentVolumeClaim in Kubernetes has a volumeMode field that accepts Filesystem or Block. What is the difference, how does a pod consume each, and when is Block the right choice?

level: middleimportance: should knowfreq 34%

basics

~20 s

Filesystem (the default) means Kubernetes formats and mounts the volume at a directory given by volumeMounts. Block exposes the raw device with no filesystem, surfaced at a path via volumeDevices; the application must handle the device itself. Block suits databases and storage systems that manage their own layout.

open as a page

How do VolumeSnapshot, VolumeSnapshotContent and VolumeSnapshotClass work together in Kubernetes, and how do you create a new PersistentVolumeClaim from a snapshot?

level: middleimportance: should knowfreq 33%

basics

~10 s

VolumeSnapshot is the namespaced request, VolumeSnapshotContent is the cluster-scoped object representing the real snapshot, VolumeSnapshotClass picks the driver and parameters - mirroring PVC/PV/StorageClass. Restore by creating a PVC with dataSource pointing at the VolumeSnapshot.

open as a page

A PersistentVolumeClaim is running out of space. What does the allowVolumeExpansion field on a Kubernetes StorageClass permit, and what are the exact steps and limits of growing an existing claim?

level: middleimportance: should knowfreq 42%

basics

~20 s

With allowVolumeExpansion: true on the claim's StorageClass, you edit the PVC's spec.resources.requests.storage upward and the CSI driver grows the disk, then the filesystem. Shrinking is never allowed, and the flag only affects claims whose class had it set.

open as a page

One PersistentVolumeClaim omits the storageClassName field entirely; another sets storageClassName to the empty string "". Explain what Kubernetes does in each case, and how a cluster declares which storage tier is the default.

level: middleimportance: should knowfreq 45%

basics

~20 s

Omitting storageClassName lets the DefaultStorageClass admission controller fill in the cluster's default class, marked by the annotation storageclass.kubernetes.io/is-default-class: "true". Setting it to "" opts out of dynamic provisioning entirely, so the claim can only bind to a pre-created PV that also has no class.

open as a page

Compare the Kubernetes volume types emptyDir, hostPath, and generic ephemeral volumes (the ephemeral.volumeClaimTemplate field in a Pod spec). What is each for, and why is hostPath discouraged in multi-tenant clusters?

level: middleimportance: should knowfreq 46%

basics

~20 s

emptyDir is per-Pod scratch space managed by kubelet. hostPath mounts an arbitrary node directory into the container, which pins the Pod to that node and can expose the host, so it is restricted. Generic ephemeral volumes create a real provisioned claim that is deleted with the Pod.

open as a page

In Kubernetes, what does an emptyDir volume with medium: Memory give you, and how do sizeLimit and memory limits bound it?

level: middleimportance: should knowfreq 46%

basics

~20 s

medium: Memory mounts the emptyDir as RAM-backed tmpfs, which is fast and never touches the node disk. The kubelet caps it at the smallest of sizeLimit, the pod's memory limit and node allocatable, and every byte written counts as memory use.

open as a page

What does a Kubernetes projected volume do, and why would a pod project a serviceAccountToken with its own audience and expirationSeconds?

level: middleimportance: should knowfreq 41%

basics

~20 s

A projected volume merges Secrets, ConfigMaps, downward API fields and service account tokens into one read-only directory. A projected serviceAccountToken is short-lived, scoped to one audience and rotated by the kubelet, so it is safer to hand to an external service.

open as a page

Several pods on different nodes need to write to the same persistent volume in Kubernetes at the same time. What kind of storage backend can actually support that, and why can a cloud block device such as AWS EBS or Azure Disk not do it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

You need a shared network filesystem with its own coordination: NFS, CephFS, AWS EFS, Azure Files, Google Filestore, Portworx or Longhorn. Block devices cannot do it because each node's filesystem caches metadata assuming exclusive ownership, so two writers corrupt it.

open as a page

Walk through everything that happens from the moment a user creates a PersistentVolumeClaim until the container has the volume mounted, when a Container Storage Interface (CSI) driver backs it. Where does that flow typically get stuck?

level: seniorimportance: should knowfreq 42%

basics

~20 s

PVC created; external-provisioner calls CreateVolume and creates the PV; the PVC binds; the scheduler places the pod; the attach controller creates a VolumeAttachment and external-attacher calls ControllerPublishVolume; kubelet calls NodeStageVolume then NodePublishVolume. It sticks most often at attach.

open as a page

A Kubernetes StatefulSet's replica checkout-db-2 has corrupted data. How do you replace its PersistentVolumeClaim with one restored from a VolumeSnapshot or cloned from a healthy replica?

level: seniorimportance: should knowfreq 40%

basics

~20 s

An existing claim cannot be repointed, because dataSource is used only at creation. Scale the StatefulSet so the pod goes away, delete the claim, recreate one with the same name whose dataSource names the snapshot or source PVC, then scale back up.

open as a page

In Kubernetes, what does a StatefulSet's `persistentVolumeClaimRetentionPolicy` control, what are its defaults, and what happens to the per-replica claims when you scale a StatefulSet from 5 replicas down to 3?

level: seniorimportance: should knowfreq 34%

basics

~20 s

It sets whether the generated PVCs are deleted when the StatefulSet is deleted (whenDeleted) or when it scales down (whenScaled). Both default to Retain, so scaling 5 to 3 leaves the claims for ordinals 3 and 4 - and their data - in place.

open as a page

A Kubernetes PVC is mistakenly resized from 80Gi to 8000Gi and the expansion fails. How do you recover, given that claims cannot shrink?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Lower spec.resources.requests.storage to a realistic size that is still above status.capacity, which RecoverVolumeExpansionFailure allows (GA since 1.34). The resizer then retries the smaller target. A true shrink still means a new, smaller claim and a data copy.

open as a page

A document-OCR pipeline on a Kubernetes cluster that autoscales between 20 and 80 nodes needs 140Gi of per-pod scratch. When would you choose a generic ephemeral volume over emptyDir, and what can go wrong?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

Choose a generic ephemeral volume when per-pod scratch is too large for the node disk. A storage driver provisions a PVC for each pod, and it is deleted with the pod. The costs are provisioning latency, name collisions, quota and attach limits.

open as a page

You are setting the storage defaults for a Kubernetes platform many teams deploy onto. How would you decide which persistent-volume access modes to offer, which to make the default, and how would you handle teams that ask for shared read-write volumes?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Default to per-replica ReadWriteOnce block volumes: universally supported, fast, small blast radius. Offer ReadWriteOncePod for single-writer correctness. Treat ReadWriteMany as an exception requiring a justification, since it means operating a shared filesystem that becomes a dependency for every consumer.

open as a page

A running Kubernetes StatefulSet needs each replica's persistent volume grown from 100Gi to 500Gi. Why can't you simply edit the size in `volumeClaimTemplates`, and how would you carry out the change?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

volumeClaimTemplates is immutable on an existing StatefulSet, so the edit is rejected. Expand each PVC directly (the StorageClass must allow expansion), then delete the StatefulSet with --cascade=orphan and recreate it with the new template so future replicas match.

open as a page