skip to content

Several pods on different nodes need to write to the same persistent volume in Kubernetes at the same time. What kind of storage backend can actually support that, and why can a cloud block device such as AWS EBS or Azure Disk not do it?

level: seniorimportance: should knowfreq 40%

answer

  1. RWX = shared filesystem: NFS/EFS/Azure Files/Filestore/CephFS
  2. block device = sectors; ext4 caches metadata assuming sole owner
  3. two ext4 writers = double-allocated blocks, silent corruption
  4. EBS Multi-Attach = cluster-aware FS only, not RWX for ext4
  5. costs: latency, weak locking, price, shared blast radius

basics

~20 s

You need a shared network filesystem with its own coordination: NFS, CephFS, AWS EFS, Azure Files, Google Filestore, Portworx or Longhorn. Block devices cannot do it because each node's filesystem caches metadata assuming exclusive ownership, so two writers corrupt it.

solid answer

~60 s

Concurrent multi-node writes require `ReadWriteMany`, and only **shared-filesystem** backends implement it: NFS (self-hosted or managed - AWS EFS, Azure Files, Google Filestore), CephFS, Portworx shared volumes, Longhorn's RWX (an NFS layer over its block volumes), and similar. Block devices - EBS, GCE PD, Azure Disk, iSCSI LUNs, Ceph RBD - cannot. The device itself is just addressable sectors; the *filesystem* on top (ext4, xfs) is what gives it structure, and a single-node filesystem caches inodes, allocation bitmaps and journal state in the node's page cache on the assumption that nobody else is writing. Attach the same device to two nodes and each will happily allocate the same block twice. Result: corruption, usually silent for a while. AWS EBS Multi-Attach exists but is explicitly only for cluster-aware filesystems or applications that coordinate themselves - it is not a way to get RWX for ext4. The cost of RWX is real: a network filesystem in the write path, weaker POSIX locking semantics, higher latency, and often per-GB or per-throughput pricing. Before adopting it, ask whether object storage, a database, or one RWO volume per replica solves the problem instead.

code

yaml · 11 lines
yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: media-scratch
spec:
  accessModes:
    - ReadWriteMany
  resources:
    requests:
      storage: 500Gi
  storageClassName: efs-sc

go deeper

for a junior

Know that ReadWriteMany needs a shared filesystem such as NFS or EFS, and that ordinary cloud disks only do ReadWriteOnce.

for a middle

Explain why: a block device carries no coordination, and a node-local filesystem caches metadata assuming it is the only writer.

for a senior

Cover the operational reality - locking weakness, small-file latency, cost, the shared blast radius - and reroute common requests to object storage or per-replica volumes.

for a principal

Decide whether the platform offers RWX at all, who operates the file service, what SLOs and cost model apply, and which application patterns are allowed to depend on a shared POSIX namespace.

## What ReadWriteMany really demands `ReadWriteMany` means many nodes hold the volume read-write simultaneously. The hard part is not attachment; it is **coordination**. Something must arbitrate concurrent metadata updates - who owns which blocks, which directory entries exist, whose write lands last - and make each node's view consistent. That coordination has to live somewhere. In shared-filesystem backends it lives in a server or a distributed protocol. In block storage, it does not exist at all. ## Why block devices cannot be shared A block device is an array of sectors with no notion of files. Structure comes from a filesystem - typically ext4 or xfs - running **on the node**. That filesystem keeps in the node's page cache: the inode table, the free-space bitmaps, the journal state, directory indexes. It updates them lazily and writes back on its own schedule, because it assumes it is the only writer. Mount the same device on a second node with its own ext4 instance and both assumptions break simultaneously. Node A allocates block 1000 for file X; node B, reading a stale bitmap, allocates the same block for file Y. Both journals are authoritative in their own view. The result is corruption that a filesystem check cannot untangle, and it is often silent for hours before surfacing. This is why cloud block services either forbid multi-attach or gate it loudly. **AWS EBS Multi-Attach** (io1/io2 in the same availability zone, limited node count) attaches one volume to several instances, but the documentation is explicit: it is for **cluster-aware filesystems** (GFS2, OCFS2) or applications that implement their own coordination and use their own fencing. Handing it ext4 and calling it RWX produces exactly the corruption described above. **Ceph RBD** is the same story - a block image, RWO in practice - which is precisely why Ceph also ships **CephFS** for the shared-filesystem case. ## Backends that do provide RWX - **NFS** - the classic answer. A server (self-managed, an appliance, or managed as **AWS EFS**, **Azure Files** with SMB/NFS, **Google Filestore**) owns the data and serialises operations. Every client speaks the protocol rather than touching sectors. - **CephFS** - a POSIX filesystem over RADOS with metadata servers coordinating access; scales further than a single NFS server. - **Longhorn RWX** - implemented by exporting its own block volume through an NFS share pod. Convenient, but it means a share pod is in the data path and is a failure domain. - **Portworx, GlusterFS, and vendor SDS** - each with their own replication and coordination models. - **Managed multi-writer file services** generally: they are all a network filesystem with a server behind it. ## The costs you accept **Latency and throughput.** Every metadata operation is a network round trip. Workloads that stat thousands of files (PHP includes, node_modules, small-file builds) degrade badly. **Locking semantics.** NFS locking is notoriously partial; `flock`, `fcntl` and `O_EXCL` behave differently or unreliably across implementations and versions. Software that relies on exclusive-create or advisory locks for correctness - SQLite, some queue implementations, lock-file-based leader election - is unsafe here, no matter what the access mode says. **Consistency.** Client-side attribute caching means one pod may not see another's write immediately unless the mount options are tightened, which costs more performance. **Cost and operations.** Managed file services bill per GB and often per throughput unit, at a multiple of block storage. Self-hosting NFS means owning a stateful single point of failure; making it highly available is a project of its own. **Failure blast radius.** With RWO-per-replica, one bad volume affects one pod. With RWX, the shared filesystem is a dependency for every replica. ## Ask whether you need it at all Most RWX requests in practice fall into a few patterns with better answers: - *User uploads shared across replicas* -> object storage (S3/GCS/Azure Blob) with a presigned-URL flow. This is the single most common misuse of RWX. - *Shared configuration or static assets* -> bake into the image, or ConfigMaps, or an init-container sync; read-mostly data can also use `ReadOnlyMany`. - *A work queue over a shared directory* -> a real queue or a database table. - *Per-replica state that merely looks shared* -> a StatefulSet where each replica gets its own RWO volume. Genuine RWX cases do exist: legacy applications that cannot be changed and expect a shared POSIX path, media pipelines where several workers must read and write large files in place, and shared scratch space for tightly coupled batch jobs. Adopt it deliberately for those, and size the backend for the metadata rate, not just the capacity.

  • AWS EBS supports Multi-Attach. Does that give you ReadWriteMany for a normal application?
    No. Multi-Attach lets one io1/io2 volume attach to several instances in the same availability zone, but AWS restricts it to cluster-aware filesystems such as GFS2 or OCFS2, or applications that coordinate and fence writes themselves. Mounting ext4 on two nodes over it corrupts the filesystem, because each node's ext4 caches metadata assuming exclusive ownership. For real RWX on AWS you use EFS or another shared filesystem.
  • A team asks for an RWX volume so all replicas of a web app can serve user uploads. What would you propose instead, and why?
    Object storage with presigned upload and download URLs. It removes the shared-filesystem dependency entirely, scales horizontally without a metadata bottleneck, costs far less per GB, and gives durability and lifecycle policies for free. RWX would put a network filesystem in the request path, add a shared failure domain across all replicas, and perform poorly on many small files.

A block device is a blank notebook; ext4 is one person's private index of it. Give two people their own index of the same notebook and they will both claim page 40. A shared filesystem is a librarian who owns the notebook and writes down every change on your behalf.

saying these in an interview costs you the question

  • Believing you can just request ReadWriteMany and any StorageClass will provide it.
  • Claiming EBS Multi-Attach makes ext4 safe across nodes.
  • Assuming NFS file locking is fully POSIX-compliant and safe for lock-based coordination.
  • Reaching for RWX for shared uploads when object storage is the right answer.
  • Treating RWX as free - ignoring latency, small-file metadata cost, pricing and the shared failure domain.

context