skip to content

You are setting the storage defaults for a Kubernetes platform many teams deploy onto. How would you decide which persistent-volume access modes to offer, which to make the default, and how would you handle teams that ask for shared read-write volumes?

level: principalimportance: nice to knowfreq 30%

answer

  1. default = RWO per replica via volumeClaimTemplates
  2. RWOP for single-writer correctness (fail loud, not corrupt quiet)
  3. RWX = a service to operate + shared blast radius
  4. triage: uploads -> object store, locks -> DB/queue, config -> image
  5. encode in StorageClasses: expansion, reclaim, WaitForFirstConsumer

basics

~20 s

Default to per-replica ReadWriteOnce block volumes: universally supported, fast, small blast radius. Offer ReadWriteOncePod for single-writer correctness. Treat ReadWriteMany as an exception requiring a justification, since it means operating a shared filesystem that becomes a dependency for every consumer.

solid answer

~1 min

I would make **ReadWriteOnce block storage the default**, one volume per replica via StatefulSet `volumeClaimTemplates`. It is supported by every driver, it is the fastest option, and a failure is contained to one pod. I would offer **ReadWriteOncePod** for workloads where a second writer is a correctness bug - single-instance databases, embedded stores - assuming the CSI driver supports it, and use it as the default for those templates rather than relying on convention. **ReadWriteMany** I would treat as an exception with a review, because it is not just an access mode - it is a service someone has to operate, with its own SLO, its own capacity and metadata-rate limits, and a shared failure domain across every consumer. The review asks one question: is a shared POSIX namespace a genuine requirement, or an accident of how the app was written? Uploads go to object storage; shared config goes into the image or a ConfigMap; work queues go to a real queue. If it is genuine - legacy software that cannot change, media pipelines editing large files in place - I would provide one managed file service, price it back to teams, and monitor metadata operations rather than only capacity. And I would encode all of this in StorageClasses and templates, not in a wiki page.

code

yaml · 31 lines
yaml
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: standard-block
  annotations:
    storageclass.kubernetes.io/is-default-class: "true"
provisioner: ebs.csi.aws.com
parameters:
  type: gp3
reclaimPolicy: Delete
allowVolumeExpansion: true
volumeBindingMode: WaitForFirstConsumer
---
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: retained-block
provisioner: ebs.csi.aws.com
parameters:
  type: gp3
reclaimPolicy: Retain
allowVolumeExpansion: true
volumeBindingMode: WaitForFirstConsumer
---
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: shared-file-exception
provisioner: efs.csi.aws.com
reclaimPolicy: Retain
allowVolumeExpansion: false

go deeper

for a junior

Focus on the recall layer: RWO is the normal default, RWX needs special storage. The strategic framing is not expected here.

for a middle

Explain the per-replica RWO pattern via volumeClaimTemplates and give the common alternatives to RWX (object storage, ConfigMaps, a real queue).

for a senior

Argue the tradeoffs concretely - blast radius, metadata throughput, locking semantics - and describe the StorageClass settings that make the defaults safe.

for a principal

Frame it as which storage services the platform operates: defaults encoded in classes and templates, an exception path that stays visible, chargeback, SLOs, and metrics that reveal drift toward shared storage.

## The decision is about operational surface, not YAML Every access mode you offer is a commitment: something has to provision it, someone gets paged when it breaks, and every workload that adopts it inherits its failure characteristics. So the question is not "which modes exist" but "which capabilities is the platform willing to operate". ## Default: ReadWriteOnce, one volume per replica This is the default for four reasons. **Universality.** Every CSI driver supports RWO block. Nothing about a workload's portability depends on an unusual capability. **Performance.** A locally attached block device with a local filesystem is the fastest thing available, with no network round trip for metadata operations. **Blast radius.** If one volume degrades, one replica degrades. Contrast with a shared filesystem, whose failure takes every consumer with it. **Shape.** Per-replica volumes force the workload into the shape Kubernetes handles well: a StatefulSet with `volumeClaimTemplates`, stable identities, ordered rollouts, and no multi-attach stalls. Teams that want a rolling update on a single shared volume are describing a workload that Kubernetes will fight; steering them to per-replica volumes fixes the class of problem rather than the instance. ## Single-writer correctness: ReadWriteOncePod RWO is node-scoped, so two pods co-scheduled on one node can both write. For most workloads that is a surprise; for a database it is data loss. Where a second writer is a correctness bug, RWOP (stable in 1.29, requires a CSI driver advertising `SINGLE_NODE_SINGLE_WRITER`) turns a silent corruption into a clean scheduling failure. That trade - fail loudly rather than corrupt quietly - is almost always the right one for stateful singletons, and it belongs in the platform's database templates by default rather than as advice. Be honest about what it does not buy: availability is unchanged, and update strategy still has to avoid overlap. ## ReadWriteMany: an exception with a service behind it When a team asks for RWX, three things are actually being requested: a shared POSIX namespace, a file service to run it, and an implicit coupling between all consumers. The last two are the platform's problem forever. The triage question is what the shared directory is *for*. - **User uploads or generated artefacts shared between replicas** - object storage. This is the most common request and almost never needs a filesystem. Presigned URLs remove the data path from the application entirely. - **Shared read-only assets or config** - bake into the image, or use a ConfigMap, or `ReadOnlyMany` where the backend supports it. Read-only sharing carries none of RWX's coordination cost. - **Coordination through a shared directory (lock files, spool dirs)** - a database or a queue. Worth stating firmly: NFS locking semantics are partial and vary by implementation, so lock-file coordination over RWX is unsafe *even though* the access mode permits it. This is the case where saying yes hurts most. - **Genuinely shared mutable files** - legacy applications that cannot be modified, media pipelines where multiple workers edit large files in place, shared scratch for tightly coupled batch jobs. These are real, and deserve a supported answer. For the genuine cases, pick **one** implementation - a managed file service where the cloud offers one, otherwise a well-understood CephFS or NFS deployment - and support it properly: capacity and metadata-rate monitoring (small-file workloads exhaust metadata throughput long before capacity), backup and restore that has actually been tested, a documented SLO that is honestly lower than block storage's, and chargeback so the cost lands on the team choosing it. ## Encode decisions in StorageClasses Make the right thing the easy thing. A small set of named StorageClasses with a sensible default, a documented `reclaimPolicy` (`Delete` for scratch, `Retain` for anything a mistaken `kubectl delete pvc` should not vaporise), `allowVolumeExpansion: true` so growth is not an outage, and `volumeBindingMode: WaitForFirstConsumer` so topology-constrained volumes are provisioned in a zone where the pod can actually run. Ship golden templates for the common shapes - stateless with no volume, singleton with RWOP, clustered StatefulSet with per-replica RWO - so teams copy a working pattern instead of inventing one. A policy engine can then enforce the boundary: RWX claims require an approved StorageClass, or an annotation recording the exception. The point is not bureaucracy; it is that the platform team learns about every new dependency on the shared file service before it is in production. ## What I would measure Count of RWX claims by team and trend over time (a rising number means the guidance is not landing). Metadata operations per second on the file service against its ceiling. Multi-attach events and stuck-attach durations, which indicate workloads with the wrong shape. Time to restore a volume from snapshot, tested, not assumed. ## The framing that matters Access modes look like a per-claim field, but the platform-level decision is which storage *services* exist. Offer few, make the safest one the default, encode it in classes and templates, and keep the exception path visible enough that you always know who depends on the shared filesystem.

  • A team insists their legacy application cannot be changed and requires a shared POSIX directory across replicas. How do you proceed?
    Accept it as a genuine case rather than arguing, and make it supportable: provision it on the single approved file service, record the exception so the dependency is visible, and check the workload's file-access pattern against the backend's limits - especially small-file metadata rates and any reliance on file locking, which NFS implements only partially. Then set expectations explicitly: a lower storage SLO than block, per-GB chargeback, and a tested restore path.
  • Why make ReadWriteOncePod the default for stateful singletons rather than just documenting the risk?
    Because the ReadWriteOnce failure mode is silent: two pods co-scheduled on one node both write, and the corruption surfaces long after the cause. RWOP converts that into a scheduling failure the team sees immediately. Defaults are the only guidance that scales - documentation is read once and templates are copied forever - so encoding it in the golden template is what actually changes outcomes.

Offering RWX is like adding a shared kitchen to an office block: convenient, and now you own cleaning, capacity, disputes and the outage when the plumbing fails for everyone at once.

saying these in an interview costs you the question

  • Offering ReadWriteMany as a general-purpose option without owning the file service behind it.
  • Assuming the access mode is the whole decision, ignoring StorageClass reclaim policy, expansion and binding mode.
  • Approving RWX for lock-file coordination, which NFS semantics cannot make safe.
  • Treating documentation as a substitute for defaults encoded in StorageClasses and templates.
  • Sizing a shared file service on capacity alone while ignoring metadata operation rates.

context