skip to content

A PersistentVolumeClaim in Kubernetes has a volumeMode field that accepts Filesystem or Block. What is the difference, how does a pod consume each, and when is Block the right choice?

level: middleimportance: should knowfreq 34%

answer

  1. Filesystem = format + mount, volumeMounts/mountPath
  2. Block = raw device, volumeDevices/devicePath
  3. Block users: Ceph OSD, KubeVirt disks, raw-device DB engines
  4. no subPath, no fsGroup, self-managed resize
  5. volumeMode is immutable once bound; driver must support it

basics

~20 s

Filesystem (the default) means Kubernetes formats and mounts the volume at a directory given by volumeMounts. Block exposes the raw device with no filesystem, surfaced at a path via volumeDevices; the application must handle the device itself. Block suits databases and storage systems that manage their own layout.

solid answer

~60 s

`spec.volumeMode` chooses what the pod actually receives. **`Filesystem`** (default): the storage stack ensures a filesystem exists - formatting the device on first use if needed - and the kubelet mounts it into the container at the path in `volumeMounts`. The container sees an ordinary directory. This is what nearly every workload wants. **`Block`**: no formatting and no mount. The raw device is surfaced inside the container at the path given by `volumeDevices[].devicePath`, e.g. `/dev/xvda`. The application must read and write the device directly, or format it itself. Block is worth it when the application manages its own on-disk layout and the filesystem layer is overhead or an obstacle: databases with direct-I/O storage engines, software-defined storage such as Ceph OSDs, and virtualisation platforms backing VM disks. The constraints follow from having no filesystem: `subPath` does not apply, `fsGroup` and filesystem ownership handling do nothing, filesystem-level resize is not performed for you, and the container needs the privileges to open a block device. The mode must be supported by the driver and is fixed at PV creation - you cannot flip a bound volume between modes.

code

yaml · 28 lines
yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: raw-data
spec:
  accessModes:
    - ReadWriteOnce
  volumeMode: Block
  resources:
    requests:
      storage: 100Gi
  storageClassName: fast-block
---
apiVersion: v1
kind: Pod
metadata:
  name: store
spec:
  containers:
    - name: engine
      image: example/engine:1.0
      volumeDevices:
        - name: data
          devicePath: /dev/xvda
  volumes:
    - name: data
      persistentVolumeClaim:
        claimName: raw-data

go deeper

for a junior

Know that Filesystem is the default and gives a mounted directory, while Block gives a raw device consumed through volumeDevices.

for a middle

Add the mechanics: who formats, which pod-spec field applies, and the loss of subPath, fsGroup and automatic filesystem resize.

for a senior

Justify it for real workloads (Ceph OSDs, KubeVirt, raw-device engines) and cover the operational cost: backup, privileges, immutability, driver support.

for a principal

Decide whether raw block is offered at all: which StorageClasses expose it, the privilege model it demands, and whether the performance case justifies the loss of standard backup and inspection tooling.

## Two ways to hand storage to a container `volumeMode` appears on both PersistentVolumes and PersistentVolumeClaims, and a claim binds only to a PV with a matching mode. **Filesystem** is the default and the one almost everyone uses. The provisioner creates the volume, the attach layer makes the device visible on the node, and the kubelet checks whether it carries a filesystem. If it does not, it formats it (ext4 by default for most drivers; a StorageClass parameter such as `csi.storage.k8s.io/fstype` selects otherwise), then mounts it and bind-mounts it into the container at the `volumeMounts[].mountPath`. Inside the container it is simply a directory. **Block** skips both the formatting and the mounting. The device itself is made available to the container at a path you specify under `volumeDevices[].devicePath` - not under `volumeMounts`, which is the syntactic tell. What the container sees at that path is a block special file. There is no directory, no inode structure, nothing to `ls`. ## What the pod spec looks like Filesystem: ```yaml volumeMounts: - name: data mountPath: /var/lib/app ``` Block: ```yaml volumeDevices: - name: data devicePath: /dev/xvda ``` The two lists are mutually exclusive per volume, and using the wrong one for the mode makes the pod fail to start. ## Why anyone wants raw block A filesystem is an abstraction with a cost: page cache and write-back behaviour you do not control, journaling that duplicates the work an application's own write-ahead log already does, metadata operations that add latency, and semantics that make direct I/O and precise durability guarantees awkward. Software that already implements its own storage engine can go faster and be more predictable without it. The genuine users are: - **Databases with raw-device support** - some engines can address a device directly and manage their own extents and caching. - **Software-defined storage** - Ceph OSDs, for example, consume raw devices and lay out their own format; running them on a filesystem inside a filesystem is wasteful. - **Virtualisation platforms** such as KubeVirt, where the volume backs a virtual machine's disk. The guest OS puts its own filesystem inside; wrapping that in a host filesystem adds a pointless layer. - **Performance and benchmarking work** where you want to measure the device, not the filesystem. For a normal application that writes files, Block offers nothing and costs a lot of complexity. ## Constraints that come with Block - **`subPath` is meaningless.** There is no path structure inside a raw device to select. - **`fsGroup` and filesystem ownership/permission handling do not apply.** Access is governed by the device node's permissions and the container's capabilities; workloads typically need elevated privileges to open the device. - **Resizing is only half automatic.** Expanding the underlying volume grows the device; there is no filesystem for Kubernetes to grow, so anything that needs to notice the new capacity is the application's problem. - **Backup and inspection are harder.** You cannot exec in and copy files. Snapshots (via CSI VolumeSnapshots) or application-level dumps become the mechanism. - **Driver support is required.** The CSI driver must advertise block volume support; not all do, particularly file-based ones. It makes no sense at all for NFS-style backends, which *are* filesystems. - **The mode is immutable.** It is set at PV and PVC creation and cannot be changed on a bound volume; switching means creating a new volume and migrating the data. ## How it interacts with access modes The two fields are orthogonal but often confused. `accessModes` says how many nodes may attach the volume and in what direction; `volumeMode` says whether the pod receives a mounted filesystem or a raw device. A raw block volume is usually `ReadWriteOnce` because the backends that offer raw block are block backends. In principle a block device can be attached to several nodes (`ReadWriteMany` block), but writing to it from multiple nodes is only safe when the consumer is cluster-aware - a clustered filesystem or a storage system with its own coordination - because ordinary filesystems assume exclusive ownership and will corrupt data otherwise. ## The interview shape Say what each mode hands to the container and which pod-spec field consumes it, give one or two honest use cases for Block, and state at least two of the constraints (no `subPath`, no `fsGroup`, self-managed resize, immutability). That demonstrates you know it exists and also know it is not a general-purpose optimisation.

  • Can you switch an existing PersistentVolumeClaim from volumeMode Filesystem to Block?
    No. volumeMode is set when the PV and PVC are created and is immutable on a bound volume, and a claim only binds to a PV whose mode matches. Changing it means provisioning a new volume with the desired mode and migrating the data at the application level - there is no in-place conversion, since the two representations of the data are fundamentally different.
  • Why do fsGroup and subPath have no effect on a raw block volume?
    Both are filesystem concepts. fsGroup works by changing group ownership and permissions of files on a mounted filesystem, and subPath selects a subdirectory within one. A raw block volume has no filesystem for Kubernetes to walk, so there is nothing to chown and no subdirectory to select; access is governed by the device node and the container's privileges instead.

Filesystem mode hands you a furnished flat - walls, shelves, a place for everything. Block mode hands you the empty shell and the keys: more freedom, and every fitting is now your job.

saying these in an interview costs you the question

  • Thinking Block mode is a general performance win for ordinary file-writing applications.
  • Putting a raw block volume under volumeMounts instead of volumeDevices.
  • Expecting fsGroup or subPath to work on a block volume.
  • Assuming Kubernetes will grow anything for you after expanding a raw block volume.
  • Believing volumeMode can be changed on a bound claim.

context