skip to content

For a database on Kubernetes, what do you give up and gain by using local NVMe PersistentVolumes instead of network-attached block storage?

level: middleimportance: nice to knowfreq 31%

answer

  1. speed versus mobility
  2. local volume needs nodeAffinity
  3. node loss loses that copy
  4. WaitForFirstConsumer binding
  5. replication must supply durability

basics

~20 s

Local NVMe gives much lower, steadier I/O latency, but the volume is tied to one node. If that node is lost, the replica's data goes with it, and the pod cannot move. Database replication has to provide the durability.

solid answer

~40 s

A **local** PersistentVolume points at a disk on one machine and must carry `nodeAffinity` for that node. Every pod using it is pinned there. You gain **latency and throughput**: each I/O stays inside the machine instead of crossing the network, which helps write-heavy databases and log-structured systems such as Kafka. You lose **mobility and durability**. A drain cannot move the pod, so maintenance either waits for the node to return or rebuilds the replica elsewhere. If the node dies, that copy of the data is gone and the claim stays bound to a volume nobody can reach. Local volumes are usually served by a StorageClass with the `kubernetes.io/no-provisioner` provisioner and `volumeBindingMode: WaitForFirstConsumer`. They make sense only when the database replicates itself and you have automated rebuilding a replica.

code

yaml · 27 lines
yaml
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: ocr-local-nvme
provisioner: kubernetes.io/no-provisioner
volumeBindingMode: WaitForFirstConsumer
---
apiVersion: v1
kind: PersistentVolume
metadata:
  name: ocr-nvme-worker-17
spec:
  capacity:
    storage: 1490Gi
  accessModes:
  - ReadWriteOnce
  storageClassName: ocr-local-nvme
  local:
    path: /mnt/nvme0
  nodeAffinity:
    required:
      nodeSelectorTerms:
      - matchExpressions:
        - key: kubernetes.io/hostname
          operator: In
          values:
          - worker-17

go deeper

for a junior

Remember that a local volume is tied to one node, so the pod using it can only run on that node.

for a middle

Explain the latency gain and the loss of mobility, and why local volumes need nodeAffinity and WaitForFirstConsumer binding.

for a senior

Plan node loss and drains: who deletes the stale claim, how long a replica rebuild takes, and whether the operator waits or recreates.

for a principal

Decide whether faster I/O is worth coupling data to machines, given node replacement frequency, rebuild time and the team's recovery maturity.

## Two kinds of disk behind a claim A **PersistentVolumeClaim** asks for storage. What backs it changes how a database behaves on Kubernetes. - **Network-attached block storage** (a cloud block volume, or a SAN or distributed storage system on premises) lives outside the node. The node attaches it over the network, and it can be detached and attached to another node when the pod moves. - A **local PersistentVolume** is a disk or partition on the node itself, often an NVMe drive. The PersistentVolume uses the `local` volume source and **must** have `nodeAffinity` naming the node. The API server rejects a local volume without it. For the document-OCR pipeline on a 3-control-plane, 27-worker self-managed cluster, suppose the Kafka brokers and the 412 GiB PostgreSQL text store are candidates for NVMe-equipped workers. ## What you gain - **Lower latency**: each read or write goes to a device in the same machine. Network storage adds at least one network round trip, and often a replication step inside the storage system. - **Steadier latency**: no contention with other tenants on a shared storage network. - **Higher throughput and IOPS** for the price, since local NVMe is part of the machine. - **Less dependence** on a storage system that can itself fail or slow down. This matters most for write-heavy workloads, databases that call `fsync` often, and log-structured systems. ## What you give up | Concern | Network block storage | Local NVMe PersistentVolume | |---|---|---| | Pod moves to another node | Volume detaches and reattaches | Impossible; pod is pinned | | Node is drained | Pod moves, after a detach wait | Pod waits or replica is rebuilt | | Node is lost | Data survives on the storage system | That replica's data is gone | | Resize or snapshot | Usually supported by the CSI driver | Depends on the provisioner; often not | | Latency | Network hop per I/O | In-machine | | Provisioning | Dynamic from a StorageClass | Pre-created or by a local provisioner | The hardest row is **node loss**. The claim stays bound to a PersistentVolume whose node no longer exists. The pod stays Pending because its only allowed node is gone. Someone, or an operator, has to delete the claim so a new one can be created on another node, and the database has to rebuild that replica from its peers. ## How to set it up 1. Create a **StorageClass** with provisioner `kubernetes.io/no-provisioner` (or a local-volume provisioner) and `volumeBindingMode: WaitForFirstConsumer`, so a claim is bound only after the scheduler has chosen a node for the pod. 2. Create the local PersistentVolumes with `nodeAffinity` for each NVMe node, or let a static provisioner discover the disks. 3. Put the database pods on those nodes with taints and tolerations, and spread replicas across nodes. 4. Make sure the data service **replicates** (PostgreSQL standbys, Kafka replication factor 3) and that rebuilding a lost replica is automated and tested. ## Questions to ask before choosing - How long does rebuilding one replica take at today's data size, and at next year's? - How often are nodes replaced, rebooted or re-imaged, and does that automation wait for the database? - Does the operator recreate a replica on another node by itself, or does a person have to delete the claim? - Is the latency gain measured on this workload, or assumed? If the rebuild takes hours and nodes are replaced weekly, the replica set spends a lot of time degraded, and the speed gain may not be worth it. ## Where it fits - **Good fit**: a replicated database with an operator that can rebuild a replica on another node, a dedicated pool of NVMe workers, and a team that tests node loss. - **Poor fit**: a single-instance database, a service whose only durability is the disk, or a cluster where nodes are replaced often by automation. CloudNativePG shows the tradeoff in its API. Its `nodeMaintenanceWindow` has `reusePVC`: `true` waits for the node to come back and keeps the local data, while `false` recreates the instance, with new storage, on another node when the cluster has more than one instance. It is the same choice between waiting and rebuilding, made explicit.

  • Why does a local-volume StorageClass need `volumeBindingMode: WaitForFirstConsumer`?
    With `Immediate`, the claim could bind to a local volume on a node where the pod cannot run, for example one without enough CPU or with the wrong taints, and the pod would stay Pending. Waiting for the first consumer lets the scheduler choose a node for the pod first and then bind a local volume on that same node.
  • A worker with a local PostgreSQL replica is gone for good. How do you recover that replica?
    Delete the claim that is bound to the unreachable volume, and remove the orphaned PersistentVolume, so a new claim can bind on another node. Then let the database rebuild the replica from the primary. An operator such as CloudNativePG can do this for you. Recovery time depends on the data size and the network, so size it before choosing local disks.

saying these in an interview costs you the question

  • A local PersistentVolume can move with the pod to another node.
  • Local NVMe is safe for a single-instance database because disks rarely fail.
  • A local volume works without nodeAffinity if the path exists everywhere.
  • Network block storage is always faster because it is SSD-backed.
  • Draining a node with local volumes reschedules the database elsewhere automatically.