When a pod in a StatefulSet is rescheduled to a different node, how does it end up with the same data it had before, and what can prevent that from happening?
answer
- Same ordinal -> same PVC -> same PV
- Detach from old node before attach
- Multi-Attach = old node unresolved
- Zonal disk pins the pod to its zone
- New PV name = re-provisioned, not reattached
basics
~20 sThe replacement pod keeps the same ordinal name, so it binds the same ordinal-suffixed PersistentVolumeClaim; the volume is detached from the old node and attached to the new one. Zone pinning, per-node attach limits, and a stuck detach from an unreachable node can block it.
solid answer
~60 sIdentity does the work. The StatefulSet controller recreates the deleted pod with the same name and ordinal, and the ordinal-suffixed PVC (`data-pg-1`) already exists and stays Bound to its PersistentVolume. The new pod references the same claim, so Kubernetes detaches the volume from the old node and attaches it to the new one, then stages and mounts it. Nothing is copied - the same disk moves. What blocks it in practice: - **Topology.** A zonal block volume can only attach to nodes in its zone. If the scheduler has no fitting node there, the pod stays Pending with a volume node affinity conflict. - **Unreachable old node.** With a ReadWriteOnce volume, the old attachment must be released first. A node that goes NotReady leaves the pod Terminating and the volume attached, producing `Multi-Attach error` until the node is removed or force-detached. - **Attach limits** on the target node, or a broken CSI node plugin there. StatefulSets are deliberately conservative here: it will not start a replacement while the old pod's state is unknown, because two writers on one filesystem is worse than downtime.
code
bash · 4 lineskubectl get pod pg-1 -o wide
kubectl get pvc data-pg-1 -o jsonpath='{.spec.volumeName}{"\n"}'
kubectl get volumeattachments -o wide | grep pvc-1a2b
kubectl describe pod pg-1 | grep -A5 Eventsgo deeper
Say that the replacement pod has the same name, so it gets the same PVC and therefore the same data.
Add detach-before-attach for ReadWriteOnce, zone pinning of the PV, and how to verify the PV name did not change.
Diagnose the blockers - unreachable node and Multi-Attach, node affinity conflict, attach limits, unhealthy node plugin - and apply the safe remediations rather than force-deleting pods.
Weigh reattachment against application-level re-replication: recovery time, zone-bound placement, capacity headroom per zone, and how those choices shape the failure domain of the data tier.
## The mechanism Three stable facts combine. 1. **Stable pod identity.** A StatefulSet pod is named `<sts>-<ordinal>`. If `pg-1` is deleted, the controller creates a new pod also called `pg-1` - not a random suffix as a Deployment would. 2. **Deterministic claim names.** The claim generated from the template is `data-pg-1`. The replacement pod therefore resolves to the same PVC object. 3. **PVC/PV binding is durable.** The claim stays Bound to its PersistentVolume across pod churn; the PV holds the backend volume handle. So reattachment is not a data copy or a sync - it is the same block device being detached from node A and attached to node B, then staged (mounted at a global staging path) and published (bind-mounted into the pod directory) by the CSI node plugin on node B. ## Why the old attachment must be released first Most block volumes are `ReadWriteOnce`: attachable to one node at a time. Kubernetes will not attach elsewhere until the previous attachment is gone, and it treats that seriously because the failure mode is filesystem corruption, not a restart. The painful case is a node that stops responding. Its kubelet may still be alive and writing; the control plane cannot tell. So: - pods on it go `Terminating` and stay there, - the `VolumeAttachment` remains, - the StatefulSet does **not** create the replacement pod, because that would risk two pods with the same identity writing one volume. Resolutions: delete the Node object once you are certain the machine is gone (cloud controller managers do this automatically for terminated instances), or apply the non-graceful node shutdown taint (`node.kubernetes.io/out-of-service`) so Kubernetes force-detaches and lets the pod move. Force-deleting the pod alone (`--force --grace-period=0`) removes the API object without proving the container stopped - it is the tempting wrong answer. ## Topology: the other common blocker Cloud block volumes are zonal. The PV carries `nodeAffinity` restricting it to that zone, and the scheduler honours it. If zone A has no schedulable capacity, `pg-1` is Pending with `node(s) had volume node affinity conflict` - and no amount of spare capacity in zones B and C helps, because the disk cannot follow. That has two design consequences: - use `volumeBindingMode: WaitForFirstConsumer` so each replica's disk is created in the zone its pod actually landed in, spreading replicas and their disks across zones instead of pinning them all to one; - keep enough headroom per zone (or a node group per zone that can scale) so a lost node can be replaced in the same zone. ## Other things that block the move - **Per-node attach limits.** Cloud instances cap attached volumes; the limit is reported in the `CSINode` object and enforced by the scheduler. A node hosting many stateful pods can refuse more. - **CSI node plugin unhealthy on the target node.** Attach succeeds, mount never does; the pod sits in `ContainerCreating` with mount errors on that node only. - **Deleted PVC.** If someone deleted `data-pg-1` while the StatefulSet was scaled down, recreating the pod provisions a fresh empty volume, and the replica silently starts with no data. For a database that is a real incident, usually masked as "the replica is re-syncing". ## Verifying it in practice ``` kubectl get pod pg-1 -o wide # which node now kubectl get pvc data-pg-1 # still Bound, same PV kubectl get volumeattachments | grep <pv> kubectl describe pod pg-1 | tail -20 # attach/mount events ``` If the PV name under `data-pg-1` changed, you did not reattach - you re-provisioned, and the old data is either orphaned or gone. ## The judgement to voice Reattachment gives fast recovery when the volume can follow the pod, but it couples failover to the storage layer: recovery time includes detach plus attach plus filesystem mount plus engine recovery, and it is bounded by zone. Systems that replicate at the application layer (Kafka, Cassandra, most modern databases) can instead bootstrap a fresh replica on any node, trading rebuild time and network for freedom of placement. Real designs use both: reattach when possible, re-replicate when the zone is gone.
- A StatefulSet pod is Pending with 'node(s) had volume node affinity conflict'. What is wrong and how do you fix it?Its PersistentVolume is pinned to one zone and no schedulable node in that zone fits the pod - often after a zone-wide capacity crunch or a node group scaled to zero there. Short term you restore capacity in that zone, for example by scaling the zone's node group. Long term, provision with WaitForFirstConsumer so disks follow pod placement, spread replicas across zones, and keep per-zone headroom for one node failure.
- Why won't the StatefulSet controller just start a replacement pod when a node goes NotReady?Because NotReady means the control plane lost contact, not that the workload stopped. The old kubelet may still be running the container and writing to a ReadWriteOnce volume. Starting a second pod with the same identity could give two writers on one filesystem, so Kubernetes waits until the node's state is resolved - by node deletion, by the out-of-service taint, or by the node returning.
saying these in an interview costs you the question
- Believing the data is copied or synced to the new node
- Force-deleting a Terminating pod as the standard fix for a stuck volume
- Assuming a rescheduled pod can land in any zone regardless of where its disk lives
- Not noticing that a new PV name means a fresh empty volume, not reattachment
- Ignoring per-node attach limits when packing many stateful pods onto few nodes