skip to content

Distinguish two Kubernetes failures: a pod Pending with `pod has unbound immediate PersistentVolumeClaims`, versus a Deployment whose replica count never rises and for which no pod object is ever created. What causes each, and how do you confirm it?

level: seniorimportance: should knowfreq 48%

answer

  1. pod exists + Pending → scheduler; no pod → admission
  2. Immediate vs WaitForFirstConsumer binding mode
  3. WFFC Pending is normal — read the POD's events
  4. describe rs → FailedCreate → quota / LimitRange / webhook
  5. quota on a resource ⇒ every pod must specify it

basics

~20 s

An unbound PVC means storage was never provisioned — check the PVC events and StorageClass. A missing pod entirely means admission rejected creation: describe the ReplicaSet for FailedCreate naming a ResourceQuota, LimitRange, or webhook. One is a scheduling problem; the other happens before scheduling.

solid answer

~60 s

**Unbound PVC (pod exists, Pending).** The scheduler will not bind a pod whose PVC is unbound with `volumeBindingMode: Immediate`. Run `kubectl get pvc` and `kubectl describe pvc`: typical causes are a StorageClass name that does not exist or a missing default class (`no persistent volumes available for this claim and no storage class is set`), a provisioner that is failing (`ProvisioningFailed`), quota exhaustion at the cloud provider, or no matching static PV. With `volumeBindingMode: WaitForFirstConsumer` the PVC stays `Pending` **by design** until a pod is scheduled — the scheduler picks the node first, then the volume is provisioned in that node's zone. Seeing `waiting for first consumer to be created` is normal; if it persists, the pod is blocked for some *other* reason. **No pod at all.** Scheduling never entered the picture: pod creation was rejected. `kubectl describe replicaset <rs>` shows `FailedCreate`, e.g. `exceeded quota: team-a, requested: requests.cpu=2, used: 18, limited: 20`, a LimitRange violation, a Pod Security admission denial, or an admission webhook error. Check `kubectl get resourcequota -n <ns> -o yaml` for used-versus-hard.

code

bash · 5 lines
bash
kubectl get pods -n team-a            # does a pod exist at all?
kubectl describe pvc data-api-0 -n team-a
kubectl get storageclass
kubectl describe rs -n team-a -l app=api | sed -n '/Events/,$p'
kubectl describe resourcequota -n team-a

go deeper

for a junior

Know that a pod will not start while its PVC is unbound, and that kubectl describe pvc shows why binding failed.

for a middle

Explain Immediate versus WaitForFirstConsumer binding and why WFFC exists, and know that ResourceQuota rejections appear as FailedCreate on the ReplicaSet.

for a senior

Drive the triage from the pod-exists-or-not boundary, read CSI provisioning errors, recognise the quota-forces-requests rule, and check webhooks and Pod Security as admission causes.

for a principal

Set namespace policy so these failures are self-explaining: LimitRange defaults alongside every quota, storage classes standardised with WFFC for zonal volumes, webhook failurePolicy chosen deliberately, and alerting on FailedCreate/ProvisioningFailed events.

## Two failures that look similar and are not Both present as "my workload is not running", but they sit on opposite sides of the scheduling boundary. Establishing which one you have is the first move: **does a pod object exist?** ``` kubectl get pods -n team-a ``` If a `Pending` pod is listed, the API accepted it and the scheduler is stuck. If nothing is listed, the pod was never created and the scheduler has never seen it. ## Case 1: the pod exists and its PVC is unbound A **PersistentVolumeClaim (PVC)** is a request for storage; a **PersistentVolume (PV)** is the actual storage. Binding is the matching of one to the other, either by finding a suitable pre-created PV or by **dynamic provisioning** through a StorageClass, which names a provisioner (a CSI driver) that creates the volume on demand. ### volumeBindingMode changes everything - **`Immediate`** — the volume is provisioned and bound as soon as the PVC is created, before any pod exists. The scheduler treats an unbound Immediate PVC as a hard blocker: `pod has unbound immediate PersistentVolumeClaims`. The problem is entirely in the storage layer. - **`WaitForFirstConsumer` (WFFC)** — binding is deliberately deferred until a pod that uses the PVC is being scheduled. The PVC sitting in `Pending` with `waiting for first consumer to be created` is **correct behaviour, not a fault**. WFFC exists because zonal storage cannot be moved: binding a volume in zone `a` before scheduling risks the scheduler wanting zone `b`, producing the classic `node(s) had volume node affinity conflict`. WFFC lets the scheduler choose a node first, then provisions the volume there. A very common misdiagnosis: someone sees a WFFC PVC `Pending` and chases the storage layer, when in fact the pod cannot be scheduled for an unrelated reason (insufficient CPU, a taint), so the volume is never provisioned. Read the **pod's** events, not the PVC's, in that case. ### Diagnosing an Immediate binding failure ``` kubectl get pvc -n team-a kubectl describe pvc data-api-0 -n team-a kubectl get storageclass ``` What the events tell you: - `no persistent volumes available for this claim and no storage class is set` — the PVC named no class and the cluster has no default class, or `storageClassName: ""` disabled dynamic provisioning and no matching static PV exists. - `storageclass.storage.k8s.io "fast-ssd" not found` — a typo or a class that exists only in another cluster. - `ProvisioningFailed` with a driver message — the CSI driver could not create the volume: cloud quota exhausted, wrong zone, IAM permissions missing, unsupported size or access mode. - Nothing at all — the CSI controller pods may be down; check them. Also check **access modes**: a `ReadWriteMany` claim against a block-storage provisioner that only supports `ReadWriteOnce` will never bind. And note that resizing and access-mode mismatches produce distinct events worth reading rather than guessing. ## Case 2: no pod object at all Here the failure is at **admission**, before the object is persisted. The controller that would create the pod records the rejection. ``` kubectl get deploy api -n team-a kubectl describe rs $(kubectl get rs -n team-a -l app=api -o name | tail -1) -n team-a kubectl get events -n team-a --sort-by=.lastTimestamp ``` The ReplicaSet event is `FailedCreate`, and the message names the cause: - **ResourceQuota** — `exceeded quota: team-a, requested: requests.cpu=2, used: requests.cpu=18, limited: requests.cpu=20`. A `ResourceQuota` caps a namespace's aggregate requests, limits, and object counts. Inspect it with `kubectl describe resourcequota -n team-a`, which prints Used versus Hard per resource. A critical wrinkle: **if a quota constrains a resource, every pod in that namespace must specify it** — a pod with no CPU request is rejected outright with `must specify requests.cpu`. This is why a namespace works fine until someone adds a quota and unrelated deployments start failing. - **LimitRange** — sets defaults and min/max per pod/container in the namespace. A container requesting more than the max, or less than the min, is rejected. LimitRange also *injects* defaults, which is the usual remedy for the quota wrinkle above. - **Pod Security admission** — `violates PodSecurity "restricted:latest"` when the pod requests privileges the namespace's policy forbids. - **Validating/mutating webhooks** — a policy engine (image-signature checks, required labels) denying the pod, or a webhook that is simply unreachable, which fails closed if `failurePolicy: Fail`. Note the asymmetry with Deployments: a Deployment's own creation succeeds because quota is enforced on **pods**, not on the Deployment object. So the Deployment exists and looks healthy while its ReplicaSet quietly fails to create pods — you must look one level down. ## The triage rule 1. Pod exists and is Pending → scheduler-side. Read the pod's events; if they point at volumes, move to the PVC and StorageClass. 2. PVC Pending with WFFC → check the *pod's* scheduling events first; the storage is waiting on placement. 3. No pod exists → admission-side. Describe the ReplicaSet (or Job/StatefulSet) for `FailedCreate` and read the quota/LimitRange/webhook message verbatim. Being able to state that boundary crisply — objects that exist versus objects that were never allowed to exist — is what the question is really testing.

  • A PVC has been Pending for an hour with `waiting for first consumer to be created`, and its pod is also Pending. Where is the real problem?
    Almost certainly in pod scheduling, not storage. With `WaitForFirstConsumer` the volume is only provisioned once the scheduler picks a node, so the PVC message is expected. Read the pod's own FailedScheduling event — insufficient CPU or memory, an untolerated taint, or an affinity mismatch is blocking placement, and the PVC will bind as soon as the pod can be placed.
  • After an administrator adds a ResourceQuota to a namespace, unrelated Deployments stop creating pods. Why?
    If a quota constrains `requests.cpu` or `requests.memory`, every pod in that namespace must explicitly specify that resource; pods without it are rejected at admission with a message like `must specify requests.cpu`. Existing manifests that omitted requests suddenly become invalid. The standard fix is a LimitRange in the namespace that injects default requests and limits.
  • Why does the Deployment object itself look healthy when quota blocks pod creation?
    Quota is enforced against pods at admission, not against Deployments, so the Deployment and its ReplicaSet are created successfully. Only the pod creation call is rejected, and the failure is recorded as a `FailedCreate` event on the ReplicaSet. That is why you have to describe the ReplicaSet rather than the Deployment to see the real error.

saying these in an interview costs you the question

  • Treating a WaitForFirstConsumer PVC in Pending as a storage fault
  • Looking for a Pending pod when admission rejected pod creation and no pod exists
  • Describing the Deployment instead of the ReplicaSet when hunting a FailedCreate
  • Not knowing that a quota on a resource forces every pod to specify that resource
  • Assuming any StorageClass can satisfy ReadWriteMany regardless of the underlying provisioner

context