skip to content

After a ResourceQuota is added to its Kubernetes namespace, a feature-flag evaluation service's Deployment rollout stalls with no new pods, although kubectl apply succeeded. How do you diagnose and fix it?

level: seniorimportance: must knowfreq 52%

answer

  1. apply stores the Deployment only
  2. no pod exists, not Pending
  3. FailedCreate on the ReplicaSet
  4. must specify vs exceeded quota
  5. replicas plus surge times request

basics

~20 s

Quota rejects pods when the ReplicaSet controller creates them, so the error appears as FailedCreate events on the ReplicaSet, not at apply time. It is usually 'must specify' (containers lack requests or limits) or 'exceeded quota' (no headroom for surge pods).

solid answer

~40 s

`kubectl apply` only created or updated the Deployment. The pods are created later by the ReplicaSet controller, and that is when the `ResourceQuota` admission plugin rejects them, so no pod exists, not even a `Pending` one. I would run `kubectl describe replicaset` for the new ReplicaSet (or check the Deployment's `ReplicaFailure` condition) and read the `FailedCreate` event. `failed quota: ... must specify limits.memory for: flag-eval` means the quota caps memory and a container declares no memory limit; the fix is to declare it or add LimitRange defaults. `exceeded quota: ... requested: requests.cpu=700m, used: requests.cpu=6300m, limited: requests.cpu=6800m` means there is no room for a surge pod; the fix is to raise the quota to cover replicas plus surge, or to allow `maxUnavailable` so old pods free headroom. `kubectl describe resourcequota` confirms Used against Hard.

code

bash · 3 lines
bash
kubectl rollout status deployment/flag-eval -n flags
kubectl get events -n flags --field-selector reason=FailedCreate
kubectl describe resourcequota -n flags

go deeper

for a junior

Remember that Deployments create pods through ReplicaSets, so a quota error shows up as a FailedCreate event, not as an apply error.

for a middle

Explain the two rejection messages: the CPU/memory declaration rule and the used-plus-requested check, including why LimitRange defaults satisfy the first.

for a senior

Work the surge arithmetic from Used, Hard and the rollout strategy, and choose between raising the cap, allowing unavailability or resizing requests.

for a principal

Treat these stalls as a platform signal: quotas should be sized for peak rollout plus autoscaling headroom, and quota changes should be announced like API changes.

## Why kubectl apply looked fine A **Deployment** does not create pods itself. `kubectl apply` stores the Deployment. The Deployment controller then creates or scales a **ReplicaSet**, and the ReplicaSet controller creates the **pods**. A **ResourceQuota** is enforced by the `ResourceQuota` **admission plugin** in kube-apiserver on each create, so the rejection hits the ReplicaSet controller's pod create, several steps after your `apply` returned success. That has two visible effects: - **No pod object exists.** A rejected create is never stored, so `kubectl get pods` shows nothing new. That differs from a scheduling problem, where a `Pending` pod exists with an event explaining why. - **The error is recorded on the ReplicaSet.** The controller emits a `Warning` event with reason `FailedCreate` and a message starting `Error creating:`, and the Deployment gains a `ReplicaFailure` condition. If nothing changes, the Deployment's `Progressing` condition eventually turns false once `progressDeadlineSeconds` (600 by default) passes. ## Diagnosis, step by step 1. `kubectl rollout status deployment/flag-eval -n flags` to confirm the rollout is stuck rather than slow. 2. `kubectl describe deployment flag-eval -n flags` to read the conditions and find the new ReplicaSet's name. 3. `kubectl describe replicaset <new-rs> -n flags`, or `kubectl get events -n flags --field-selector reason=FailedCreate`, to read the exact rejection. 4. `kubectl describe resourcequota -n flags` to compare `Used` with `Hard` for every quota in the namespace; a pod must fit under all of them. ## The two messages you will see | Message fragment | Meaning | Fix | |---|---|---| | `failed quota: flags-compute: must specify limits.memory for: flag-eval` | the quota caps a CPU or memory key and container `flag-eval` does not declare it | declare the value in the pod template, or add LimitRange defaults | | `exceeded quota: flags-compute, requested: requests.cpu=700m, used: requests.cpu=6300m, limited: requests.cpu=6800m` | the pod would take usage past the cap | raise the cap, reduce requests, or change the rollout strategy | ### "must specify": the declaration rule When a quota caps `cpu`, `memory`, or any `requests.`/`limits.` form of them, **every container and init container** in a new pod it matches must declare that exact value. `limits.memory` in the quota requires a memory **limit**, not just a request. The rule covers only CPU and memory: quotas on `requests.ephemeral-storage` or on extended resources do not force containers to declare them. Two things can supply the value: - **The pod template itself.** This is the most explicit fix. - **A LimitRange with defaults.** The `LimitRanger` admission plugin mutates pods before quota validation runs, so injected defaults count. How defaults are chosen belongs to LimitRange. One exception: when a pod sets pod-level `spec.resources` (the `PodLevelResources` feature, beta and on by default since 1.34), the per-container declaration check is skipped. ### "exceeded quota": surge arithmetic Take the incident numbers: `flag-eval` runs 9 replicas at 700m CPU request each, so `Used` is 9 × 700m = **6300m** against a `requests.cpu` cap of **6800m**. - The strategy is `maxSurge: 25%` and `maxUnavailable: 0`. 25% of 9 is 2.25; surge rounds **up**, so the controller wants 3 extra pods. - The first surge pod would need 6300m + 700m = **7000m**, which is above 6800m, so it is rejected. - With `maxUnavailable: 0`, no old pod may be removed until a new one is ready. Nothing can move, and the rollout stalls. Your fixes: - **Raise the cap** to cover the peak: (9 + 3) × 700m = **8400m**. - **Allow unavailability**: `maxUnavailable: 1` lets an old pod go first and frees 700m for each new pod, at the cost of capacity during the rollout. - **Reduce the request** if 700m is oversized, which is a sizing decision outside quota. The same rejection appears when a HorizontalPodAutoscaler raises `replicas` past what the quota allows: the ReplicaSet controller's creates fail in the same way. ## Less common causes - **`status unknown for quota`** right after a quota is created: the ResourceQuota controller has not filled in `status` yet. It is transient. - **Another quota** in the namespace, for example a `count/replicasets.apps` or `pods` cap, rejects the create even though the compute quota has room. - **An object-count cap on the Deployment's kind** would fail the `kubectl apply` itself, which is how you tell it apart from the pod-level failures above.

  • The quota caps only requests.cpu. Must containers now also declare a memory request?
    No. The declaration rule applies only to the CPU and memory keys the matching quota actually caps. A quota with only `requests.cpu` requires each container to declare a CPU request and nothing else. Adding `limits.memory` later would start rejecting any container without a memory limit, which is why tightening a quota can break Deployments that were fine yesterday.
  • How is a quota rejection different from a pod stuck Pending for lack of node capacity?
    A quota rejection happens in kube-apiserver before the pod is stored, so there is no pod object; the evidence is a `FailedCreate` event on the ReplicaSet. A capacity problem produces a real pod in `Pending`, with a `FailedScheduling` event from kube-scheduler. Having quota left does not guarantee capacity, so both can happen in the same namespace.
  • Do finished Job pods or pods stuck deleting block new pods under quota?
    Pods in phase `Succeeded` or `Failed` are not charged. A pod whose deletion grace period has already passed, for example on a lost node, also stops being charged, so it does not block replacements. A pod still inside its grace period is counted, so a slow shutdown can briefly hold quota during a rollout.

saying these in an interview costs you the question

  • The failure should appear as an error from kubectl apply.
  • Look for a Pending pod to find the quota error.
  • A limits.memory quota is satisfied by a memory request alone.
  • Quota allows a temporary overrun while a rollout is in progress.
  • A quota on GPUs forces every container to declare a GPU value.
  • The fix is always to delete the quota and recreate it.