skip to content

After a new Kubernetes LimitRange lands mid-upgrade on a 48-node cluster, an IoT ingest gateway Deployment keeps losing replicas with no failing Pods visible. How do you diagnose it and roll such a LimitRange out safely?

level: seniorimportance: should knowfreq 36%

answer

  1. replacement Pods never get stored
  2. look at the ReplicaSet
  3. FailedCreate forbidden message
  4. drain recreates Pods
  5. inventory before enforcing

basics

~10 s

Drained Pods are being replaced by Pods that LimitRanger refuses at admission, so they never exist. Read FailedCreate events on the ReplicaSet, fix the bound or the workload, and check existing workloads before enforcing.

solid answer

~50 s

A LimitRange acts only when a Pod is created, so the running gateway Pods, which have a 1Gi limit above the new `max.memory` of 768Mi, stayed up. Each node drain evicted one, and the replacement Pod was refused with `403 Forbidden`. A refused Pod is never stored, so there are no Pod events to find. The evidence is on the ReplicaSet: `kubectl describe rs` shows `FailedCreate ... pods "ingest-gateway-..." is forbidden: maximum memory usage per Container is 768Mi, but limit is 1Gi`, and the Deployment has a `ReplicaFailure` condition. To restore capacity, loosen or remove the LimitRange, since it has no runtime effect, or fit the workload to the bound. To roll one out safely, compare every Pod and template against the new bounds first, dry-run a Pod with `--dry-run=server` in a staging namespace that already holds the new LimitRange, report violations before enforcing, freeze such changes during drains and rollouts, and watch for `FailedCreate` afterwards.

code

bash · 4 lines
bash
kubectl get deploy ingest-gateway -n telemetry-ingest
kubectl describe rs -n telemetry-ingest -l app=ingest-gateway | grep -A3 FailedCreate
kubectl get events -n telemetry-ingest --field-selector reason=FailedCreate
kubectl describe limitrange -n telemetry-ingest

go deeper

for a junior

Remember that a Pod refused at admission never exists, so look at the ReplicaSet's events rather than for a failing Pod.

for a middle

Explain why the running Pods survived while their replacements failed: LimitRanger checks only Pod creation, and a drain turns every eviction into a new creation.

for a senior

Show the whole diagnosis, from ReplicaFailure to the FailedCreate text to comparing the bounds with the workload. Separate LimitRange refusals from Pod-validation refusals caused by an injected default limit.

for a principal

Treat bound changes as contract changes with the namespace owners. Require an inventory and an audit period, and block such changes during cluster-wide disruption windows.

## The situation A platform team applies a new LimitRange to the `telemetry-ingest` namespace with `max.memory: 768Mi` for containers. The IoT telemetry ingest gateway Deployment runs 9 replicas with `limits.memory: 1Gi`. Meanwhile, a version upgrade is draining nodes across a 48-node cluster. Six nodes in, the Deployment shows `READY 5/9`, yet `kubectl get pods` shows nothing in `Pending`, `CrashLoopBackOff` or `Error`. The four missing Pods simply do not exist. ## Why nothing looks broken A LimitRange acts only at **admission**, and only on **Pod** creation (plus PVC checks): - The Deployment and its ReplicaSet were accepted long ago. LimitRanger does not look at them, and it did not re-check them when the LimitRange appeared. - The five surviving Pods predate the LimitRange. Ordinary Pod updates are not re-validated, so they run on with a 1Gi limit that now breaks `max`. - Each drain evicted a gateway Pod. The old Pod spent its **25-second preStop budget** finishing device connections and then exited. The ReplicaSet controller then tried to create a replacement, and kube-apiserver refused it with `403 Forbidden` before it was ever stored. - A refused Pod leaves no Pod object, so there are no Pod events. The only evidence sits on the **ReplicaSet**. The upgrade turned a latent policy violation into lost capacity. Any other event that recreates Pods would do the same: a rollout, a node failure, or a scale-up. ## Diagnosis, step by step 1. `kubectl get deploy ingest-gateway -n telemetry-ingest` shows available below desired. `kubectl describe deploy` shows a `ReplicaFailure` condition copied from the ReplicaSet. 2. `kubectl describe rs <current-replicaset> -n telemetry-ingest` shows `Warning FailedCreate ... Error creating: pods "ingest-gateway-6f8c9d7b54-" is forbidden: maximum memory usage per Container is 768Mi, but limit is 1Gi`. 3. `kubectl get events -n telemetry-ingest --field-selector reason=FailedCreate` finds the same across every controller in the namespace. 4. `kubectl describe limitrange -n telemetry-ingest` (or `kubectl describe namespace telemetry-ingest`) shows the new bound and whether more than one LimitRange is present. 5. Compare the surviving Pods against it, for example with `kubectl get pods -n telemetry-ingest -o 'custom-columns=NAME:.metadata.name,MEM:.spec.containers[*].resources.limits.memory'`. That shows how many more Pods will be refused on their next recreation. The error text names the rule: `minimum ... usage per Container`, `maximum ... usage per Container`, `max limit to request ratio`, or the `Pod`/`PersistentVolumeClaim` variants. A message like `must be less than or equal to cpu limit` is different. It comes from Pod validation after LimitRanger injected a `default` limit below a declared request. ## Restoring capacity - **Fastest**: loosen or delete the LimitRange. It has no runtime effect, so the ReplicaSet controller's next create attempt succeeds once the API server sees the change. - **Correct**: agree on the gateway's real memory need. Then either lower its limit to fit (only if the 1Gi was padding) or raise `max` for this namespace. - Keep an eye on the drain. When the gateway has a PodDisruptionBudget, the missing replicas eventually stop further evictions, which stalls the upgrade on the next node that hosts a gateway Pod. ## Rolling a LimitRange out safely 1. **Inventory first.** List every Pod and every pod template in the namespace, and compare each against the proposed `min`, `max` and `maxLimitRequestRatio`. Also check containers that declare a request above the proposed `default` limit. 2. **Dry-run the refusal.** Apply the proposed LimitRange to a staging namespace first, then run `kubectl create --dry-run=server -f` there on a Pod built from each workload's template. Built-in admission runs on server-side dry-run against the LimitRanges that exist in that namespace, so the output is exactly the error a real create would get. 3. **Report before enforcing.** A policy engine in audit mode can list violations without refusing anything, and that list becomes the work order for the owning teams. 4. **Never change it during a disruption.** Freeze LimitRange changes while upgrades, drains or large rollouts are running, because those are exactly the moments that recreate Pods. 5. **Roll out namespace by namespace**, and treat each rollout as a change to the contract with that namespace's owners. 6. **Watch for `FailedCreate`** across the cluster for a day afterwards. | Mistake | Symptom | |---|---| | `max` below an existing workload's limit | silent capacity loss on the next Pod recreation | | `default` limit below a declared request | `must be less than or equal to ... limit` on new Pods | | `maxLimitRequestRatio` with no `default` limit, and containers that set only a request | refused: no limit is specified | | two `Container` LimitRanges | which defaults apply depends on list order |

  • Why does kubectl apply --dry-run=server on the Kubernetes Deployment not reveal the LimitRange conflict?
    LimitRanger checks Pods and PersistentVolumeClaims, not Deployments. A dry run of the Deployment only admits the Deployment object. To see the refusal in advance, put the LimitRange in a staging namespace and dry-run a Pod built from the template there with `kubectl create --dry-run=server -f pod.yaml`. The Pod admission chain runs against that namespace's LimitRanges and returns the same `forbidden` message a real create would get.
  • A Kubernetes Pod is refused with "must be less than or equal to cpu limit of 300m" right after a LimitRange was added. What caused it?
    It is not a LimitRange bound message. The container declared a CPU request, for example 350m, but no limit, so LimitRanger injected `default.cpu: 300m`. Pod validation then refused a request larger than its limit. The fix is to declare the limit in the workload, or to raise the namespace's `default` above the requests that teams actually declare.

saying these in an interview costs you the question

  • Refused replacement Pods will show up as Pending with scheduling events.
  • The Deployment would have been refused if it broke the LimitRange.
  • Running Pods above the new max are evicted once the LimitRange exists.
  • Deleting the LimitRange requires restarting existing Pods to take effect.
  • A server-side dry run of the Deployment proves its Pods will be admitted.