A platform team wants every namespace to have sensible per-container resource defaults plus a hard cap on total consumption. Explain what the Kubernetes LimitRange and ResourceQuota objects each do, and how they interact.
answer
- LimitRange = per object; Quota = per namespace total
- default / defaultRequest / min / max / maxLimitRequestRatio
- quota on a resource → declaring it becomes mandatory
- quota = admission only, never evicts, counts declared not used
- FailedCreate lands on the ReplicaSet, not the Deployment
basics
~20 sLimitRange is per-object: it defaults and bounds an individual container's or Pod's requests and limits at admission. ResourceQuota is per-namespace aggregate: it caps total requests, limits and object counts. A quota on a resource makes declaring it mandatory; LimitRange defaults satisfy that.
solid answer
~50 s**LimitRange** is namespace-scoped and acts on *each* object. It can supply `defaultRequest` and `default` (limit) for containers that omit them, enforce `min`/`max` bounds, cap `maxLimitRequestRatio` to stop wild overcommit, and set min/max storage for PVCs. It is a mutating-then-validating admission step: defaults are injected, then bounds are checked, and violating Pods are rejected. **ResourceQuota** is namespace-scoped and acts on the *sum*. It caps `requests.cpu`, `requests.memory`, `limits.cpu`, `limits.memory`, storage, and object counts such as `pods`, `services`, `configmaps`. Scope selectors let you write different quotas for terminating versus long-running Pods, or per PriorityClass. The interaction is the classic gotcha: **if a quota constrains a resource, every Pod in that namespace must declare it**, or creation is rejected. A LimitRange supplying defaults is what makes that livable. Also, quota is enforced only at admission — lowering a quota never evicts running Pods, it just blocks new ones.
code
yaml · 32 linesapiVersion: v1
kind: LimitRange
metadata:
name: defaults
namespace: team-a
spec:
limits:
- type: Container
defaultRequest:
cpu: "100m"
memory: "128Mi"
default:
cpu: "500m"
memory: "512Mi"
max:
cpu: "4"
memory: "8Gi"
maxLimitRequestRatio:
cpu: "4"
---
apiVersion: v1
kind: ResourceQuota
metadata:
name: team-a-quota
namespace: team-a
spec:
hard:
requests.cpu: "20"
requests.memory: 40Gi
limits.cpu: "40"
limits.memory: 80Gi
pods: "100"go deeper
State the split — LimitRange sets per-container defaults and bounds, ResourceQuota caps the namespace total — and that both are namespace-scoped.
Add the fields that matter (defaultRequest, default, max, maxLimitRequestRatio; requests/limits keys and object counts) and the mandatory-declaration interaction.
Explain the operational failure modes: silent FailedCreate on ReplicaSets, stalled HPA scale-ups, admission-only enforcement, and alerting on quota utilisation.
Position it as multi-tenancy policy — how namespaces are bootstrapped, how quota interacts with PriorityClass and cluster capacity planning, and where quota should give way to cost showback instead.
## Two objects, two scopes Both objects live inside a namespace and both are enforced by admission controllers in the API server, but they answer different questions. LimitRange answers "is this *one* container reasonable, and what should it get if the author said nothing?" ResourceQuota answers "has this namespace consumed more than its share *in total*?" ## LimitRange in detail A LimitRange contains a list of `limits` entries, each with a `type` — `Container`, `Pod`, or `PersistentVolumeClaim` — and any of these fields: - `defaultRequest`: the request injected when the container omits one. - `default`: the limit injected when the container omits one. - `min` / `max`: the inclusive bounds a container (or Pod total) may declare. - `maxLimitRequestRatio`: the largest allowed limit-to-request ratio, which is how you stop a team declaring a 100m request with a 16-core limit and destabilising nodes. Admission runs in two phases. First, defaulting mutates the Pod spec, filling in anything missing. Second, validation checks the resulting values against min/max/ratio. Rejection surfaces as an API error at Pod create time, so a Deployment whose template violates a LimitRange shows up as failing Pod creation in the ReplicaSet's events rather than as a Deployment error — worth knowing when debugging "my Deployment has zero pods and no message". One subtlety: LimitRange defaults apply only to containers that omit the field, and interact with the general Kubernetes rule that a missing request defaults to the limit — the LimitRange's `defaultRequest` is applied first when present. ## ResourceQuota in detail A ResourceQuota's `spec.hard` map caps aggregate consumption across the namespace. Common keys: - Compute: `requests.cpu`, `requests.memory`, `limits.cpu`, `limits.memory` (the shorthands `cpu` and `memory` mean the request forms). - Storage: `requests.storage`, `persistentvolumeclaims`, and per-StorageClass variants. - Object counts: `pods`, `services`, `services.loadbalancers`, `secrets`, `configmaps`, `count/deployments.apps`. `status.used` tracks current consumption, so `kubectl describe quota` is the fastest way to see how close a namespace is. Scope selectors refine what counts: `Terminating`/`NotTerminating` split by whether `activeDeadlineSeconds` is set, `BestEffort`/`NotBestEffort` split by QoS, and `scopeSelector` with `PriorityClass` lets you reserve headroom for high-priority workloads or forbid low-tier teams from using a critical priority class at all. The critical rule: **if a quota constrains `requests.cpu` (or any compute resource), every new Pod in that namespace must specify that resource, or it is rejected.** Kubernetes cannot count what has not been declared. This is why quota and LimitRange are almost always deployed as a pair — the LimitRange guarantees no Pod ever arrives resource-less, so the quota never rejects an ordinary workload on a technicality. ## Enforcement semantics and failure modes Quota is checked at admission only. Consequences: - Lowering a quota below current usage does **not** evict anything. Existing Pods keep running; the namespace simply cannot create more until usage drops. `status.used` will legitimately exceed `spec.hard`. - Quota counts *declared* requests and limits, not actual usage. A namespace can be at 100% quota while consuming nothing. - Under quota pressure, controllers do not fail loudly at their own level. A Deployment scaling from 3 to 10 sits at 3 with `FailedCreate` events on the ReplicaSet. HorizontalPodAutoscaler scale-ups fail the same silent way. Alerting on quota utilisation is worth doing for exactly this reason. - Object-count quotas (`pods`, `secrets`) protect the control plane and etcd from runaway automation, which is a different risk from compute exhaustion and worth setting independently. ## Putting it together A typical namespace bootstrap ships three things: a LimitRange with modest `defaultRequest`/`default` and a `maxLimitRequestRatio` around 4, a ResourceQuota capping requests and limits plus a Pod count, and a default PriorityClass policy. That combination guarantees every workload is schedulable by default, no single container can hog a node, and one team cannot starve the cluster — without any application team having to think about it.
- A namespace has a ResourceQuota on requests.cpu and a Deployment suddenly stops creating Pods, but the Deployment itself shows no error. Where do you look?At the ReplicaSet, not the Deployment. Quota rejection happens when the ReplicaSet controller tries to create Pods, so the signal is `FailedCreate` events and a failure condition on the ReplicaSet, with the Deployment merely showing fewer ready replicas than desired. `kubectl describe quota` then shows which key is exhausted.
- If you reduce a ResourceQuota below what the namespace is already using, what happens to the running Pods?Nothing — quota is enforced only at admission time, so existing Pods keep running and `status.used` simply reports more than `spec.hard`. The namespace is frozen for new creations until usage falls under the new cap, which makes shrinking a quota a soft, gradual action rather than an eviction event.
saying these in an interview costs you the question
- Saying ResourceQuota caps a single container's usage — that is LimitRange's job.
- Believing a lowered quota evicts or kills running Pods.
- Thinking quota tracks measured usage rather than declared requests and limits.
- Forgetting that a compute quota makes resource declaration mandatory, so resource-less Pods are rejected outright.
- Expecting the quota rejection error to appear on the Deployment rather than on the ReplicaSet's Pod creation.