skip to content

ResourceQuota

A ResourceQuota caps what one namespace may consume: CPU and memory requests and limits, storage, and object counts. Once it names a compute resource, every new pod must declare that resource or be rejected - which interviewers ask as a broken-deploys story.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

4

What is a Kubernetes ResourceQuota, and which kinds of consumption can its spec.hard field cap for one namespace?

level: juniorimportance: must knowfreq 64%

answer

  1. namespace-wide ceilings
  2. compute, storage, counts, extended
  3. count/<resource>.<group>
  4. checked at admission, 403
  5. status.used vs spec.hard

basics

~20 s

A ResourceQuota is a namespaced object whose spec.hard caps the namespace's total CPU and memory requests and limits, storage requests and object counts. The API server rejects any create that would push usage past a cap.

solid answer

~30 s

A `ResourceQuota` is a namespaced object that sets aggregate ceilings for everything in that namespace. Its `spec.hard` map takes four families of keys: compute (`requests.cpu`, `requests.memory`, `limits.cpu`, `limits.memory`), storage (`requests.storage`, `persistentvolumeclaims`, and per-StorageClass variants), object counts (`pods`, `services.loadbalancers`, or the generic `count/<resource>.<group>` such as `count/jobs.batch`), and extended resources in the `requests.` form. The ResourceQuota controller keeps `status.used` current, and the `ResourceQuota` admission plugin in kube-apiserver rejects a create with 403 Forbidden if it would exceed any cap. It is an admission-time cap: it never evicts running pods, never throttles CPU and never reserves capacity on nodes.

code

yaml · 16 lines
yaml
apiVersion: v1
kind: ResourceQuota
metadata:
  name: flags-compute
  namespace: flags
spec:
  hard:
    requests.cpu: 6800m
    requests.memory: 13Gi
    limits.memory: 19Gi
    pods: "23"
    persistentvolumeclaims: "7"
    requests.storage: 340Gi
    count/configmaps: "41"
    count/jobs.batch: "17"
    services.loadbalancers: "1"

go deeper

for a junior

Recall that a ResourceQuota is per namespace and name the four kinds of caps: compute requests and limits, storage, object counts and extended resources. Know that an over-quota create fails with 403.

for a middle

Explain the split between the controller that computes status.used and the admission plugin that checks each create, and why a new cpu or memory cap forces every container to declare that value.

for a senior

Show you treat quota as an admission cap, not capacity: remaining quota does not mean a pod will fit, and tightening a quota never removes running workloads.

for a principal

Frame quota as a governance tool: which families of keys each tenant gets, who may edit quota objects, and how object counts protect shared control-plane storage.

## What a ResourceQuota is A **namespace** is a name scope that most Kubernetes objects live in. A **ResourceQuota** is a namespaced object (`apiVersion: v1`, `kind: ResourceQuota`) that caps the **aggregate** consumption of one namespace. It exists so that one team sharing a cluster cannot create enough pods, volumes or objects to crowd out everyone else. It has two halves: - `spec.hard`: the caps you declare, as a map from a **resource name** to a quantity. - `status.hard` and `status.used`: what the control plane is enforcing and what it has counted so far. A namespace can hold several ResourceQuota objects. A new object must fit under **every** quota that tracks it. ## The four families of spec.hard keys | Family | Example keys | What is summed | |---|---|---| | Compute | `requests.cpu`, `requests.memory`, `limits.cpu`, `limits.memory` (bare `cpu`/`memory` mean the `requests.` form) | the requests or limits declared by non-terminal pods | | Storage | `requests.storage`, `persistentvolumeclaims`, `<class>.storageclass.storage.k8s.io/requests.storage` | capacity requested by PersistentVolumeClaims, overall or per StorageClass | | Object counts | `pods`, `services`, `services.loadbalancers`, `services.nodeports`, `count/<resource>.<group>` | how many objects of a kind exist | | Extended resources | `requests.nvidia.com/gpu` | device requests; only the `requests.` form is counted | The generic object-count syntax drops the group for core resources and includes it otherwise: `count/configmaps`, `count/secrets`, `count/deployments.apps`, `count/jobs.batch`. It also works for custom resources, using the plural name and the API group. Two accounting details matter in practice: - Pods in phase `Succeeded` or `Failed` are **not** charged, so finished Job pods do not eat compute quota. - A pod that is stuck deleting past its grace period stops being charged, so a lost node does not block a replacement. ## How enforcement works 1. You create the quota, for example with `kubectl apply` or `kubectl create quota`. 2. The **ResourceQuota controller** in kube-controller-manager lists the namespace, computes usage and writes `status.hard` and `status.used`. It also does a full recalculation on a resync period (five minutes by default, set with `--resource-quota-sync-period`). 3. On every create (and on updates that change usage, such as growing a PVC), the **`ResourceQuota` admission plugin** in kube-apiserver adds the new object's usage to `status.used` and compares it with each cap. 4. If any cap would be exceeded, the request fails with **403 Forbidden** and a message such as `exceeded quota: flags-compute, requested: requests.cpu=700m, used: requests.cpu=6300m, limited: requests.cpu=6800m`. Otherwise the object is saved and `status.used` is bumped. One side effect catches people out: once a quota caps `cpu` or `memory` in any form, **each container in a new pod must declare that value** or the pod is rejected. Defaults for containers that declare nothing come from a LimitRange, which is a separate object. ## What a ResourceQuota does not do - **It is not retroactive.** Creating `pods: "23"` in a namespace that already runs 31 pods evicts nothing; it only blocks new pods until usage falls below the cap. - **It is not a reservation.** Having quota left does not mean a node has room; a pod can pass admission and still sit `Pending`. - **It is not a runtime limit.** CPU throttling and memory kills come from container `resources.limits`, enforced on the node, not from the quota. - **It is not cluster-wide.** It only counts objects in its own namespace, and cluster-scoped objects are never counted. ## How it fits with the other namespace controls ResourceQuota is one of several objects that shape a namespace, and interviewers like to hear them kept apart: | Control | Scope | Acts when | |---|---|---| | ResourceQuota | the namespace's **sum** of usage and objects | a create or usage-changing update is admitted | | LimitRange | bounds and defaults for **each** pod, container or claim | a single object is admitted | | container `resources.limits` | one running container | on the node, at runtime | A quota answers "how much may this team have in total?". It says nothing about how big one pod may be. That is why platform teams usually ship a quota and a LimitRange together: the quota caps the total, and the LimitRange keeps single objects sensible and fills in values that containers leave out. ## Reading a quota `kubectl describe resourcequota <name> -n <ns>` prints each resource with its `Used` and `Hard` values side by side. That view is the first thing to check when creates start failing, for example when a feature-flag evaluation service's rollout suddenly stops creating pods. `kubectl get resourcequota -n <ns>` gives the same comparison in a compact table.

  • Does a new ResourceQuota take effect the instant it is created?
    Almost. The ResourceQuota controller must first fill in `status.hard` and `status.used`. Until it does, creates that touch the quota's resources are rejected with a `status unknown for quota` error, which normally clears within seconds. After that, admission updates `status.used` on each admitted create, and the controller recalculates on its resync period, five minutes by default.
  • Can a ResourceQuota cap GPUs or other device-plugin resources?
    Yes, using the `requests.` form, for example `requests.nvidia.com/gpu: "3"`. Extended resources cannot be overcommitted, so quota only counts requests for them; a `limits.` key for an extended resource is not counted. Unlike CPU and memory, a GPU quota does not force every container in the namespace to declare a GPU value.

A ResourceQuota is like a department's annual budget: purchases that would overspend it are refused at the till, but nothing already bought is taken back.

saying these in an interview costs you the question

  • A ResourceQuota evicts pods when a namespace is already over the new cap.
  • A ResourceQuota throttles CPU for a namespace that uses too much.
  • Remaining quota guarantees that a pod will be scheduled onto a node.
  • ResourceQuota applies to the whole cluster, not to one namespace.
  • Object-count quota only works for pods and services, not other kinds.
  • Completed Job pods keep consuming the namespace's compute quota.
open as a page

After a ResourceQuota is added to its Kubernetes namespace, a feature-flag evaluation service's Deployment rollout stalls with no new pods, although kubectl apply succeeded. How do you diagnose and fix it?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Quota rejects pods when the ReplicaSet controller creates them, so the error appears as FailedCreate events on the ReplicaSet, not at apply time. It is usually 'must specify' (containers lack requests or limits) or 'exceeded quota' (no headroom for surge pods).

open as a page

You own ResourceQuota policy for tenant namespaces on a 16-node regulated-workload Kubernetes cluster whose etcd database has grown to 4.2 GB. How do you design the quotas, and what do they deliberately not protect?

level: principalimportance: should knowfreq 38%

basics

~20 s

Give every tenant namespace a standard quota set: compute caps sized against allocatable, per-class storage, object counts that protect etcd, and scoped priority budgets. Quotas cap namespace totals only, not runtime usage, object size, API traffic or cluster-scoped objects.

open as a page

How do a Kubernetes ResourceQuota's scopes and scopeSelector fields narrow which pods it counts, and what do the Terminating and BestEffort scopes match?

level: middleimportance: nice to knowfreq 29%

basics

~20 s

Scopes make a ResourceQuota count only matching pods: Terminating means spec.activeDeadlineSeconds is set, BestEffort means no CPU or memory requests or limits, and PriorityClass matches priorityClassName. A pod must match every listed scope and selector.

open as a page