skip to content

What is a Kubernetes ResourceQuota, and which kinds of consumption can its spec.hard field cap for one namespace?

level: juniorimportance: must knowfreq 64%

answer

  1. namespace-wide ceilings
  2. compute, storage, counts, extended
  3. count/<resource>.<group>
  4. checked at admission, 403
  5. status.used vs spec.hard

basics

~20 s

A ResourceQuota is a namespaced object whose spec.hard caps the namespace's total CPU and memory requests and limits, storage requests and object counts. The API server rejects any create that would push usage past a cap.

solid answer

~30 s

A `ResourceQuota` is a namespaced object that sets aggregate ceilings for everything in that namespace. Its `spec.hard` map takes four families of keys: compute (`requests.cpu`, `requests.memory`, `limits.cpu`, `limits.memory`), storage (`requests.storage`, `persistentvolumeclaims`, and per-StorageClass variants), object counts (`pods`, `services.loadbalancers`, or the generic `count/<resource>.<group>` such as `count/jobs.batch`), and extended resources in the `requests.` form. The ResourceQuota controller keeps `status.used` current, and the `ResourceQuota` admission plugin in kube-apiserver rejects a create with 403 Forbidden if it would exceed any cap. It is an admission-time cap: it never evicts running pods, never throttles CPU and never reserves capacity on nodes.

code

yaml · 16 lines
yaml
apiVersion: v1
kind: ResourceQuota
metadata:
  name: flags-compute
  namespace: flags
spec:
  hard:
    requests.cpu: 6800m
    requests.memory: 13Gi
    limits.memory: 19Gi
    pods: "23"
    persistentvolumeclaims: "7"
    requests.storage: 340Gi
    count/configmaps: "41"
    count/jobs.batch: "17"
    services.loadbalancers: "1"

go deeper

for a junior

Recall that a ResourceQuota is per namespace and name the four kinds of caps: compute requests and limits, storage, object counts and extended resources. Know that an over-quota create fails with 403.

for a middle

Explain the split between the controller that computes status.used and the admission plugin that checks each create, and why a new cpu or memory cap forces every container to declare that value.

for a senior

Show you treat quota as an admission cap, not capacity: remaining quota does not mean a pod will fit, and tightening a quota never removes running workloads.

for a principal

Frame quota as a governance tool: which families of keys each tenant gets, who may edit quota objects, and how object counts protect shared control-plane storage.

## What a ResourceQuota is A **namespace** is a name scope that most Kubernetes objects live in. A **ResourceQuota** is a namespaced object (`apiVersion: v1`, `kind: ResourceQuota`) that caps the **aggregate** consumption of one namespace. It exists so that one team sharing a cluster cannot create enough pods, volumes or objects to crowd out everyone else. It has two halves: - `spec.hard`: the caps you declare, as a map from a **resource name** to a quantity. - `status.hard` and `status.used`: what the control plane is enforcing and what it has counted so far. A namespace can hold several ResourceQuota objects. A new object must fit under **every** quota that tracks it. ## The four families of spec.hard keys | Family | Example keys | What is summed | |---|---|---| | Compute | `requests.cpu`, `requests.memory`, `limits.cpu`, `limits.memory` (bare `cpu`/`memory` mean the `requests.` form) | the requests or limits declared by non-terminal pods | | Storage | `requests.storage`, `persistentvolumeclaims`, `<class>.storageclass.storage.k8s.io/requests.storage` | capacity requested by PersistentVolumeClaims, overall or per StorageClass | | Object counts | `pods`, `services`, `services.loadbalancers`, `services.nodeports`, `count/<resource>.<group>` | how many objects of a kind exist | | Extended resources | `requests.nvidia.com/gpu` | device requests; only the `requests.` form is counted | The generic object-count syntax drops the group for core resources and includes it otherwise: `count/configmaps`, `count/secrets`, `count/deployments.apps`, `count/jobs.batch`. It also works for custom resources, using the plural name and the API group. Two accounting details matter in practice: - Pods in phase `Succeeded` or `Failed` are **not** charged, so finished Job pods do not eat compute quota. - A pod that is stuck deleting past its grace period stops being charged, so a lost node does not block a replacement. ## How enforcement works 1. You create the quota, for example with `kubectl apply` or `kubectl create quota`. 2. The **ResourceQuota controller** in kube-controller-manager lists the namespace, computes usage and writes `status.hard` and `status.used`. It also does a full recalculation on a resync period (five minutes by default, set with `--resource-quota-sync-period`). 3. On every create (and on updates that change usage, such as growing a PVC), the **`ResourceQuota` admission plugin** in kube-apiserver adds the new object's usage to `status.used` and compares it with each cap. 4. If any cap would be exceeded, the request fails with **403 Forbidden** and a message such as `exceeded quota: flags-compute, requested: requests.cpu=700m, used: requests.cpu=6300m, limited: requests.cpu=6800m`. Otherwise the object is saved and `status.used` is bumped. One side effect catches people out: once a quota caps `cpu` or `memory` in any form, **each container in a new pod must declare that value** or the pod is rejected. Defaults for containers that declare nothing come from a LimitRange, which is a separate object. ## What a ResourceQuota does not do - **It is not retroactive.** Creating `pods: "23"` in a namespace that already runs 31 pods evicts nothing; it only blocks new pods until usage falls below the cap. - **It is not a reservation.** Having quota left does not mean a node has room; a pod can pass admission and still sit `Pending`. - **It is not a runtime limit.** CPU throttling and memory kills come from container `resources.limits`, enforced on the node, not from the quota. - **It is not cluster-wide.** It only counts objects in its own namespace, and cluster-scoped objects are never counted. ## How it fits with the other namespace controls ResourceQuota is one of several objects that shape a namespace, and interviewers like to hear them kept apart: | Control | Scope | Acts when | |---|---|---| | ResourceQuota | the namespace's **sum** of usage and objects | a create or usage-changing update is admitted | | LimitRange | bounds and defaults for **each** pod, container or claim | a single object is admitted | | container `resources.limits` | one running container | on the node, at runtime | A quota answers "how much may this team have in total?". It says nothing about how big one pod may be. That is why platform teams usually ship a quota and a LimitRange together: the quota caps the total, and the LimitRange keeps single objects sensible and fills in values that containers leave out. ## Reading a quota `kubectl describe resourcequota <name> -n <ns>` prints each resource with its `Used` and `Hard` values side by side. That view is the first thing to check when creates start failing, for example when a feature-flag evaluation service's rollout suddenly stops creating pods. `kubectl get resourcequota -n <ns>` gives the same comparison in a compact table.

  • Does a new ResourceQuota take effect the instant it is created?
    Almost. The ResourceQuota controller must first fill in `status.hard` and `status.used`. Until it does, creates that touch the quota's resources are rejected with a `status unknown for quota` error, which normally clears within seconds. After that, admission updates `status.used` on each admitted create, and the controller recalculates on its resync period, five minutes by default.
  • Can a ResourceQuota cap GPUs or other device-plugin resources?
    Yes, using the `requests.` form, for example `requests.nvidia.com/gpu: "3"`. Extended resources cannot be overcommitted, so quota only counts requests for them; a `limits.` key for an extended resource is not counted. Unlike CPU and memory, a GPU quota does not force every container in the namespace to declare a GPU value.

A ResourceQuota is like a department's annual budget: purchases that would overspend it are refused at the till, but nothing already bought is taken back.

saying these in an interview costs you the question

  • A ResourceQuota evicts pods when a namespace is already over the new cap.
  • A ResourceQuota throttles CPU for a namespace that uses too much.
  • Remaining quota guarantees that a pod will be scheduled onto a node.
  • ResourceQuota applies to the whole cluster, not to one namespace.
  • Object-count quota only works for pods and services, not other kinds.
  • Completed Job pods keep consuming the namespace's compute quota.