skip to content

You own ResourceQuota policy for tenant namespaces on a 16-node regulated-workload Kubernetes cluster whose etcd database has grown to 4.2 GB. How do you design the quotas, and what do they deliberately not protect?

level: principalimportance: should knowfreq 38%

answer

  1. bundle per namespace, platform-owned
  2. sum of quotas vs allocatable
  3. counts not bytes
  4. tiers via limitedResources
  5. list what quota misses

basics

~20 s

Give every tenant namespace a standard quota set: compute caps sized against allocatable, per-class storage, object counts that protect etcd, and scoped priority budgets. Quotas cap namespace totals only, not runtime usage, object size, API traffic or cluster-scoped objects.

solid answer

~40 s

I would provision a quota bundle with every tenant namespace, owned by the platform team and not editable by the tenant through RBAC. **Compute**: `requests.cpu`/`requests.memory` sized against the 16 nodes' allocatable, with a deliberate choice about whether the sum may exceed allocatable, plus `limits.*` caps that set the overcommit ceiling. **Storage**: per-StorageClass `requests.storage` and `persistentvolumeclaims`. **Objects**: `count/secrets`, `count/configmaps`, `count/jobs.batch` and counts for chatty custom resources, because the etcd database is already 4.2 GB. **Tiers**: `PriorityClass`-scoped quotas plus `limitedResources` so high-priority pods need an explicit grant. I would be explicit about the gaps: quota does not bound object size, API request rate, node-level contention or cluster-scoped objects, and a namespace without a quota is uncapped.

code

yaml · 16 lines
yaml
apiVersion: v1
kind: ResourceQuota
metadata:
  name: tenant-objects
  namespace: flags
spec:
  hard:
    count/secrets: "29"
    count/configmaps: "41"
    count/jobs.batch: "17"
    count/replicasets.apps: "57"
    services.loadbalancers: "1"
    services.nodeports: "0"
    gold.storageclass.storage.k8s.io/requests.storage: 120Gi
    gold.storageclass.storage.k8s.io/persistentvolumeclaims: "3"
    requests.nvidia.com/gpu: "2"

go deeper

for a junior

Recall that each namespace needs its own quota objects, and that a namespace without one has no caps at all.

for a middle

Explain which spec.hard keys protect compute, storage classes, network resources and etcd, and why the extended-resource form must start with requests.

for a senior

Show how you size quotas for peak rollout and autoscaling, and how you detect tenants nearing their caps before deploys fail.

for a principal

Own the tradeoff between summing quotas within allocatable and oversubscribing, and be explicit about what quota leaves to other controls: object size, API rate, runtime contention and isolation.

## Framing the decision A **ResourceQuota** caps one namespace's totals at admission time. On a shared cluster, a set of quotas is the platform team's **budget policy**: it decides how finite capacity and control-plane storage are split among tenants. No single design is right; the choices below are tradeoffs to make on purpose. The setting is a 16-node regulated-workload cluster whose etcd database has already reached **4.2 GB**, so control-plane storage is as scarce as CPU. ## The standard quota bundle Every tenant namespace gets the same set of objects when it is created, applied by whatever provisions namespaces (a GitOps controller or a platform API): | Quota | Keys | Purpose | |---|---|---| | compute | `requests.cpu`, `requests.memory`, `limits.cpu`, `limits.memory` | share of node allocatable and the overcommit ceiling | | storage | `<class>.storageclass.storage.k8s.io/requests.storage`, `persistentvolumeclaims` | stop one tenant filling the expensive tier | | objects | `pods`, `count/secrets`, `count/configmaps`, `count/jobs.batch`, `count/<plural>.<group>` | protect etcd and the controllers that watch these kinds | | network | `services.loadbalancers`, `services.nodeports` | scarce external addresses and node ports | | devices | `requests.nvidia.com/gpu` | ration accelerators (only the `requests.` form is counted) | | tiers | `pods`, `requests.cpu` scoped by `PriorityClass` | small budgets for high-priority work | ## Compute: fit the sum or oversubscribe it Quota is **not** a reservation. Deciding how the sum of all `requests.cpu` caps relates to the 16 nodes' allocatable CPU is the main call: 1. **Sum at or below allocatable.** Anything admitted within quota will almost always find room (fragmentation aside). This is predictable, which suits regulated workloads, but idle quota strands capacity. 2. **Sum above allocatable.** Utilization is higher, but a pod admitted within quota can still sit `Pending` when everyone bursts at once. You then need node autoscaling or a priority scheme to decide who waits. For a regulated cluster I would start near option 1 for production tiers and oversubscribe only a batch tier. The ratio between `limits.*` and `requests.*` caps sets how much burst a namespace can promise; per-container request and limit policy belongs to workload sizing, not to the quota. ## Object counts: protecting etcd A 4.2 GB etcd database means every extra object costs storage, watch traffic and compaction work. Count quotas cap the growth: - `count/secrets` and `count/configmaps` stop generators that create a new object on every deploy. - `count/jobs.batch` limits pile-ups from CronJobs that never clean up. - `count/replicasets.apps` bounds old ReplicaSets kept by generous revision history; leave headroom above what each Deployment's `revisionHistoryLimit` retains plus its live ReplicaSets, or rollouts fail to create their new ReplicaSet. - `count/<plural>.<group>` covers custom resources, often the fastest-growing kind. Be honest about the limit: **quota counts objects, not bytes**. Forty-one very large ConfigMaps pass a `count/configmaps: "41"` cap. Size limits and cleanup need other controls, such as a policy engine at admission. ## Priority tiers and making quota mandatory Priority values and preemption belong to the scheduler. Quota's job is to **ration** the tiers: - Create `PriorityClass`-scoped quotas so each namespace has a small high-tier budget and a larger default one. - Configure `limitedResources` with `matchScopes` in the API server's `ResourceQuotaConfiguration`, so a high-tier pod is rejected unless its namespace has a quota covering that class. Access to the tier becomes an explicit grant. ## Governance - **Ownership.** Tenants get read access to ResourceQuota objects but no write access; otherwise the cap is advisory. - **Namespace creation is the real gate.** A namespace without a quota is uncapped, so only the provisioning path may create namespaces, and a policy engine can reject namespaces that lack the bundle. - **Declaration side effect.** A compute quota rejects containers that omit CPU or memory values, so ship LimitRange defaults in the same bundle and announce quota changes like API changes. - **Review cadence.** Revisit each tenant's caps on a fixed schedule against measured `Used` values, and treat a raise as a capacity decision with an owner, not a ticket rubber-stamp. - **Observability.** Alert on `Used`/`Hard` ratios (kube-state-metrics exposes quota objects) before tenants hit a wall mid-rollout. ## What quota deliberately does not protect - **Runtime contention.** CPU throttling and OOM kills come from container limits on the node. - **API request rate.** Request floods are handled by API Priority and Fairness, not quota. - **Object size**, as noted above. - **Cluster-scoped objects** such as PersistentVolumes, ClusterRoles and CRDs themselves. - **Isolation.** Quota caps consumption; it is not a security boundary. RBAC, NetworkPolicy and Pod Security are what give a namespace real walls.

  • A tenant says a quota on count/configmaps is useless because their ConfigMaps are huge. How do you respond?
    They are right that quota counts objects, not bytes, so it cannot bound etcd growth from large objects. I would keep the count cap for runaway generators and add a size rule at admission with a policy engine, plus cleanup of unused generated objects. The etcd-size alert should also track which namespaces write most, so the conversation is based on data.
  • How would you size a tenant's compute quota for rollouts and autoscaling rather than steady state?
    Size it for the peak: maximum autoscaled replicas plus the rollout surge, times the per-pod request, for every workload in the namespace. Otherwise rollouts stall with exceeded-quota errors at the worst moment. Then check the sum of those peaks against allocatable, and oversubscribe only tiers where a Pending pod is acceptable.

Quota design is like issuing credit limits to departments: the limits can add up to more cash than the company holds, and whether you allow that is a policy choice, not an accounting fact.

saying these in an interview costs you the question

  • If every namespace stays within quota, every pod will schedule.
  • Object-count quotas cap how much space a tenant uses in etcd.
  • Tenants can own and tune their own ResourceQuota objects safely.
  • A ResourceQuota makes a namespace a security boundary.
  • Quota protects the API server from a tenant's request floods.
  • New namespaces are capped automatically without a provisioning step.