You own ResourceQuota policy for tenant namespaces on a 16-node regulated-workload Kubernetes cluster whose etcd database has grown to 4.2 GB. How do you design the quotas, and what do they deliberately not protect?
answer
- bundle per namespace, platform-owned
- sum of quotas vs allocatable
- counts not bytes
- tiers via limitedResources
- list what quota misses
basics
~20 sGive every tenant namespace a standard quota set: compute caps sized against allocatable, per-class storage, object counts that protect etcd, and scoped priority budgets. Quotas cap namespace totals only, not runtime usage, object size, API traffic or cluster-scoped objects.
solid answer
~40 sI would provision a quota bundle with every tenant namespace, owned by the platform team and not editable by the tenant through RBAC. **Compute**: `requests.cpu`/`requests.memory` sized against the 16 nodes' allocatable, with a deliberate choice about whether the sum may exceed allocatable, plus `limits.*` caps that set the overcommit ceiling. **Storage**: per-StorageClass `requests.storage` and `persistentvolumeclaims`. **Objects**: `count/secrets`, `count/configmaps`, `count/jobs.batch` and counts for chatty custom resources, because the etcd database is already 4.2 GB. **Tiers**: `PriorityClass`-scoped quotas plus `limitedResources` so high-priority pods need an explicit grant. I would be explicit about the gaps: quota does not bound object size, API request rate, node-level contention or cluster-scoped objects, and a namespace without a quota is uncapped.
code
yaml · 16 linesapiVersion: v1
kind: ResourceQuota
metadata:
name: tenant-objects
namespace: flags
spec:
hard:
count/secrets: "29"
count/configmaps: "41"
count/jobs.batch: "17"
count/replicasets.apps: "57"
services.loadbalancers: "1"
services.nodeports: "0"
gold.storageclass.storage.k8s.io/requests.storage: 120Gi
gold.storageclass.storage.k8s.io/persistentvolumeclaims: "3"
requests.nvidia.com/gpu: "2"go deeper
Recall that each namespace needs its own quota objects, and that a namespace without one has no caps at all.
Explain which spec.hard keys protect compute, storage classes, network resources and etcd, and why the extended-resource form must start with requests.
Show how you size quotas for peak rollout and autoscaling, and how you detect tenants nearing their caps before deploys fail.
Own the tradeoff between summing quotas within allocatable and oversubscribing, and be explicit about what quota leaves to other controls: object size, API rate, runtime contention and isolation.
## Framing the decision A **ResourceQuota** caps one namespace's totals at admission time. On a shared cluster, a set of quotas is the platform team's **budget policy**: it decides how finite capacity and control-plane storage are split among tenants. No single design is right; the choices below are tradeoffs to make on purpose. The setting is a 16-node regulated-workload cluster whose etcd database has already reached **4.2 GB**, so control-plane storage is as scarce as CPU. ## The standard quota bundle Every tenant namespace gets the same set of objects when it is created, applied by whatever provisions namespaces (a GitOps controller or a platform API): | Quota | Keys | Purpose | |---|---|---| | compute | `requests.cpu`, `requests.memory`, `limits.cpu`, `limits.memory` | share of node allocatable and the overcommit ceiling | | storage | `<class>.storageclass.storage.k8s.io/requests.storage`, `persistentvolumeclaims` | stop one tenant filling the expensive tier | | objects | `pods`, `count/secrets`, `count/configmaps`, `count/jobs.batch`, `count/<plural>.<group>` | protect etcd and the controllers that watch these kinds | | network | `services.loadbalancers`, `services.nodeports` | scarce external addresses and node ports | | devices | `requests.nvidia.com/gpu` | ration accelerators (only the `requests.` form is counted) | | tiers | `pods`, `requests.cpu` scoped by `PriorityClass` | small budgets for high-priority work | ## Compute: fit the sum or oversubscribe it Quota is **not** a reservation. Deciding how the sum of all `requests.cpu` caps relates to the 16 nodes' allocatable CPU is the main call: 1. **Sum at or below allocatable.** Anything admitted within quota will almost always find room (fragmentation aside). This is predictable, which suits regulated workloads, but idle quota strands capacity. 2. **Sum above allocatable.** Utilization is higher, but a pod admitted within quota can still sit `Pending` when everyone bursts at once. You then need node autoscaling or a priority scheme to decide who waits. For a regulated cluster I would start near option 1 for production tiers and oversubscribe only a batch tier. The ratio between `limits.*` and `requests.*` caps sets how much burst a namespace can promise; per-container request and limit policy belongs to workload sizing, not to the quota. ## Object counts: protecting etcd A 4.2 GB etcd database means every extra object costs storage, watch traffic and compaction work. Count quotas cap the growth: - `count/secrets` and `count/configmaps` stop generators that create a new object on every deploy. - `count/jobs.batch` limits pile-ups from CronJobs that never clean up. - `count/replicasets.apps` bounds old ReplicaSets kept by generous revision history; leave headroom above what each Deployment's `revisionHistoryLimit` retains plus its live ReplicaSets, or rollouts fail to create their new ReplicaSet. - `count/<plural>.<group>` covers custom resources, often the fastest-growing kind. Be honest about the limit: **quota counts objects, not bytes**. Forty-one very large ConfigMaps pass a `count/configmaps: "41"` cap. Size limits and cleanup need other controls, such as a policy engine at admission. ## Priority tiers and making quota mandatory Priority values and preemption belong to the scheduler. Quota's job is to **ration** the tiers: - Create `PriorityClass`-scoped quotas so each namespace has a small high-tier budget and a larger default one. - Configure `limitedResources` with `matchScopes` in the API server's `ResourceQuotaConfiguration`, so a high-tier pod is rejected unless its namespace has a quota covering that class. Access to the tier becomes an explicit grant. ## Governance - **Ownership.** Tenants get read access to ResourceQuota objects but no write access; otherwise the cap is advisory. - **Namespace creation is the real gate.** A namespace without a quota is uncapped, so only the provisioning path may create namespaces, and a policy engine can reject namespaces that lack the bundle. - **Declaration side effect.** A compute quota rejects containers that omit CPU or memory values, so ship LimitRange defaults in the same bundle and announce quota changes like API changes. - **Review cadence.** Revisit each tenant's caps on a fixed schedule against measured `Used` values, and treat a raise as a capacity decision with an owner, not a ticket rubber-stamp. - **Observability.** Alert on `Used`/`Hard` ratios (kube-state-metrics exposes quota objects) before tenants hit a wall mid-rollout. ## What quota deliberately does not protect - **Runtime contention.** CPU throttling and OOM kills come from container limits on the node. - **API request rate.** Request floods are handled by API Priority and Fairness, not quota. - **Object size**, as noted above. - **Cluster-scoped objects** such as PersistentVolumes, ClusterRoles and CRDs themselves. - **Isolation.** Quota caps consumption; it is not a security boundary. RBAC, NetworkPolicy and Pod Security are what give a namespace real walls.
- A tenant says a quota on count/configmaps is useless because their ConfigMaps are huge. How do you respond?They are right that quota counts objects, not bytes, so it cannot bound etcd growth from large objects. I would keep the count cap for runaway generators and add a size rule at admission with a policy engine, plus cleanup of unused generated objects. The etcd-size alert should also track which namespaces write most, so the conversation is based on data.
- How would you size a tenant's compute quota for rollouts and autoscaling rather than steady state?Size it for the peak: maximum autoscaled replicas plus the rollout surge, times the per-pod request, for every workload in the namespace. Otherwise rollouts stall with exceeded-quota errors at the worst moment. Then check the sum of those peaks against allocatable, and oversubscribe only tiers where a Pending pod is acceptable.
Quota design is like issuing credit limits to departments: the limits can add up to more cash than the company holds, and whether you allow that is a policy choice, not an accounting fact.
saying these in an interview costs you the question
- If every namespace stays within quota, every pod will schedule.
- Object-count quotas cap how much space a tenant uses in etcd.
- Tenants can own and tune their own ResourceQuota objects safely.
- A ResourceQuota makes a namespace a security boundary.
- Quota protects the API server from a tenant's request floods.
- New namespaces are capped automatically without a provisioning step.