In a Kubernetes Pod specification, what is the difference between a container's resource requests and its resource limits, and what does each one actually control?
answer
- requests = scheduler; limits = kernel
- CPU compressible → throttle
- memory incompressible → OOMKill 137
- request also = cpu.shares
- overcommit = limits > requests
basics
~20 sA request is what the scheduler reserves to place the Pod on a node. A limit is the hard ceiling the node enforces at runtime: exceed a CPU limit and the container is throttled, exceed a memory limit and it is OOMKilled.
solid answer
~50 sThey serve two different consumers. **Requests** are a scheduling contract: kube-scheduler sums a Pod's container requests and only places it on a node whose allocatable capacity minus already-committed requests still fits. Requests also become CPU shares, so under contention CPU is split roughly in proportion to requests. **Limits** are a runtime contract enforced by the kernel through cgroups: a CPU limit becomes CFS bandwidth quota, so the container is throttled once it burns its quota within a 100 ms period; a memory limit becomes the memory cgroup maximum, and crossing it triggers the cgroup OOM killer, killing the container with exit code 137 and reason OOMKilled. The scheduler never looks at limits and the kernel never treats requests as a ceiling. That gap is what allows overcommit: nodes are packed by requests but may actually consume up to the sum of limits.
code
yaml · 15 linesapiVersion: v1
kind: Pod
metadata:
name: api
spec:
containers:
- name: api
image: example/api:1.4
resources:
requests:
cpu: "250m"
memory: "512Mi"
limits:
cpu: "1"
memory: "512Mi"go deeper
Be able to state the one-line split — requests are for scheduling, limits are the runtime cap — and name the two failure modes: CPU throttling and memory OOMKill.
Add the mechanics: requests also become CPU shares, limits become CFS quota and the memory cgroup max, and the request/limit gap is what creates overcommit.
Tie it to operations — how you pick the numbers from observed usage, why memory request is usually set equal to limit, and how you spot throttling and OOMKills in metrics.
Frame it as cluster economics: request accuracy drives bin packing and cost, the limits policy drives blast radius and tail latency, and defaults must be enforced platform-wide rather than left to each team.
## The two fields Each container in a Pod may declare `resources.requests` and `resources.limits`, holding CPU, memory, ephemeral storage and optionally extended resources such as GPUs. Both are optional and independent. CPU is expressed in cores: `1` is one core, `500m` is half a core. Memory is bytes with suffixes, where `Mi`/`Gi` are powers of 1024 and `M`/`G` are powers of 1000. ## What requests do Requests drive **scheduling**. kube-scheduler computes a Pod's effective request and filters out nodes where that would exceed allocatable capacity minus the requests of Pods already assigned there. This is pure bookkeeping against *declared* numbers — the scheduler does not measure real usage. A node can therefore be 100% "full" by requests while idling, or heavily loaded while showing spare requestable room. On the node, the CPU request is also translated into the cgroup's CPU weight (`cpu.shares` under cgroup v1, `cpu.weight` under v2). Shares only matter when the CPU is saturated: two containers requesting 500m and 1500m that both want more CPU than exists will get roughly a 1:3 split. When the CPU is not saturated, either can use as much as it wants up to its limit. The memory request is not enforced at all at runtime. Its second job is eviction: when a node runs low on memory, the kubelet ranks Pods partly by how far their usage exceeds their memory request, and evicts the worst offenders first. ## What limits do Limits are enforced by the kernel. The CPU limit becomes CFS bandwidth control — `cpu.cfs_quota_us`/`cpu.cfs_period_us` on cgroup v1, `cpu.max` on v2. With a default 100 ms period, a limit of `500m` grants 50 ms of CPU time per 100 ms window across all threads in the container. Once spent, every runnable thread is stopped until the next period. This is *throttling*: the container is not killed, it just stalls, which shows up as latency spikes rather than errors. The memory limit becomes the memory cgroup maximum (`memory.limit_in_bytes` / `memory.max`). Memory cannot be compressed — you cannot give a process "less memory, slower" — so when an allocation would exceed the limit the kernel's OOM killer terminates a process inside that cgroup. The container exits with 137 (128 + SIGKILL), `kubectl describe pod` shows `Last State: Terminated, Reason: OOMKilled`, and the Pod's restart policy decides whether it comes back, often into CrashLoopBackOff. ## Why the distinction matters Because the scheduler uses requests and the kernel uses limits, the request-to-limit ratio is a deliberate policy knob. Requests equal to limits means no overcommit, predictable behaviour, and the highest QoS class. Requests much lower than limits packs more Pods per node and lets bursty workloads use idle capacity, but the node can then be oversubscribed and start throttling or evicting. Omitting values has consequences too. No request means the scheduler treats the container as free and will happily stack it onto a full node. No limit means the container can consume everything the node has, starving neighbours — which is exactly what LimitRange defaults exist to prevent. Note also the defaulting rule: if you set a limit but no request, Kubernetes sets the request equal to the limit. ## Practical guidance Set a memory request that reflects steady-state working set, and a memory limit at or near it so failure is a clean, attributable OOMKill rather than a node-wide memory crisis. Set a CPU request from observed usage, and treat the CPU limit as a deliberate choice: it caps blast radius but also caps burst.
- What happens if you set a limit but no request?Kubernetes defaults the request to equal the limit for that resource. So a container with only `limits.memory: 1Gi` is scheduled as if it requested 1Gi, which is conservative rather than dangerous. The reverse — a request with no limit — leaves the container uncapped at runtime.
- Can a container use more CPU than its request?Yes. The request is only a scheduling reservation and a share weight, not a ceiling. If the node has idle CPU the container can burst up to its limit, or up to the whole node if no limit is set. It is only pulled back toward its share when the CPU is genuinely contended.
A request is the seat you reserved on the train — it guarantees you get on. The limit is how far you may spread out once aboard: go past it on the armrest and you get pushed back, go past it on the emergency exit and you get thrown off.
saying these in an interview costs you the question
- Saying the scheduler uses limits to decide placement — it only ever uses requests.
- Claiming a container is killed when it exceeds its CPU limit; CPU is throttled, not killed.
- Thinking the memory request is enforced at runtime — nothing stops a container exceeding it until the limit or node pressure.
- Assuming a node reporting 100% requested CPU is actually busy, or vice versa.
- Treating requests and limits as interchangeable synonyms for 'how much resource the container gets'.