A container exceeds its CPU limit, and another container exceeds its memory limit. Describe what the Linux kernel does in each case, and how you would detect each from outside the container.
answer
- CPU compressible → CFS quota, 100 ms period
- nr_throttled / container_cpu_cfs_throttled_seconds_total
- averages hide burst throttling
- 137 = 128 + SIGKILL, reason OOMKilled
- OOMKilled vs Evicted vs node OOM
basics
~20 sCPU is compressible: the kernel's CFS bandwidth controller stalls the container's threads until the next 100 ms period, visible as throttled-seconds metrics and latency spikes. Memory is not: the cgroup OOM killer terminates the process, exit code 137, reason OOMKilled.
solid answer
~50 s**CPU limit.** The limit becomes CFS quota (`cpu.max` on cgroup v2). Every 100 ms period the container gets `limit × 100 ms` of CPU time across all its threads; once spent, runnable threads are descheduled until the period rolls over. Nothing crashes — you see latency spikes and rising `container_cpu_cfs_throttled_seconds_total` and `nr_throttled` in the cgroup's `cpu.stat`. Average CPU usage can look comfortably below the limit while a bursty, highly-threaded process is throttled hard inside individual periods. **Memory limit.** Memory cannot be handed out more slowly, so exceeding `memory.max` triggers the cgroup OOM killer, which SIGKILLs a process in that cgroup. The container exits 137; `kubectl describe pod` shows `Last State: Terminated, Reason: OOMKilled`, the restart count increments, and repeated kills produce CrashLoopBackOff. The node's `dmesg` records the kill. So the diagnostic signature differs: throttling is a *performance* symptom you must go looking for; OOMKill is a *liveness* symptom that announces itself.
code
bash · 6 lineskubectl get pod api -o jsonpath='{.status.containerStatuses[0].lastState.terminated.reason}'
kubectl get pod api -o jsonpath='{.status.containerStatuses[0].lastState.terminated.exitCode}'
kubectl exec api -- cat /sys/fs/cgroup/cpu.stat
kubectl debug node/ip-10-0-1-7 -it --image=busybox -- dmesg | grep -i 'killed process'go deeper
Know the two outcomes by name: CPU over limit means throttling, memory over limit means the container is killed and restarted.
Explain the cgroup mechanics — CFS quota over a 100 ms period versus memory.max and the OOM killer — and which kubectl output shows each.
Own the diagnosis: throttled-period ratios, working-set versus limit graphs, distinguishing cgroup OOM from node OOM from eviction, and the remediation for each.
Turn it into policy — whether CPU limits are set at all for latency-tier services, how thread pools are sized against limits, and what alerting on throttling and OOMKills the platform ships by default.
## Compressible versus incompressible Kubernetes treats CPU as *compressible* and memory as *incompressible*. You can give a process less CPU by making it wait — it runs slower but stays correct. You cannot give a process less memory; an allocation either succeeds or the process must die. That single distinction explains every behavioural difference below. ## What happens at the CPU limit The kubelet writes the CPU limit into the container's cgroup as CFS bandwidth control: `cpu.cfs_quota_us` and `cpu.cfs_period_us` on cgroup v1, or the single `cpu.max` file on v2. The period defaults to 100 ms. A limit of `500m` therefore means 50 ms of CPU time per 100 ms window, summed across every thread in the container. When the quota is exhausted the kernel throttles: all runnable threads in that cgroup are removed from the run queue until the next period begins. Nothing is signalled, no error is returned, no log line appears in the application. The process simply loses time. The pathology that surprises people is that **throttling is bursty and averages hide it**. A JVM or Go service with sixteen worker threads and a `1` CPU limit exhausts its 100 ms quota in about 6 ms of wall clock if all sixteen threads run at once — then stalls for 94 ms. Averaged over a minute, CPU usage might read 30% of the limit while p99 latency is dominated by throttle stalls. Garbage collectors, thread pools that fan out, and startup paths are classic triggers. **Detecting it:** the cgroup exposes `cpu.stat` with `nr_periods`, `nr_throttled` and `throttled_time`/`throttled_usec`. cAdvisor surfaces these as `container_cpu_cfs_periods_total` and `container_cpu_cfs_throttled_seconds_total`. The useful signal is the ratio of throttled periods to total periods per container; anything consistently non-trivial on a latency-sensitive service warrants raising the limit, reducing thread-pool sizes to match the limit, or removing the CPU limit and relying on requests for fairness. Note that older kernels (before roughly 4.18/5.4, depending on backports) had CFS accounting bugs that throttled workloads well below their quota, so on very old nodes kernel version is part of the diagnosis. ## What happens at the memory limit The memory limit lands in the memory cgroup as `memory.limit_in_bytes` (v1) or `memory.max` (v2). When a charge would push the cgroup over its limit, the kernel first tries to reclaim — dropping page cache, writing back dirty pages, swapping if swap is enabled and permitted. If reclaim cannot free enough, the cgroup OOM killer picks a victim process inside that cgroup by `oom_score` and SIGKILLs it. From Kubernetes' side: the container's exit code is 137 (128 + 9 for SIGKILL), and the container status shows `reason: OOMKilled`. Crucially, if the killed process was not PID 1, the container may keep running in a degraded state — which is why some workloads look "fine" while quietly losing a worker. The restart policy then applies, and repeated kills back off into CrashLoopBackOff. Distinguish three OOM flavours during diagnosis: 1. **Container cgroup OOM** — the Pod exceeded its own limit. Reason `OOMKilled`; blame the workload or the limit. 2. **Node-level OOM** — the node ran out of memory as a whole and the kernel chose a victim by `oom_score_adj`, which the kubelet set from QoS. A well-behaved Pod can be killed here through no fault of its own. 3. **kubelet eviction** — the kubelet noticed `memory.available` crossing its threshold and gracefully evicted Pods first. The Pod status shows `Evicted` with a reason message, not `OOMKilled`. **Detecting it:** `kubectl describe pod` and the container's `lastState.terminated` for the terminated state; `kube_pod_container_status_last_terminated_reason` in kube-state-metrics for alerting; `container_memory_working_set_bytes` compared against `container_spec_memory_limit_bytes` for the run-up; and `dmesg` or the node journal for the kernel's own record naming the killed process and cgroup. ## Practical rule of thumb If latency is bad but nothing restarts, suspect CPU throttling. If containers restart with 137, it is memory — and the next question is whether the working set genuinely grew (leak, cache, larger request payloads) or the limit was simply set below what the runtime needs, including JVM non-heap or Go runtime overhead beyond the heap.
- A service shows average CPU well under its limit but a high throttled-periods ratio. How is that possible?CFS accounts quota in 100 ms windows, not in minutes. A multi-threaded process can burn an entire window's quota in a few milliseconds when several threads run concurrently, then sit stalled for the rest of the window. Averaged over a scrape interval the usage looks low, yet request latency is dominated by those stalls.
- A container is OOMKilled but the Pod keeps running with no restart. What does that tell you?The cgroup OOM killer chose a child process rather than PID 1, so the container's main process survived and Kubernetes never saw a container exit. The workload is silently degraded — a lost worker or forked helper. You find it in the node's dmesg and in memory metrics dropping abruptly, not in the restart count.
saying these in an interview costs you the question
- Saying a container is killed when it exceeds its CPU limit.
- Saying memory over-usage causes throttling or swapping that keeps the container alive.
- Diagnosing throttling from average CPU utilisation instead of the throttled-periods ratio.
- Confusing an Evicted Pod (kubelet, graceful, node pressure) with OOMKilled (kernel, SIGKILL, own limit).
- Assuming exit code 137 always means the memory limit was hit — any SIGKILL, including a liveness-probe restart, produces 137, but only a cgroup OOM sets reason OOMKilled.