In the cgroup v2 memory controller on Linux, what is the difference between writing a limit to memory.high and writing one to memory.max?
answer
- one throttles, one kills
- usage may exceed the softer one
- a penalty sleep, not a refusal
- memory.events counts which you hit
- both default to the string max
basics
~20 smemory.high is a throttle: the kernel reclaims hard and stalls the offending processes, but never kills them, so usage may sit above it. memory.max is a wall: when reclaim fails to get under it, the cgroup OOM killer kills a process inside.
solid answer
~50 sBoth are set in bytes (or `max`, the default) on a cgroup under `/sys/fs/cgroup`, but they enforce differently. `memory.high` is a soft throttle — once `memory.current` goes above it, the kernel applies reclaim pressure to that cgroup and puts allocating tasks to sleep for a penalty period proportional to the overage. Usage can legitimately exceed `memory.high`; nothing is killed. `memory.max` is the hard limit: on an allocation the kernel cannot satisfy by reclaiming, it invokes the cgroup out-of-memory killer and kills a process inside that cgroup. So a workload sitting above `memory.high` degrades — more reclaim, more page faults, higher latency — while a workload that hits `memory.max` disappears. The counters in `memory.events` (`high`, `max`, `oom`, `oom_kill`) tell you which of the two you are hitting, and `memory.pressure` shows how much time tasks are stalled on memory. A common pattern is to set `high` below `max` so you get a visible warning band before anything dies.
code
bash · 9 linesecho "+memory" > /sys/fs/cgroup/cgroup.subtree_control
mkdir -p /sys/fs/cgroup/svc
# warn band at 200 MB, hard wall at 256 MB
echo 200M > /sys/fs/cgroup/svc/memory.high
echo 256M > /sys/fs/cgroup/svc/memory.max
echo $$ > /sys/fs/cgroup/svc/cgroup.procs
cat /sys/fs/cgroup/svc/memory.current /sys/fs/cgroup/svc/memory.eventsgo deeper
Know that a cgroup can cap memory and that the file to set is memory.max, expressed in bytes, and that a process going past it is killed rather than told no.
Explain both files precisely: memory.high reclaims and throttles without killing, memory.max invokes the cgroup OOM killer. Mention that memory.events tells you which one fired.
Be ready to choose values for a real service, defend a gap between high and max as an early-warning band, and use memory.pressure to distinguish a throttled workload from a slow one.
Own the policy across a fleet: where the warning band sits relative to the hard cap, whether swap is permitted inside limits, and whether a cgroup dies as a unit or one process at a time.
## The knobs, and what each one is for The cgroup v2 memory controller exposes a small set of files in every cgroup that has `memory` enabled by its parent's `cgroup.subtree_control`: - `memory.current` — bytes currently charged to this cgroup and its descendants. - `memory.min` — hard protection; memory below this is never reclaimed from the cgroup, even under global pressure. - `memory.low` — best-effort protection; reclaim avoids this cgroup while it is under the value, unless there is no alternative. - `memory.high` — throttle limit. - `memory.max` — hard limit. - `memory.swap.max` — cap on swap usage for the cgroup. - `memory.events` — a counter file with `low`, `high`, `max`, `oom` and `oom_kill` fields. - `memory.stat` — a detailed breakdown (anon, file, slab, and much more). The four limits form a ladder: `min` and `low` protect memory *for* the cgroup, `high` and `max` protect the rest of the system *from* it. ## memory.high: throttle, do not kill When `memory.current` exceeds `memory.high`, the kernel does two things. First, it runs direct reclaim against that cgroup to try to bring usage back down — writing back dirty page cache, dropping clean pages, swapping anonymous memory if swap is available and permitted. Second, and this is the part people miss, it *penalises the allocating task*: on return to userspace after an allocation that pushed the cgroup over, the task is put to sleep, and the sleep grows with how far over the limit the cgroup is. The result is a workload that keeps running but gets slower. That is deliberate — it converts a hard failure into back-pressure, giving an operator or a supervisor time to react. Because there is no enforcement wall, `memory.current` can and does sit above `memory.high` indefinitely if the workload keeps allocating faster than reclaim frees. The cost is that the throttling is invisible in ordinary CPU metrics: the process is not consuming CPU, it is sleeping and thrashing. This is what `memory.pressure` (the PSI file) is for — it reports the share of wall time tasks in the cgroup spent stalled waiting on memory. ## memory.max: the wall `memory.max` is enforced at charge time. When a page charge would push the cgroup over the limit, the kernel first tries reclaim within the cgroup. If reclaim cannot free enough, the cgroup out-of-memory killer runs and kills a process **inside that cgroup** — the rest of the host is untouched. The killed process dies from `SIGKILL`, which it cannot catch, so it writes nothing to its own log about why it went away; the evidence is in `memory.events` (`oom_kill` increments) and in the kernel log. Two related knobs matter here. `memory.oom.group` makes the kill atomic for the whole cgroup — either every process in it dies or none does, which is what you want when the processes are one logical unit and a survivor is useless. And `memory.swap.max` decides whether anonymous memory can be pushed to swap before the limit is hit; setting it to `0` means anonymous pressure goes straight to OOM. ```bash cat /sys/fs/cgroup/svc/memory.events # low 0 # high 41 # max 3 # oom 1 # oom_kill 1 ``` Reading that file is the fastest way to answer "was this cgroup throttled, squeezed, or killed?" — `high` counts times usage went over the throttle, `max` counts times an allocation was blocked at the hard limit, and `oom_kill` counts processes actually killed. ## Choosing values The useful configuration is both, with a gap: `memory.high` at the level where you want to be told something is wrong, `memory.max` a margin above it as the backstop that protects the host. A workload with a genuine leak will cross `high`, stall visibly in `memory.pressure`, and only die at `max` — by which time you have signal. Setting only `max` gives you a binary outcome: fine, then dead. Setting only `high` protects nothing, because a runaway allocator can outrun reclaim and drag the whole machine into pressure. One caveat worth stating in an interview: `memory.high` throttling is not free. A workload parked just above `high` burns real time in reclaim, and if that memory is page cache it needs, the cgroup can thrash — repeatedly evicting and re-reading the same pages. If `memory.pressure` shows sustained stall, the correct fix is more memory or less work, not a higher throttle. ## Defaults Both files default to the literal string `max`, meaning no limit, and both accept `max` to reset. Values are written in bytes, with the usual `K`/`M`/`G` suffixes accepted. Limits are hierarchical: a child cannot escape an ancestor's `memory.max`, so the effective limit is the smallest along the path from the root.
- A cgroup shows memory.current well above its memory.high value. Is that a misconfiguration?No — `memory.high` is not enforced as a ceiling. It triggers reclaim and a throttling sleep for allocating tasks, but usage is allowed to sit above it. What it does mean is that the cgroup is being actively squeezed, so check `memory.pressure` and the `high` counter in `memory.events` to see how much time is being lost to it.
- How do memory.min and memory.low differ from these two limits?They point the other way: `min` and `low` protect memory *for* the cgroup against reclaim, rather than capping it. Memory under `memory.min` is never reclaimed from the cgroup even under host pressure; `memory.low` is best-effort protection that reclaim will honour unless there is nothing else to take.
- What does memory.oom.group change about a hit on memory.max?It makes the kill cover the whole cgroup instead of one process. When set, hitting the hard limit kills every process in the cgroup as a unit. That is the right setting when the processes form one logical service and a surviving helper with its main process gone is worse than a clean death.
saying these in an interview costs you the question
- Thinks memory.high is a hard ceiling that cannot be exceeded
- Says exceeding memory.high triggers an OOM kill
- Believes hitting memory.max kills a process outside the cgroup
- Assumes an unset limit means the parent's value applies literally
- Says allocations past memory.max just fail with ENOMEM