What changed between cgroup v1 and cgroup v2 on Linux, and how does that affect container resource limits?
answer
- v1: hierarchy per controller; v2: one unified tree
- subtree_control + no internal processes rule
- memory.max/high/low, cpu.max, io.max, PSI, cgroup.kill
- v2 fixes buffered-writeback attribution
- no --oom-kill-disable / kernel-memory on v2
basics
~20 sv1 had a separate hierarchy per controller with inconsistent interfaces; v2 has one unified hierarchy, one consistent API (memory.max, cpu.max, io.max), pressure metrics, and coordinated memory-plus-I/O accounting so buffered writes are attributed correctly. Some v1-only options, such as disabling the OOM killer, no longer exist.
solid answer
~50 scgroup v1 mounted one hierarchy per controller - /sys/fs/cgroup/memory, /cpu, /blkio - so a process could sit at unrelated positions in each tree, and the controllers could not cooperate. Interfaces were inconsistent (memory.limit_in_bytes, cpu.cfs_quota_us, blkio.throttle.read_bps_device). v2 is a single unified hierarchy: one tree, controllers enabled per subtree via cgroup.subtree_control, consistent naming (memory.max/high/low, cpu.max/cpu.weight, io.max), and the rule that only leaf cgroups hold processes. It adds memory.high for throttling instead of killing, memory.events counters, PSI pressure files, cgroup.kill, and - crucially - memory and I/O controllers that cooperate, so buffered writeback is charged to the right cgroup instead of escaping I/O limits. Practical container impact: v2 is the default on current distributions and is required for rootless containers with resource limits; `--oom-kill-disable` and the old kernel-memory limit are not supported on v2; `--cpu-shares` is translated into cpu.weight. Check with `docker info` before assuming which one you are debugging.
code
bash · 7 linesstat -fc %T /sys/fs/cgroup # cgroup2fs = v2, tmpfs = v1 layout
docker info --format '{{.CgroupVersion}} {{.CgroupDriver}}'
# v2
cat /sys/fs/cgroup/memory.max /sys/fs/cgroup/cpu.max
# v1
cat /sys/fs/cgroup/memory/memory.limit_in_bytes
cat /sys/fs/cgroup/cpu/cpu.cfs_quota_us /sys/fs/cgroup/cpu/cpu.cfs_period_usgo deeper
Know that two versions exist, that v2 is unified and current, and name one file from each layout.
Contrast the hierarchies and the renamed interfaces, and know which docker flags behave differently on v2.
Explain writeback attribution, memory.high versus memory.max, PSI, delegation and rootless implications, and write version-aware tooling.
Weigh fleet migration: kernel and distro floors, monitoring rewrites, features gained (I/O QoS, PSI, safe delegation) versus v1-only options being dropped.
## The v1 design and its problems In cgroup v1 every controller had its own hierarchy, mounted separately: /sys/fs/cgroup/memory, /sys/fs/cgroup/cpu, /sys/fs/cgroup/cpuacct, /sys/fs/cgroup/blkio, and so on. A process could be in group A in the memory tree and group B in the blkio tree. That flexibility was mostly a mistake: it made delegation to unprivileged users unsafe, and it prevented controllers from cooperating. The canonical casualty is buffered writeback - the memory controller owned the page cache, the blkio controller owned the device queue, and neither knew which cgroup dirtied a page, so writeback I/O was effectively unattributable and blkio throttling only reliably affected direct I/O. The interfaces were also inconsistent in naming and units: memory.limit_in_bytes, memory.memsw.limit_in_bytes for memory-plus-swap, cpu.cfs_quota_us with cpu.cfs_period_us, cpu.shares defaulting to 1024, blkio.throttle.read_bps_device. Accounting was split (cpuacct separate from cpu), and kernel memory accounting was notoriously buggy. ## What v2 changed One unified hierarchy, mounted once at /sys/fs/cgroup with filesystem type cgroup2. Controllers are turned on for children by writing to cgroup.subtree_control, which makes delegation to an unprivileged user or a container safe and explicit. The "no internal processes" rule says a cgroup with enabled controllers for its children may not itself hold processes, forcing a clean leaf structure. Interfaces were normalized: - Memory: memory.max (hard limit), memory.high (throttling target - allocations slow down and reclaim is forced rather than the process being killed), memory.low and memory.min (protection under pressure), memory.current, memory.stat, memory.events with oom and oom_kill counters, memory.swap.max as a separate knob instead of the confusing memsw combination. - CPU: cpu.max as "quota period", cpu.weight (1-10000, default 100) replacing cpu.shares, and cpu.stat carrying usage plus nr_throttled and throttled_usec. - I/O: io.max for per-device bandwidth and IOPS caps, io.stat, io.weight (needs the BFQ scheduler), plus io.latency and io.cost for quality-of-service policies that v1 never had. - New capabilities: PSI (pressure stall information) files - cpu.pressure, memory.pressure, io.pressure - quantifying time lost to resource starvation, and cgroup.kill to terminate an entire subtree atomically. Because the memory and io controllers now live in the same tree, cgroup writeback attributes dirty page flushes to the cgroup that dirtied them (with filesystem support, such as ext4 and btrfs), so I/O limits finally apply to ordinary buffered writes. ## What it means for containers Current distributions boot with v2 by default. Docker and containerd support both; `docker info` reports the cgroup version and driver, and the systemd driver is the recommended pairing on v2 because systemd owns the tree and delegates a subtree to the runtime. Behavioral differences that bite: - `--oom-kill-disable` and `--kernel-memory` are v1-only and are rejected or ignored on v2. There is no way to disable the OOM killer for a cgroup on v2 by design. - `--cpu-shares` still works but is converted into a cpu.weight value; do not expect the literal number in the file. - Swap is configured separately (memory.swap.max) rather than as a combined memory+swap limit, so old `--memory-swap` reasoning translates differently. - Rootless containers with working resource limits effectively require v2 plus systemd delegation; on v1 an unprivileged user cannot safely be given controller access. - Monitoring and any script that reads /sys/fs/cgroup paths must handle both layouts; tools that only knew memory.limit_in_bytes silently report nothing on v2. ## Debugging posture Start by establishing which version the host runs (`stat -fc %T /sys/fs/cgroup` returns cgroup2fs for v2), then read the right file names. On v2, memory.events and the PSI files are usually the fastest route from a vague "the container is slow or dying" report to a specific answer: oom_kill counts prove kills, memory.pressure and cpu.pressure quantify stalls.
- How do you limit a container's disk throughput, and why did that often fail to bite on cgroup v1?Use per-device caps: --device-write-bps /dev/nvme0n1:10mb and its read and IOPS siblings, which land in io.max on v2 (blkio.throttle.* on v1). On v1 the memory and blkio controllers lived in separate hierarchies, so pages dirtied by a container and flushed later by kernel writeback threads could not be attributed back to it - throttling therefore applied mainly to direct I/O and buffered writes escaped. v2's unified hierarchy enables cgroup writeback, charging those flushes to the originating cgroup on filesystems that support it. Note limits are per block device, so target the device that actually backs the container's data.
- What is memory.high, and why is it interesting for containers?memory.high is a v2 throttling threshold: when usage exceeds it the kernel forces aggressive reclaim and slows the allocating process rather than killing it, while memory.max remains the hard limit that triggers an OOM kill. It gives a soft landing - back pressure and a chance to shed load or be noticed by monitoring before the process is SIGKILLed. Docker does not expose it directly, so it is usually set by systemd or by writing the file, but knowing it exists explains why v2 offers options v1 simply did not have.
saying these in an interview costs you the question
- Believing v2 merely renamed files with no behavioral change
- Expecting --oom-kill-disable to work on a v2 host
- Assuming blkio limits on v1 restrict ordinary buffered writes
- Reading only memory.limit_in_bytes in tooling and reporting nothing on v2 hosts
- Confusing cpu.weight (relative) with cpu.max (absolute quota)