What are Linux control groups (cgroups), and what actually happens in the kernel when a container is started with `--memory=512m --cpus=1.5`?
answer
- namespaces = what you see, cgroups = how much you get
- one cgroup per container, runc writes the files
- --memory -> memory.max, --cpus -> cpu.max quota/period
- systemd driver: system.slice/docker-<id>.scope
- /proc/meminfo not namespaced -> free lies
basics
~20 scgroups are the kernel feature that limits and accounts resources for a group of processes. The runtime creates a cgroup per container and writes the flags into its control files: on cgroup v2, memory.max becomes 536870912 and cpu.max becomes '150000 100000' (150 ms of CPU per 100 ms period).
solid answer
~50 scgroups group processes and let the kernel cap and account their CPU, memory, I/O and PID usage. They are the *how much* half of containment; namespaces are the *what you can see* half. When you run a container, the runtime (runc) creates a cgroup for it and writes your flags into control files under /sys/fs/cgroup. On cgroup v2, `--memory=512m` becomes `memory.max = 536870912`, and `--cpus=1.5` becomes `cpu.max = "150000 100000"` - 150 ms of CPU time per 100 ms period, i.e. one and a half cores' worth. On v1 the same values land in memory.limit_in_bytes and cpu.cfs_quota_us / cpu.cfs_period_us in separate hierarchies. The cgroup path depends on the driver: with the systemd driver it is under system.slice/docker-<id>.scope. Because containers usually get a cgroup namespace, the container sees its own cgroup as the root of /sys/fs/cgroup. `docker stats` just reads these files, and `docker update` can change limits on a live container. Note /proc/meminfo is not namespaced, so `free` inside the container still reports host memory.
code
bash · 5 linesdocker run --rm -it --memory=512m --cpus=1.5 alpine sh -c '
cat /sys/fs/cgroup/memory.max; # 536870912
cat /sys/fs/cgroup/cpu.max; # 150000 100000
cat /sys/fs/cgroup/memory.current;
cat /sys/fs/cgroup/cpu.stat'go deeper
Define cgroups as kernel resource limiting and accounting, and contrast them with namespaces in one sentence.
Map the common docker flags to the concrete v2 control files, and know where the container's cgroup lives and how to read it back.
Discuss the systemd versus cgroupfs driver, live updates during incidents, and the /proc mismatch that makes runtimes mis-size themselves.
Reason about the whole enforcement layer: which controllers you standardize on, delegation and driver consistency across the fleet, and how limits interact with node capacity planning.
## What cgroups are A control group is a kernel-managed collection of processes with resource controllers attached. Controllers implement limiting, accounting and, for some resources, prioritization: memory, cpu, io, pids, cpuset, hugetlb and others. The interface is a filesystem, normally mounted at /sys/fs/cgroup: you create a directory, write PIDs into cgroup.procs, and write numbers into control files. Children inherit and cannot exceed ancestors' limits. Pair this with namespaces to explain containers in one sentence: namespaces decide what a process can *see* (its own PIDs, mounts, network interfaces), cgroups decide how much it can *consume*. Neither is a security boundary on its own, but cgroups are what stop one container from starving the host. ## What the runtime does `docker run` hands a spec to containerd and then runc. runc creates a cgroup for the container, writes the resource limits, moves the container's init process into it, and executes the workload. Every descendant process is automatically in that cgroup, so a fork bomb inside the container still counts against the same limits (with the pids controller, `--pids-limit` caps process count outright). The path depends on the cgroup driver. With the systemd driver - the default on modern distros, and the recommended one because systemd wants to own the tree - a container's cgroup looks like /sys/fs/cgroup/system.slice/docker-<container-id>.scope. With the cgroupfs driver it is /sys/fs/cgroup/docker/<container-id>. `docker info` reports both the driver and the cgroup version. ## Flag to file mapping (cgroup v2) - `--memory=512m` -> memory.max = 536870912 (bytes). Exceeding it triggers reclaim, then an OOM kill inside the cgroup. - `--memory-reservation` -> memory.low, a soft target honored under host pressure. - `--memory-swap` -> together with --memory determines memory.swap.max. - `--cpus=1.5` -> cpu.max = "150000 100000": quota microseconds per period microseconds. The default period is 100 ms, so 150000 means the group may consume 1.5 CPU-seconds per wall second. - `--cpu-shares` -> cpu.weight, a *relative* share used only when CPUs are contended. - `--cpuset-cpus=0-3` -> cpuset.cpus, pinning to specific CPUs rather than metering time. - `--pids-limit` -> pids.max. - `--device-write-bps` -> io.max entries per block device. On cgroup v1 the same intents map to per-controller hierarchies: /sys/fs/cgroup/memory/.../memory.limit_in_bytes, /sys/fs/cgroup/cpu/.../cpu.cfs_quota_us and cpu.cfs_period_us, /sys/fs/cgroup/blkio/... . ## Seeing it from inside Modern runtimes give the container a cgroup namespace, so /sys/fs/cgroup inside shows the container's own cgroup as the root - `cat /sys/fs/cgroup/memory.max` reads back the limit. Accounting is live: memory.current is current charge, cpu.stat carries usage plus throttling counters, memory.events counts OOM events. `docker stats` is a thin reader over these files, which is why its memory number includes page cache. ## What cgroups do not do They do not virtualize /proc. Inside the container, /proc/meminfo and /proc/cpuinfo still describe the host, so `free`, core-count probes and older runtimes over-detect resources - the classic case of a JVM sizing its heap for a 64 GB host inside a 512 MB container. Modern runtimes read cgroup files instead (the JVM's container support does this by default), and lxcfs can fake /proc, but the underlying mismatch is worth naming. ## Changing limits later `docker update --memory 1g --cpus 2 <container>` rewrites the control files in place, no restart needed (some parameters, such as lowering memory below current usage, can fail or trigger reclaim). This is a genuinely useful incident tool: relieve a throttled container immediately, then fix the manifest.
- How do cgroups differ from namespaces?Namespaces isolate the view: PID, mount, network, UTS, IPC and user namespaces decide which processes, filesystems and interfaces a container can see. cgroups isolate consumption: they meter and cap CPU time, memory, I/O and process counts for a group of tasks. A container is both together plus a root filesystem; neither alone is sufficient, and neither is by itself a security boundary.
- Why does `free -m` inside a memory-limited container report the host's total RAM?cgroups constrain allocation but do not virtualize /proc, and /proc/meminfo is not namespaced, so it still describes the host. Programs must read the cgroup files (memory.max, memory.current) to learn their real budget - which is what modern runtimes such as the JVM's container support do - or run with something like lxcfs that synthesizes container-aware /proc entries.
saying these in an interview costs you the question
- Saying cgroups isolate visibility, confusing them with namespaces
- Believing --cpus pins the container to specific cores (that is --cpuset-cpus)
- Assuming limits inside the container are visible through free or /proc/cpuinfo
- Thinking limits require a container restart to change
- Treating cgroups as a security boundary that prevents escapes