skip to content

What are Linux control groups (cgroups), and what actually happens in the kernel when a container is started with `--memory=512m --cpus=1.5`?

level: middleimportance: must knowfreq 58%

answer

  1. namespaces = what you see, cgroups = how much you get
  2. one cgroup per container, runc writes the files
  3. --memory -> memory.max, --cpus -> cpu.max quota/period
  4. systemd driver: system.slice/docker-<id>.scope
  5. /proc/meminfo not namespaced -> free lies

basics

~20 s

cgroups are the kernel feature that limits and accounts resources for a group of processes. The runtime creates a cgroup per container and writes the flags into its control files: on cgroup v2, memory.max becomes 536870912 and cpu.max becomes '150000 100000' (150 ms of CPU per 100 ms period).

solid answer

~50 s

cgroups group processes and let the kernel cap and account their CPU, memory, I/O and PID usage. They are the *how much* half of containment; namespaces are the *what you can see* half. When you run a container, the runtime (runc) creates a cgroup for it and writes your flags into control files under /sys/fs/cgroup. On cgroup v2, `--memory=512m` becomes `memory.max = 536870912`, and `--cpus=1.5` becomes `cpu.max = "150000 100000"` - 150 ms of CPU time per 100 ms period, i.e. one and a half cores' worth. On v1 the same values land in memory.limit_in_bytes and cpu.cfs_quota_us / cpu.cfs_period_us in separate hierarchies. The cgroup path depends on the driver: with the systemd driver it is under system.slice/docker-<id>.scope. Because containers usually get a cgroup namespace, the container sees its own cgroup as the root of /sys/fs/cgroup. `docker stats` just reads these files, and `docker update` can change limits on a live container. Note /proc/meminfo is not namespaced, so `free` inside the container still reports host memory.

code

bash · 5 lines
bash
docker run --rm -it --memory=512m --cpus=1.5 alpine sh -c '
  cat /sys/fs/cgroup/memory.max;   # 536870912
  cat /sys/fs/cgroup/cpu.max;      # 150000 100000
  cat /sys/fs/cgroup/memory.current;
  cat /sys/fs/cgroup/cpu.stat'

go deeper

for a junior

Define cgroups as kernel resource limiting and accounting, and contrast them with namespaces in one sentence.

for a middle

Map the common docker flags to the concrete v2 control files, and know where the container's cgroup lives and how to read it back.

for a senior

Discuss the systemd versus cgroupfs driver, live updates during incidents, and the /proc mismatch that makes runtimes mis-size themselves.

for a principal

Reason about the whole enforcement layer: which controllers you standardize on, delegation and driver consistency across the fleet, and how limits interact with node capacity planning.

## What cgroups are A control group is a kernel-managed collection of processes with resource controllers attached. Controllers implement limiting, accounting and, for some resources, prioritization: memory, cpu, io, pids, cpuset, hugetlb and others. The interface is a filesystem, normally mounted at /sys/fs/cgroup: you create a directory, write PIDs into cgroup.procs, and write numbers into control files. Children inherit and cannot exceed ancestors' limits. Pair this with namespaces to explain containers in one sentence: namespaces decide what a process can *see* (its own PIDs, mounts, network interfaces), cgroups decide how much it can *consume*. Neither is a security boundary on its own, but cgroups are what stop one container from starving the host. ## What the runtime does `docker run` hands a spec to containerd and then runc. runc creates a cgroup for the container, writes the resource limits, moves the container's init process into it, and executes the workload. Every descendant process is automatically in that cgroup, so a fork bomb inside the container still counts against the same limits (with the pids controller, `--pids-limit` caps process count outright). The path depends on the cgroup driver. With the systemd driver - the default on modern distros, and the recommended one because systemd wants to own the tree - a container's cgroup looks like /sys/fs/cgroup/system.slice/docker-<container-id>.scope. With the cgroupfs driver it is /sys/fs/cgroup/docker/<container-id>. `docker info` reports both the driver and the cgroup version. ## Flag to file mapping (cgroup v2) - `--memory=512m` -> memory.max = 536870912 (bytes). Exceeding it triggers reclaim, then an OOM kill inside the cgroup. - `--memory-reservation` -> memory.low, a soft target honored under host pressure. - `--memory-swap` -> together with --memory determines memory.swap.max. - `--cpus=1.5` -> cpu.max = "150000 100000": quota microseconds per period microseconds. The default period is 100 ms, so 150000 means the group may consume 1.5 CPU-seconds per wall second. - `--cpu-shares` -> cpu.weight, a *relative* share used only when CPUs are contended. - `--cpuset-cpus=0-3` -> cpuset.cpus, pinning to specific CPUs rather than metering time. - `--pids-limit` -> pids.max. - `--device-write-bps` -> io.max entries per block device. On cgroup v1 the same intents map to per-controller hierarchies: /sys/fs/cgroup/memory/.../memory.limit_in_bytes, /sys/fs/cgroup/cpu/.../cpu.cfs_quota_us and cpu.cfs_period_us, /sys/fs/cgroup/blkio/... . ## Seeing it from inside Modern runtimes give the container a cgroup namespace, so /sys/fs/cgroup inside shows the container's own cgroup as the root - `cat /sys/fs/cgroup/memory.max` reads back the limit. Accounting is live: memory.current is current charge, cpu.stat carries usage plus throttling counters, memory.events counts OOM events. `docker stats` is a thin reader over these files, which is why its memory number includes page cache. ## What cgroups do not do They do not virtualize /proc. Inside the container, /proc/meminfo and /proc/cpuinfo still describe the host, so `free`, core-count probes and older runtimes over-detect resources - the classic case of a JVM sizing its heap for a 64 GB host inside a 512 MB container. Modern runtimes read cgroup files instead (the JVM's container support does this by default), and lxcfs can fake /proc, but the underlying mismatch is worth naming. ## Changing limits later `docker update --memory 1g --cpus 2 <container>` rewrites the control files in place, no restart needed (some parameters, such as lowering memory below current usage, can fail or trigger reclaim). This is a genuinely useful incident tool: relieve a throttled container immediately, then fix the manifest.

  • How do cgroups differ from namespaces?
    Namespaces isolate the view: PID, mount, network, UTS, IPC and user namespaces decide which processes, filesystems and interfaces a container can see. cgroups isolate consumption: they meter and cap CPU time, memory, I/O and process counts for a group of tasks. A container is both together plus a root filesystem; neither alone is sufficient, and neither is by itself a security boundary.
  • Why does `free -m` inside a memory-limited container report the host's total RAM?
    cgroups constrain allocation but do not virtualize /proc, and /proc/meminfo is not namespaced, so it still describes the host. Programs must read the cgroup files (memory.max, memory.current) to learn their real budget - which is what modern runtimes such as the JVM's container support do - or run with something like lxcfs that synthesizes container-aware /proc entries.

saying these in an interview costs you the question

  • Saying cgroups isolate visibility, confusing them with namespaces
  • Believing --cpus pins the container to specific cores (that is --cpuset-cpus)
  • Assuming limits inside the container are visible through free or /proc/cpuinfo
  • Thinking limits require a container restart to change
  • Treating cgroups as a security boundary that prevents escapes

context