skip to content

questions

6

On a Linux host, what is the difference between a namespace and a cgroup, and which of the two would you use to stop one process starving the machine of memory?

level: juniorimportance: must knowfreq 68%

answer

  1. two independent knobs, not one
  2. one is about visibility
  3. the other is about accounting
  4. eight namespace types, one hierarchy
  5. stopping a fork bomb needs pids.max

basics

~20 s

Namespaces control what a process can see: its own PIDs, mounts, network stack, hostname. Cgroups control how much it may consume: CPU time, memory, I/O, process count. Starving the machine is a resource problem, so the answer is a cgroup.

solid answer

~50 s

They are two unrelated kernel mechanisms that people conflate because they are usually used together. A **namespace** virtualises a global kernel resource so the processes inside it get their own instance of it — the kernel has separate namespace types for mounts, PIDs, network, UTS (hostname), IPC, users, cgroups and time. A process enters one via `clone(2)` with a `CLONE_NEW*` flag, `unshare(2)`, or `setns(2)` on a file under `/proc/<pid>/ns/`. A **cgroup** is a node in a hierarchy under `/sys/fs/cgroup` that accounts for and caps resource use through controllers — `cpu`, `memory`, `io`, `pids`, `cpuset` and others — using files like `memory.max` and `cpu.max`. Namespaces answer "what does this process see?"; cgroups answer "how much may it take?". A process locked into every namespace type but no cgroup can still fork-bomb the box or exhaust RAM, so memory starvation is fixed with `memory.max` on a cgroup, never with a namespace.

code

bash · 9 lines
bash
# Isolation: new PID + mount namespace, /proc reflects the new PID namespace
unshare --pid --mount --fork --mount-proc ps -e

# Limits: cap a cgroup at 100 MB of memory and half a CPU
echo "+cpu +memory" > /sys/fs/cgroup/cgroup.subtree_control
mkdir -p /sys/fs/cgroup/demo
echo 100M > /sys/fs/cgroup/demo/memory.max
echo "50000 100000" > /sys/fs/cgroup/demo/cpu.max
echo $$ > /sys/fs/cgroup/demo/cgroup.procs

go deeper

for a junior

Be able to state the split in one sentence — namespaces change what a process can see, cgroups change how much it can use — and name a few namespace types plus one cgroup limit such as memory.

for a middle

Explain the entry points (clone with CLONE_NEW* flags, unshare, setns) and the cgroup v2 file interface: cgroup.procs, cgroup.controllers, memory.max, cpu.max. Say how you check which namespace a process is in.

for a senior

Show what each mechanism fails to do alone: namespaces bound nothing, cgroups hide nothing. Be ready to say which of the two you would reach for given a real symptom on a shared host.

for a principal

Own the policy question: which isolation boundary a workload actually needs, what namespaces buy you versus a stronger boundary, and where per-workload resource caps should be set and enforced across a fleet.

## Two orthogonal mechanisms The Linux kernel keeps many resources that are conceptually *global*: the process table, the mount table, the network stack, the hostname, the System V IPC keyspace, the UID space. Namespaces let the kernel keep several independent copies of those global resources and give each process a pointer to the copy it should use. Nothing about that limits how much CPU or memory a process may take. Control groups solve the opposite problem. A cgroup does not hide anything; it groups processes so the kernel can *account* for what they use and *cap* it. A process in a cgroup with `memory.max` set can still enumerate every process on the host and see every mount — it simply cannot allocate past its cap. Both mechanisms are per-process attributes inherited across `fork(2)` and preserved across `execve(2)`, and both are exposed through the filesystem, which is why they get filed together in people's heads. ## What a namespace does The kernel currently offers these types: **mount** (`CLONE_NEWNS`) — a private copy of the mount table; **PID** (`CLONE_NEWPID`) — its own PID numbering, so the first process created inside is PID 1; **network** (`CLONE_NEWNET`) — its own interfaces, routing tables, netfilter rules and socket port space; **UTS** (`CLONE_NEWUTS`) — its own hostname and domain name; **IPC** (`CLONE_NEWIPC`) — its own System V IPC objects and POSIX message queues; **user** (`CLONE_NEWUSER`) — its own UID/GID mapping and capability set; **cgroup** (`CLONE_NEWCGROUP`) — a virtualised view of the cgroup hierarchy root; and **time** (`CLONE_NEWTIME`, Linux 5.6+) — its own boot and monotonic clock offsets. Every namespace a process belongs to is visible as a magic symlink under `/proc/<pid>/ns/`. Two processes are in the same namespace exactly when those symlinks resolve to the same inode number, which is the standard way to compare them: ```bash readlink /proc/self/ns/net /proc/1/ns/net ``` The three ways in are `clone(2)`/`clone3(2)` with `CLONE_NEW*` flags when creating a child, `unshare(2)` to detach the calling process into fresh namespaces, and `setns(2)` to join an existing one through an open file descriptor. ## What a cgroup does On a current distribution the unified hierarchy (cgroup v2) is mounted at `/sys/fs/cgroup`. Each directory is a cgroup; creating a subdirectory creates a child. The processes belonging to a cgroup are listed in its `cgroup.procs` file, and you move a process by writing its PID there. Which controllers a cgroup may use is listed in `cgroup.controllers`, and which of them it hands down to its children is written to `cgroup.subtree_control`. Controllers expose their own knobs: `memory.max` and `memory.high` for memory, `cpu.max` (a quota and a period) and `cpu.weight` for CPU, `io.max` for block I/O, `pids.max` for the number of tasks. Usage is readable back through `memory.current`, `cpu.stat`, `pids.current`. ```bash echo "+cpu +memory +pids" > /sys/fs/cgroup/cgroup.subtree_control mkdir /sys/fs/cgroup/demo echo 100M > /sys/fs/cgroup/demo/memory.max echo "50000 100000" > /sys/fs/cgroup/demo/cpu.max # 50 ms per 100 ms = half a core ``` ## Why the distinction matters in practice The clean way to remember it is *isolation versus limits*. Namespaces are a **security and correctness** feature: they stop a process seeing or touching things that are not its own. Cgroups are a **resource management** feature: they stop a process taking more than its share. Neither substitutes for the other. Put a workload in fresh PID, mount, network and user namespaces but no cgroup, and it can still allocate until the host runs out of memory, spin every core, or spawn processes until the system-wide PID limit is hit — a namespace has no notion of "too much". Conversely, cap a workload with `memory.max` and `pids.max` but give it no namespaces, and it can read `/proc` to enumerate every other process on the box, send signals to them, and see every mounted filesystem. They do meet at the edges. The **cgroup namespace** exists precisely because a process could otherwise read `/proc/self/cgroup` and learn its absolute path in the host's hierarchy; the namespace rebases that view so the process sees its own cgroup as the root. And memory accounting is a cgroup property, not a namespace property — being in a PID namespace does not give a process its own memory budget. ## The stock interview follow-through When an interviewer asks this, they usually want you to land the summary in one line — *namespaces change what you see, cgroups change what you get* — and then show you know the failure mode of each on its own. Saying "they are both how containers work" without separating them is the weak answer.

  • A process runs in its own PID, mount and network namespaces but belongs to no restricted cgroup. What can it still do to the host?
    Everything resource-related: allocate until the host is out of memory, saturate every core, fill the disk, and spawn tasks up to the system-wide limit. Namespaces do not account for anything. You need a cgroup with `memory.max`, `cpu.max` and `pids.max` to bound it.
  • What does the cgroup namespace type actually hide, given that cgroups are not an isolation feature?
    It virtualises the *path*. Without it, a process can read `/proc/self/cgroup` and learn its absolute position in the host's hierarchy, leaking the supervisor's naming and structure. Inside a cgroup namespace, that path is rebased so the process's own cgroup appears as the root.
  • Are namespace membership and cgroup membership inherited by child processes?
    Yes, both. A child created by `fork(2)` starts in exactly the same namespaces and the same cgroup as its parent, and both survive `execve(2)`. Changing either requires an explicit action — a `CLONE_NEW*` flag, `unshare(2)`, `setns(2)`, or a write to another cgroup's `cgroup.procs`.

saying these in an interview costs you the question

  • Says cgroups isolate processes from seeing each other
  • Thinks a namespace can cap CPU or memory
  • Claims a PID namespace prevents a fork bomb
  • Says cgroup v2 uses one hierarchy per controller
  • Treats namespaces and cgroups as one kernel feature

context

open as a page

In the cgroup v2 memory controller on Linux, what is the difference between writing a limit to memory.high and writing one to memory.max?

level: middleimportance: should knowfreq 42%

basics

~20 s

memory.high is a throttle: the kernel reclaims hard and stalls the offending processes, but never kills them, so usage may sit above it. memory.max is a wall: when reclaim fails to get under it, the cgroup OOM killer kills a process inside.

open as a page

On a cgroup v2 Linux host, moving a PID into a cgroup that has controllers listed in its cgroup.subtree_control fails with EBUSY. What rule is the kernel enforcing, and how does a supervisor like systemd shape its tree around it?

level: seniorimportance: should knowfreq 30%

basics

~20 s

cgroup v2 forbids a non-root cgroup from both holding processes and distributing resources to its children — the no-internal-processes rule. Processes belong in leaves; inner cgroups only enable controllers via cgroup.subtree_control. That is why systemd puts services in leaf units under slices.

open as a page

On Linux, how does a user namespace let an unprivileged user become UID 0 inside it, and why does that root not give them root on the host?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A user namespace maps IDs inside it to different IDs outside, written once to /proc/PID/uid_map and gid_map. You hold a full capability set inside, but only over resources the kernel owns through that namespace; host files are still checked against your real, unprivileged outside UID.

open as a page

On a Linux host, `ip netns list` shows a network namespace with no processes running in it at all. What keeps a namespace object alive when nothing is running inside it?

level: middleimportance: nice to knowfreq 22%

basics

~20 s

A namespace lives as long as something references it: a member process, an open file descriptor on /proc/PID/ns/*, a bind mount of that file, or a child namespace. ip netns add deliberately bind-mounts the namespace under /run/netns so it outlives the process that created it.

open as a page

A process on a Linux host gets its own mount namespace, yet it is not automatically a sealed copy of the host's mount table. Explain mount propagation — shared, private, slave and unbindable — and which type gives one-way visibility from the host into the namespace.

level: seniorimportance: nice to knowfreq 25%

basics

~20 s

Mount propagation decides whether mount and unmount events travel between mount namespaces that share a peer group. Shared propagates both ways, private neither way, slave one way only — in from the master — and unbindable additionally refuses to be used as a bind source. One-way is slave.

open as a page