skip to content

questions

27

On a Linux host, what is the difference between a namespace and a cgroup, and which of the two would you use to stop one process starving the machine of memory?

level: juniorimportance: must knowfreq 68%

answer

  1. two independent knobs, not one
  2. one is about visibility
  3. the other is about accounting
  4. eight namespace types, one hierarchy
  5. stopping a fork bomb needs pids.max

basics

~20 s

Namespaces control what a process can see: its own PIDs, mounts, network stack, hostname. Cgroups control how much it may consume: CPU time, memory, I/O, process count. Starving the machine is a resource problem, so the answer is a cgroup.

solid answer

~50 s

They are two unrelated kernel mechanisms that people conflate because they are usually used together. A **namespace** virtualises a global kernel resource so the processes inside it get their own instance of it — the kernel has separate namespace types for mounts, PIDs, network, UTS (hostname), IPC, users, cgroups and time. A process enters one via `clone(2)` with a `CLONE_NEW*` flag, `unshare(2)`, or `setns(2)` on a file under `/proc/<pid>/ns/`. A **cgroup** is a node in a hierarchy under `/sys/fs/cgroup` that accounts for and caps resource use through controllers — `cpu`, `memory`, `io`, `pids`, `cpuset` and others — using files like `memory.max` and `cpu.max`. Namespaces answer "what does this process see?"; cgroups answer "how much may it take?". A process locked into every namespace type but no cgroup can still fork-bomb the box or exhaust RAM, so memory starvation is fixed with `memory.max` on a cgroup, never with a namespace.

code

bash · 9 lines
bash
# Isolation: new PID + mount namespace, /proc reflects the new PID namespace
unshare --pid --mount --fork --mount-proc ps -e

# Limits: cap a cgroup at 100 MB of memory and half a CPU
echo "+cpu +memory" > /sys/fs/cgroup/cgroup.subtree_control
mkdir -p /sys/fs/cgroup/demo
echo 100M > /sys/fs/cgroup/demo/memory.max
echo "50000 100000" > /sys/fs/cgroup/demo/cpu.max
echo $$ > /sys/fs/cgroup/demo/cgroup.procs

go deeper

for a junior

Be able to state the split in one sentence — namespaces change what a process can see, cgroups change how much it can use — and name a few namespace types plus one cgroup limit such as memory.

for a middle

Explain the entry points (clone with CLONE_NEW* flags, unshare, setns) and the cgroup v2 file interface: cgroup.procs, cgroup.controllers, memory.max, cpu.max. Say how you check which namespace a process is in.

for a senior

Show what each mechanism fails to do alone: namespaces bound nothing, cgroups hide nothing. Be ready to say which of the two you would reach for given a real symptom on a shared host.

for a principal

Own the policy question: which isolation boundary a workload actually needs, what namespaces buy you versus a stronger boundary, and where per-workload resource caps should be set and enforced across a fleet.

## Two orthogonal mechanisms The Linux kernel keeps many resources that are conceptually *global*: the process table, the mount table, the network stack, the hostname, the System V IPC keyspace, the UID space. Namespaces let the kernel keep several independent copies of those global resources and give each process a pointer to the copy it should use. Nothing about that limits how much CPU or memory a process may take. Control groups solve the opposite problem. A cgroup does not hide anything; it groups processes so the kernel can *account* for what they use and *cap* it. A process in a cgroup with `memory.max` set can still enumerate every process on the host and see every mount — it simply cannot allocate past its cap. Both mechanisms are per-process attributes inherited across `fork(2)` and preserved across `execve(2)`, and both are exposed through the filesystem, which is why they get filed together in people's heads. ## What a namespace does The kernel currently offers these types: **mount** (`CLONE_NEWNS`) — a private copy of the mount table; **PID** (`CLONE_NEWPID`) — its own PID numbering, so the first process created inside is PID 1; **network** (`CLONE_NEWNET`) — its own interfaces, routing tables, netfilter rules and socket port space; **UTS** (`CLONE_NEWUTS`) — its own hostname and domain name; **IPC** (`CLONE_NEWIPC`) — its own System V IPC objects and POSIX message queues; **user** (`CLONE_NEWUSER`) — its own UID/GID mapping and capability set; **cgroup** (`CLONE_NEWCGROUP`) — a virtualised view of the cgroup hierarchy root; and **time** (`CLONE_NEWTIME`, Linux 5.6+) — its own boot and monotonic clock offsets. Every namespace a process belongs to is visible as a magic symlink under `/proc/<pid>/ns/`. Two processes are in the same namespace exactly when those symlinks resolve to the same inode number, which is the standard way to compare them: ```bash readlink /proc/self/ns/net /proc/1/ns/net ``` The three ways in are `clone(2)`/`clone3(2)` with `CLONE_NEW*` flags when creating a child, `unshare(2)` to detach the calling process into fresh namespaces, and `setns(2)` to join an existing one through an open file descriptor. ## What a cgroup does On a current distribution the unified hierarchy (cgroup v2) is mounted at `/sys/fs/cgroup`. Each directory is a cgroup; creating a subdirectory creates a child. The processes belonging to a cgroup are listed in its `cgroup.procs` file, and you move a process by writing its PID there. Which controllers a cgroup may use is listed in `cgroup.controllers`, and which of them it hands down to its children is written to `cgroup.subtree_control`. Controllers expose their own knobs: `memory.max` and `memory.high` for memory, `cpu.max` (a quota and a period) and `cpu.weight` for CPU, `io.max` for block I/O, `pids.max` for the number of tasks. Usage is readable back through `memory.current`, `cpu.stat`, `pids.current`. ```bash echo "+cpu +memory +pids" > /sys/fs/cgroup/cgroup.subtree_control mkdir /sys/fs/cgroup/demo echo 100M > /sys/fs/cgroup/demo/memory.max echo "50000 100000" > /sys/fs/cgroup/demo/cpu.max # 50 ms per 100 ms = half a core ``` ## Why the distinction matters in practice The clean way to remember it is *isolation versus limits*. Namespaces are a **security and correctness** feature: they stop a process seeing or touching things that are not its own. Cgroups are a **resource management** feature: they stop a process taking more than its share. Neither substitutes for the other. Put a workload in fresh PID, mount, network and user namespaces but no cgroup, and it can still allocate until the host runs out of memory, spin every core, or spawn processes until the system-wide PID limit is hit — a namespace has no notion of "too much". Conversely, cap a workload with `memory.max` and `pids.max` but give it no namespaces, and it can read `/proc` to enumerate every other process on the box, send signals to them, and see every mounted filesystem. They do meet at the edges. The **cgroup namespace** exists precisely because a process could otherwise read `/proc/self/cgroup` and learn its absolute path in the host's hierarchy; the namespace rebases that view so the process sees its own cgroup as the root. And memory accounting is a cgroup property, not a namespace property — being in a PID namespace does not give a process its own memory budget. ## The stock interview follow-through When an interviewer asks this, they usually want you to land the summary in one line — *namespaces change what you see, cgroups change what you get* — and then show you know the failure mode of each on its own. Saying "they are both how containers work" without separating them is the weak answer.

  • A process runs in its own PID, mount and network namespaces but belongs to no restricted cgroup. What can it still do to the host?
    Everything resource-related: allocate until the host is out of memory, saturate every core, fill the disk, and spawn tasks up to the system-wide limit. Namespaces do not account for anything. You need a cgroup with `memory.max`, `cpu.max` and `pids.max` to bound it.
  • What does the cgroup namespace type actually hide, given that cgroups are not an isolation feature?
    It virtualises the *path*. Without it, a process can read `/proc/self/cgroup` and learn its absolute position in the host's hierarchy, leaking the supervisor's naming and structure. Inside a cgroup namespace, that path is rebased so the process's own cgroup appears as the root.
  • Are namespace membership and cgroup membership inherited by child processes?
    Yes, both. A child created by `fork(2)` starts in exactly the same namespaces and the same cgroup as its parent, and both survive `execve(2)`. Changing either requires an explicit action — a `CLONE_NEW*` flag, `unshare(2)`, `setns(2)`, or a write to another cgroup's `cgroup.procs`.

saying these in an interview costs you the question

  • Says cgroups isolate processes from seeing each other
  • Thinks a namespace can cap CPU or memory
  • Claims a PID namespace prevents a fork bomb
  • Says cgroup v2 uses one hierarchy per controller
  • Treats namespaces and cgroups as one kernel feature

context

open as a page

On Linux, the `kill` command sends SIGTERM by default and `kill -9` sends SIGKILL. What is the difference between the two signals, and why can a process never handle SIGKILL?

level: juniorimportance: must knowfreq 85%

basics

~20 s

SIGTERM is a polite request: the process can catch it and shut down cleanly. SIGKILL is enforced by the kernel — the target never runs code for it, so buffers, lock files and in-flight work are abandoned as they are.

open as a page

On a Linux host, a program has finished running but still shows up in the process list as `<defunct>` in state Z. What is that entry, and what makes it go away?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A zombie is a process that has already terminated but whose parent has not yet collected its exit status. The kernel keeps only its process-table entry, PID and exit code; it disappears as soon as the parent waits on it.

open as a page

On a Linux server with 32 GB of RAM, `free` reports only about 200 MB free, several gigabytes under buff/cache, and 20 GB available. Is that machine short of memory, and what is the kernel doing with the RAM?

level: juniorimportance: must knowfreq 72%

basics

~20 s

No. Linux spends otherwise-idle RAM on page cache for file data and reclaims it on demand, so near-zero free memory is normal and healthy. The number that matters is available, which estimates what a new process could still get.

open as a page

On Linux, what does a process's nice value control, what range can it take, and why can an ordinary user raise a process's nice value but not lower it again?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A nice value biases how much CPU a Linux task gets when CPUs are contended. It runs from -20 (most favoured) to 19 (least favoured), default 0. Lowering it needs privilege, so an unprivileged renice is a one-way trip.

open as a page

A Linux server with 8 CPUs shows a 1-minute load average of 30, yet CPU utilisation is near idle. What does Linux's load average actually count, and what does this combination point to?

level: middleimportance: must knowfreq 68%

basics

~20 s

Linux counts both runnable tasks and tasks in uninterruptible sleep (D state) in its load average, unlike traditional Unix. High load with idle CPUs therefore means processes are blocked in the kernel waiting on storage or a hung mount, not competing for CPU.

open as a page

When a Linux host exhausts memory, how does the kernel's out-of-memory killer choose which process to kill, and what does writing to /proc/<pid>/oom_score_adj change about that choice?

level: seniorimportance: must knowfreq 58%

basics

~20 s

The kernel scores every process by how much memory freeing it would recover — mainly resident plus swap usage, relative to total memory — and kills the highest scorer. oom_score_adj, from -1000 to 1000, biases that score, with -1000 exempting a process entirely.

open as a page

In the cgroup v2 memory controller on Linux, what is the difference between writing a limit to memory.high and writing one to memory.max?

level: middleimportance: should knowfreq 42%

basics

~20 s

memory.high is a throttle: the kernel reclaims hard and stalls the offending processes, but never kills them, so usage may sit above it. memory.max is a wall: when reclaim fails to get under it, the cgroup OOM killer kills a process inside.

open as a page

A long-running command you started in an interactive SSH session keeps dying whenever the connection drops. Which signal kills it, what makes the kernel send that signal, and how would you start the job so it survives?

level: middleimportance: should knowfreq 60%

basics

~20 s

SIGHUP kills it. When the SSH connection drops the kernel hangs up the controlling terminal and signals the session, and SIGHUP's default action is to terminate. Start the job under nohup or setsid, or inside a terminal multiplexer.

open as a page

You start a long-running program from an interactive Linux shell with `&`, then the shell process is killed. The program keeps running, but its parent process ID is now 1. What happened to it, and what does that change for the process itself?

level: middleimportance: should knowfreq 45%

basics

~20 s

The program became an orphan and the kernel reparented it to PID 1 (or the nearest ancestor marked as a subreaper). It keeps running unchanged — same PID, same memory, same open files — and its exit status will be collected by its new parent.

open as a page

A process on Linux allocates a large buffer, the allocation returns successfully, and the process is later killed while writing into that buffer. Why can an allocation succeed when the memory is not actually available, and which sysctl governs that behaviour?

level: middleimportance: should knowfreq 55%

basics

~20 s

Linux overcommits: an allocation only creates a mapping, and physical pages are committed on first touch. The kernel may promise more than it can back, so the shortage surfaces later as a fault it cannot satisfy. vm.overcommit_memory selects the policy.

open as a page

A Java process on a 4 GB Linux machine shows a VSZ of 12 GB and an RSS of 800 MB in `ps` output. What does each of those two numbers measure, and why is the machine not out of memory?

level: middleimportance: should knowfreq 62%

basics

~20 s

VSZ is the size of the process's virtual address space — mappings it may use, most of which are never backed by RAM. RSS is the pages actually resident in physical memory. Only RSS consumes RAM, so 12 GB of address space on a 4 GB box is unremarkable.

open as a page

You pin a latency-sensitive Linux process to CPUs 2 and 3 with taskset. Does that reserve those CPUs for it, and what happens to the affinity when the process forks or calls exec?

level: middleimportance: should knowfreq 35%

basics

~20 s

Setting CPU affinity restricts where a task may run; it does not reserve anything. Every other task on the system can still be scheduled onto CPUs 2 and 3. Children inherit the mask across fork, and it survives exec.

open as a page

How does Linux's fair-share CPU scheduler (CFS, and EEVDF since kernel 6.6) turn a task's nice value into an actual share of CPU time?

level: middleimportance: should knowfreq 45%

basics

~20 s

Linux maps each nice value to a weight, roughly 1.25x per step with nice 0 at 1024, and hands every runnable task CPU time in proportion to its weight. Each nice step therefore shifts about 10% of the contested CPU.

open as a page

On a cgroup v2 Linux host, moving a PID into a cgroup that has controllers listed in its cgroup.subtree_control fails with EBUSY. What rule is the kernel enforcing, and how does a supervisor like systemd shape its tree around it?

level: seniorimportance: should knowfreq 30%

basics

~20 s

cgroup v2 forbids a non-root cgroup from both holding processes and distributing resources to its children — the no-internal-processes rule. Processes belong in leaves; inner cgroups only enable controllers via cgroup.subtree_control. That is why systemd puts services in leaf units under slices.

open as a page

On Linux, how does a user namespace let an unprivileged user become UID 0 inside it, and why does that root not give them root on the host?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A user namespace maps IDs inside it to different IDs outside, written once to /proc/PID/uid_map and gid_map. You hold a full capability set inside, but only over resources the kernel owns through that namespace; host files are still checked against your real, unprivileged outside UID.

open as a page

On a Linux host you send SIGKILL to a hung process, and minutes later `ps` still shows it in state D. Why has SIGKILL not taken effect, and what does the kernel do with the signal in the meantime?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The signal is recorded as pending, not lost. A task in state D is in uninterruptible sleep inside a kernel call — usually blocked on I/O — and signals are only acted on when a task heads back toward user space, which it never reaches.

open as a page

A Linux daemon is restarted and fails to bind its listening TCP port with EADDRINUSE, even though the old daemon process is gone. A helper program that the old daemon had launched is still running. How can that be, and what should the daemon have done differently?

level: seniorimportance: should knowfreq 35%

basics

~20 s

The helper inherited the listening socket. File descriptors survive fork and survive exec unless marked close-on-exec, and a socket stays bound while any process still holds it. The fix is to create such descriptors with O_CLOEXEC or SOCK_CLOEXEC, or set FD_CLOEXEC before exec.

open as a page

On a Linux server, a process sits in state D and `kill -9` has no effect on it. The host mounts a network filesystem whose server has stopped responding. What does state D mean, and why can the process not be killed?

level: seniorimportance: should knowfreq 50%

basics

~20 s

State D is uninterruptible sleep: the task is blocked inside a kernel operation that cannot be unwound safely, usually waiting on storage. SIGKILL is recorded as pending but only acted on when the task heads back to user space, which it cannot do until the operation finishes.

open as a page

A Linux application server with swap enabled becomes extremely slow while its CPUs sit mostly idle and its swap usage climbs. What is the kernel doing, and why does adding more swap not fix a memory shortage?

level: seniorimportance: should knowfreq 50%

basics

~20 s

The kernel is reclaiming pages that are still in use and immediately faulting them back from disk — thrashing. The CPUs are idle because tasks are blocked on I/O. Swap buys time and absorbs cold pages; it does not add usable memory to a working set that exceeds RAM.

open as a page

What do the Linux real-time scheduling policies SCHED_FIFO and SCHED_RR change about how a thread is scheduled, and what can a runaway SCHED_FIFO thread do to the machine?

level: seniorimportance: should knowfreq 33%

basics

~20 s

SCHED_FIFO and SCHED_RR give a thread a static priority of 1-99 that preempts every normal task; FIFO runs until it blocks or yields, RR adds a time slice between equal priorities. A runaway FIFO loop starves everything below it, which is why Linux throttles real-time tasks by default.

open as a page

On a Linux host, `ip netns list` shows a network namespace with no processes running in it at all. What keeps a namespace object alive when nothing is running inside it?

level: middleimportance: nice to knowfreq 22%

basics

~20 s

A namespace lives as long as something references it: a member process, an open file descriptor on /proc/PID/ns/*, a bind mount of that file, or a child namespace. ip netns add deliberately bind-mounts the namespace under /run/netns so it outlives the process that created it.

open as a page

In a C program on Linux, `waitpid()` stores an integer status for the child that ended. Why is that integer not simply the child's exit code, and how are you supposed to interpret it?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

The status is a packed word describing how the child ended, not just what it returned. Decode it with the macros: WIFEXITED plus WEXITSTATUS for a normal exit, WIFSIGNALED plus WTERMSIG when a signal killed it.

open as a page

You raise a limit with `ulimit` in your shell, but an already-running daemon still hits the old value. Explain how Linux resource limits (rlimits) are scoped, and the difference between a soft and a hard limit.

level: middleimportance: nice to knowfreq 34%

basics

~20 s

Rlimits are per-process attributes, inherited across fork and kept across exec. A shell's ulimit changes only that shell and children it starts afterwards, never a running daemon. The soft limit is what is enforced; the hard limit is the ceiling the soft one may be raised to.

open as a page

A process on a Linux host gets its own mount namespace, yet it is not automatically a sealed copy of the host's mount table. Explain mount propagation — shared, private, slave and unbindable — and which type gives one-way visibility from the host into the namespace.

level: seniorimportance: nice to knowfreq 25%

basics

~20 s

Mount propagation decides whether mount and unmount events travel between mount namespaces that share a peer group. Shared propagates both ways, private neither way, slave one way only — in from the master — and unbindable additionally refuses to be used as a bind source. One-way is slave.

open as a page

In a C program on Linux, what may a signal handler installed with `sigaction()` safely do, and why can calling something like `printf()` from a handler deadlock the process?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Only async-signal-safe functions, listed in the signal-safety(7) manual page — chiefly write() and _exit() — plus setting a volatile sig_atomic_t flag. printf() takes locks the interrupted code may already hold, so re-entering it can deadlock.

open as a page

What do the Linux I/O scheduling classes set by ionice (realtime, best-effort, idle) do, and why can running a backup under `ionice -c 3` have no effect at all?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

ionice tags a process's block-layer requests with an I/O class and level, so a lower-priority job yields disk service to others. It only works if the active I/O scheduler honours those tags — with the none scheduler common on NVMe, the tag is simply ignored.

open as a page