skip to content

questions

4

What are Linux namespaces, which kinds exist, and how do they make a container look like a separate machine?

level: juniorimportance: must knowfreq 70%

answer

  1. pid, net, mnt, uts, ipc, user, cgroup, time
  2. own view of a global kernel resource
  3. /proc/<pid>/ns symlinks = inode identity
  4. clone / unshare / setns
  5. isolation of visibility, not of consumption

basics

~20 s

Namespaces are a kernel feature that gives a process its own view of a global resource. Separate namespaces exist for process IDs, network stacks, mounts, hostname, IPC, users and cgroups. A container is just a process placed in a fresh set of them.

solid answer

~50 s

A namespace virtualises one kernel resource so processes inside see their own instance of it. The kinds: - **PID** — its own process tree; the first process is PID 1 and cannot see host processes. - **Network (net)** — its own interfaces, IP addresses, routes, iptables rules, port space. - **Mount (mnt)** — its own filesystem mount table, so it sees the image's root, not the host's. - **UTS** — its own hostname and domain name. - **IPC** — its own System V IPC and POSIX message queues. - **User** — its own UID/GID mapping, so root inside can be an unprivileged UID outside. - **Cgroup** — its own view of the cgroup hierarchy root. - **Time** — its own boot/monotonic clock offsets (newer kernels). `docker run` creates a fresh set (all but user, unless userns-remap is enabled) and starts your process inside them. Namespaces provide **isolation of visibility**; limiting how much CPU or memory that process may consume is a different mechanism.

go deeper

for a junior

List the namespace types with a one-line effect each, and state that a container is a process started inside a fresh set of them.

for a middle

Add clone/unshare/setns, /proc/<pid>/ns identity, and that memberships are per-type and independently shareable.

for a senior

Stress that namespaces are visibility isolation only, that the kernel is shared, and enumerate the complementary controls that make containers actually safe to run.

for a principal

Position namespaces on the isolation spectrum against VMs and sandboxed runtimes, and decide when shared-kernel isolation is insufficient for a given tenancy model.

## The core idea A Linux kernel normally exposes one global instance of each resource: one process table, one network stack, one mount table, one hostname. A **namespace** wraps one of those resources so that a set of processes sees its own independent instance. Processes in different namespaces cannot see each other's entries, and identifiers may collide harmlessly — PID 1 can exist many times, and `eth0` can exist in every container. A container is not a kernel object. It is an ordinary process (plus its children) that was started with a fresh set of namespaces, a restricted filesystem root, resource limits, and a reduced privilege set. Namespaces supply the "looks like its own machine" half. ## The namespaces, and what each hides - **PID** — a private process tree. The first process becomes PID 1 inside; host processes are invisible. Nesting is one-way: the host can see container processes under their real host PIDs, but not vice versa. Destroying the namespace (PID 1 exits) kills everything inside. - **Network (net)** — a complete private network stack: interfaces, addresses, routing table, netfilter/iptables rules, sockets, and the whole 0–65535 port range. Two containers can both bind port 8080. A fresh net namespace starts with only a loopback device; connectivity is created by moving one end of a veth pair into it and bridging the other end. - **Mount (mnt)** — a private mount table. This is what makes the container see the image's root filesystem. Note that the *namespace* isolates the mount list; the root change itself is done with `pivot_root`, and volumes are simply mounts made inside this namespace. - **UTS** — hostname and NIS domain. This is why `hostname` inside a container returns the container ID by default. (The name is historical: UNIX Time-sharing System.) - **IPC** — System V IPC objects (shared memory segments, semaphores, message queues) and POSIX message queues. Without it, one container could attach to another's shared memory by key. - **User** — maps UIDs/GIDs between inside and outside. UID 0 inside can map to an unprivileged host UID, so "root in the container" holds capabilities only within that namespace. This is the security-critical one and the foundation of rootless containers. - **Cgroup** — hides the container's position in the cgroup tree so `/proc/self/cgroup` shows a root-relative path rather than leaking the host layout. - **Time** — offsets the boot and monotonic clocks (Linux 5.6+); rarely used, mostly for checkpoint/restore. There is no per-namespace wall clock. ## Namespaces are per-resource and independently composable This is the point candidates most often miss. A process does not "have a namespace"; it has one membership per type, and each can be shared or private independently. That composability is what makes patterns like a sidecar sharing a network namespace with the main container, or a debug container sharing a PID namespace, possible at all — the mix is chosen per type. You can see the memberships as symlinks in `/proc/<pid>/ns/`. Each points to a pseudo-file whose inode number identifies the namespace, so comparing the link targets of two processes tells you exactly which namespaces they share. A namespace lives as long as it has a member process or something holding a reference (a bind mount or an open file descriptor to the ns file); otherwise the kernel destroys it. ## Creating and joining Three syscalls matter. `clone()` with `CLONE_NEW*` flags creates a child in new namespaces — this is how a runtime starts a container. `unshare()` moves the *calling* process into new namespaces, which is what the `unshare` command-line tool wraps. `setns()` joins an existing namespace given a file descriptor from `/proc/<pid>/ns/…`, which is what `nsenter` and `docker exec` use. ## What namespaces are not - **Not resource limits.** A container in its own namespaces can still consume all host CPU and memory; capping that is a separate kernel mechanism. - **Not a security boundary by themselves.** All containers share one kernel. Namespaces restrict *what you can name and see*; the remaining protection comes from dropped capabilities, seccomp filters, LSMs like AppArmor/SELinux, and read-only or `nodev` mounts. A kernel exploit crosses every namespace. - **Not virtualisation.** There is no second kernel, no emulated hardware, and no separate boot — which is why containers start in milliseconds and why the host kernel version is the one your container actually runs on. ## Seeing it concretely On the host, `lsns` lists namespaces and their member processes. `ls -l /proc/<pid>/ns` shows one process's memberships. Running `unshare --pid --fork --mount-proc bash` gives you a shell where `ps aux` shows two processes — a container-like environment built by hand in one command, which is the most convincing demonstration that there is no magic container object underneath.

  • Do namespaces limit how much CPU or memory a container can use?
    No. Namespaces control visibility and naming, not consumption; a process alone in fresh namespaces can still exhaust the host's CPU and memory. Enforcing limits is the job of a different kernel mechanism that accounts and caps resource usage per group of processes. Interviewers often pair the two because a real container runtime always applies both.
  • Why can two containers both bind to port 8080 without conflicting?
    Each has its own network namespace, and a network namespace contains a complete independent network stack including its own port number space. The binds happen in different stacks, so they never collide. A conflict appears only when both are published to the same host port, because that mapping lives in the host's namespace.

An office building where each tenant gets their own phone extension list, their own floor plan and their own mailroom — same building, same power supply, but nobody can dial or walk into anyone else's space.

saying these in an interview costs you the question

  • Saying namespaces enforce CPU/memory limits — that is a separate mechanism.
  • Claiming each container gets its own kernel or a lightweight VM.
  • Treating namespaces alone as a security boundary, ignoring capabilities, seccomp and LSMs.
  • Thinking a process has one single namespace rather than one membership per namespace type.
  • Believing PID 1 inside a container is the host's init.

context

open as a page

Inside a container, your application runs as PID 1. What does the Linux kernel treat differently about PID 1, and what problems does that cause?

level: middleimportance: must knowfreq 62%

basics

~20 s

PID 1 is special: default signal handlers are ignored, so it will not die from SIGTERM unless it handles it, and it inherits orphaned children and must reap them. An app that does neither ignores graceful shutdown and accumulates zombie processes.

open as a page

When you run a container with the default bridge network, what does the Linux kernel actually set up so it can reach the internet, and why can it still bind port 8080 while another container also uses 8080?

level: middleimportance: should knowfreq 55%

basics

~20 s

The container gets its own network namespace — a private stack with its own interfaces, routes, iptables rules and full port range. Docker creates a veth pair, puts one end inside as eth0 and attaches the other to the docker0 bridge, then NATs outbound traffic via the host.

open as a page

How can one container be made to share another container's network or process namespace, and which command lets you inspect a running container's namespaces from the host?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Namespace membership is per type, so containers can share some and not others: --network container:<name>, --pid container:<name>, --ipc container:<name>, or host for the host's own. From the host, nsenter -t <pid> -n -p -m joins a running container's namespaces via /proc/<pid>/ns/.

open as a page