skip to content

A process on a Linux host gets its own mount namespace, yet it is not automatically a sealed copy of the host's mount table. Explain mount propagation — shared, private, slave and unbindable — and which type gives one-way visibility from the host into the namespace.

level: seniorimportance: nice to knowfreq 25%

answer

  1. a copy with a relationship, not a snapshot
  2. events travel between peer groups
  3. one direction only is the useful case
  4. mountinfo tags name the group
  5. systemd leaves the root shared

basics

~20 s

Mount propagation decides whether mount and unmount events travel between mount namespaces that share a peer group. Shared propagates both ways, private neither way, slave one way only — in from the master — and unbindable additionally refuses to be used as a bind source. One-way is slave.

solid answer

~50 s

A new mount namespace starts as a *copy* of the parent's mount table, and each entry keeps its propagation type, so events can keep flowing between them. `shared` puts mounts in a peer group: a mount or unmount under that point in either namespace shows up in the other. `private` isolates it — events go neither way. `slave` makes the mount receive events from its master peer group but send none back, which is the one-way behaviour: a disk mounted on the host appears inside, while mounts made inside stay invisible outside. `unbindable` is private plus a refusal to serve as the source of a bind mount, which stops recursive bind mounts from exploding. You set them with `mount --make-shared`, `--make-private`, `--make-slave`, `--make-unbindable` and their recursive `--make-r*` forms, and read the current state in the optional fields of `/proc/self/mountinfo` (`shared:N`, `master:N`) or via `findmnt -o TARGET,PROPAGATION`. On a systemd host `/` is shared by default, which is why sandboxes explicitly re-mark their tree.

code

bash · 8 lines
bash
# what propagation is in effect right now
findmnt -o TARGET,PROPAGATION /

# host -> namespace only: receives host mounts, leaks nothing back
mount --make-rslave /mnt/hostdata

# raw view: 'shared:N' = peer group, 'master:N' with no shared = slave
grep hostdata /proc/self/mountinfo

go deeper

for a junior

Know that a mount namespace gives a process its own mount table and that the table starts as a copy of the parent's rather than as an empty one.

for a middle

Name the four propagation types and what each does to mount and unmount events, and show where to read the current type — findmnt's PROPAGATION column or the mountinfo tags.

for a senior

Diagnose from symptoms: a mount leaking out means shared was left in place; a host disk never appearing inside means the tree was made private instead of slave. Know the order in which to re-mark a new namespace.

for a principal

Own the convention for how sandboxes on your fleet mark their trees, and the tradeoff between a strictly private tree and one that can receive host-side mounts after startup.

## Why propagation exists at all When `CLONE_NEWNS` gives a process a new mount namespace, the kernel copies the current mount table into it. If that were the end of the story, the copy would be frozen: plug in a disk on the host after the copy was taken and the isolated process could never see it, which breaks any workflow where a supervisor wants to hand a filesystem to an already-running sandbox. Shared subtrees solve this. Each mount belongs to a *peer group*, and its propagation type determines how mount and unmount events move between members of that group. The types are per-mount, not per-namespace, so one namespace can be shared in one subtree and private in another. ## The four types **shared (`MS_SHARED`)** — the mount is a member of a peer group and events propagate in **both** directions. Mount something under it in one namespace and the mount appears at the corresponding point in every peer. This is the default on a systemd host for `/`, deliberately, so that things like removable media and network mounts reach services that were started earlier. **private (`MS_PRIVATE`)** — the mount belongs to no peer group. Nothing propagates in, nothing propagates out. This is the safe default for a sandbox that must not leak, and it is what `unshare --mount` applies by default when it creates the namespace. **slave (`MS_SLAVE`)** — the mount has a master peer group and receives events from it, but its own mounts and unmounts do not travel back. This is the asymmetric case: the host can hand filesystems in, and the namespace cannot push anything out. It is the setting sandbox tooling reaches for when it bind-mounts a host directory into an isolated tree. **unbindable (`MS_UNBINDABLE`)** — private, and additionally the mount may not be used as the *source* of a bind mount. Its purpose is to prevent the pathological growth you get when you recursively bind-mount a tree that contains a shared mount of itself; marking such points unbindable stops the recursion. Each type has a recursive form applied with the `r` variants, which set the type on the mount and everything under it. ## Reading the current state The authoritative source is `/proc/self/mountinfo`. Each line carries optional fields before the separator: ```bash grep ' / ' /proc/self/mountinfo # 25 1 259:2 / / rw,relatime shared:1 - ext4 /dev/nvme0n1p2 rw ``` `shared:1` means this mount is in peer group 1. A line carrying `master:1` and no `shared:` tag is a **slave** of peer group 1. A line carrying both is a slave that is itself shared with further peers. A line with neither is private. `unbindable` appears as its own tag. The friendlier view is `findmnt -o TARGET,PROPAGATION`, which prints the type per mount point directly. ## The failure modes you get asked about **"My container-style sandbox leaked a mount onto the host."** The tree was left shared. Because the copy inherits propagation, a mount made inside a shared subtree propagates straight back out. Marking the tree `rslave` or `rprivate` after creating the namespace is the fix. **"A filesystem mounted on the host after startup is invisible inside."** The tree was made fully private. Nothing propagates in. If the sandbox is supposed to see later host mounts under a particular path, that path needs to be slave rather than private. **"Unmounting on the host did not free the device."** A propagated copy of the mount still exists in another namespace, holding a reference. Until every peer's copy is gone, the filesystem stays busy — which is also why an ordinary `umount` on the host can appear to succeed while the block device is still in use. ## Setting it deliberately ```bash # make a subtree strictly one-way: host -> namespace mount --make-rslave /mnt/hostdata # fully isolate a subtree in both directions mount --make-rprivate /srv/sandbox ``` The order matters: propagation is set on mounts as they exist now, so a sandbox typically creates its mount namespace, immediately re-marks its whole tree, and only then performs the mounts it wants to keep to itself. Doing it the other way round means the first mounts have already propagated. ## The interview framing The question underneath all of this is whether you understand that a mount namespace is a *copy with a relationship*, not a snapshot. Someone who says "a new mount namespace is completely isolated" has not hit the leak, and someone who says "it is a frozen copy" has not hit the invisible-disk case. Both symptoms come from the same mechanism seen from opposite ends.

  • A sandbox made a bind mount inside its own mount namespace and it appeared on the host. What happened?
    The mount point was still `shared`, inherited from the parent namespace when the table was copied. Shared propagation is bidirectional, so the new mount propagated back into the peer group. Re-marking the tree with `mount --make-rslave` or `--make-rprivate` immediately after creating the namespace prevents it.
  • How do you tell a slave mount from a private one by reading /proc/self/mountinfo?
    By the optional fields. A slave carries a `master:N` tag naming the peer group it receives from; a private mount carries no propagation tag at all. If a line shows both `shared:N` and `master:M`, it is a slave that is itself shared with its own peers.
  • Why does unmounting a filesystem on the host sometimes leave the device busy even though the umount succeeded?
    Because a propagated copy of that mount still exists in another mount namespace and holds a reference to the superblock. The kernel only tears the filesystem down when the last mount of it goes away, so you have to find and unmount the peer — or let the namespace holding it exit.

saying these in an interview costs you the question

  • Says a new mount namespace is fully isolated by default
  • Thinks propagation is a namespace property rather than per-mount
  • Believes slave propagation works in both directions
  • Assumes unmounting on the host always frees the device
  • Confuses unbindable with read-only

context