A process on a Linux host gets its own mount namespace, yet it is not automatically a sealed copy of the host's mount table. Explain mount propagation — shared, private, slave and unbindable — and which type gives one-way visibility from the host into the namespace.
answer
- a copy with a relationship, not a snapshot
- events travel between peer groups
- one direction only is the useful case
- mountinfo tags name the group
- systemd leaves the root shared
basics
~20 sMount propagation decides whether mount and unmount events travel between mount namespaces that share a peer group. Shared propagates both ways, private neither way, slave one way only — in from the master — and unbindable additionally refuses to be used as a bind source. One-way is slave.
solid answer
~50 sA new mount namespace starts as a *copy* of the parent's mount table, and each entry keeps its propagation type, so events can keep flowing between them. `shared` puts mounts in a peer group: a mount or unmount under that point in either namespace shows up in the other. `private` isolates it — events go neither way. `slave` makes the mount receive events from its master peer group but send none back, which is the one-way behaviour: a disk mounted on the host appears inside, while mounts made inside stay invisible outside. `unbindable` is private plus a refusal to serve as the source of a bind mount, which stops recursive bind mounts from exploding. You set them with `mount --make-shared`, `--make-private`, `--make-slave`, `--make-unbindable` and their recursive `--make-r*` forms, and read the current state in the optional fields of `/proc/self/mountinfo` (`shared:N`, `master:N`) or via `findmnt -o TARGET,PROPAGATION`. On a systemd host `/` is shared by default, which is why sandboxes explicitly re-mark their tree.
code
bash · 8 lines# what propagation is in effect right now
findmnt -o TARGET,PROPAGATION /
# host -> namespace only: receives host mounts, leaks nothing back
mount --make-rslave /mnt/hostdata
# raw view: 'shared:N' = peer group, 'master:N' with no shared = slave
grep hostdata /proc/self/mountinfogo deeper
Know that a mount namespace gives a process its own mount table and that the table starts as a copy of the parent's rather than as an empty one.
Name the four propagation types and what each does to mount and unmount events, and show where to read the current type — findmnt's PROPAGATION column or the mountinfo tags.
Diagnose from symptoms: a mount leaking out means shared was left in place; a host disk never appearing inside means the tree was made private instead of slave. Know the order in which to re-mark a new namespace.
Own the convention for how sandboxes on your fleet mark their trees, and the tradeoff between a strictly private tree and one that can receive host-side mounts after startup.
## Why propagation exists at all When `CLONE_NEWNS` gives a process a new mount namespace, the kernel copies the current mount table into it. If that were the end of the story, the copy would be frozen: plug in a disk on the host after the copy was taken and the isolated process could never see it, which breaks any workflow where a supervisor wants to hand a filesystem to an already-running sandbox. Shared subtrees solve this. Each mount belongs to a *peer group*, and its propagation type determines how mount and unmount events move between members of that group. The types are per-mount, not per-namespace, so one namespace can be shared in one subtree and private in another. ## The four types **shared (`MS_SHARED`)** — the mount is a member of a peer group and events propagate in **both** directions. Mount something under it in one namespace and the mount appears at the corresponding point in every peer. This is the default on a systemd host for `/`, deliberately, so that things like removable media and network mounts reach services that were started earlier. **private (`MS_PRIVATE`)** — the mount belongs to no peer group. Nothing propagates in, nothing propagates out. This is the safe default for a sandbox that must not leak, and it is what `unshare --mount` applies by default when it creates the namespace. **slave (`MS_SLAVE`)** — the mount has a master peer group and receives events from it, but its own mounts and unmounts do not travel back. This is the asymmetric case: the host can hand filesystems in, and the namespace cannot push anything out. It is the setting sandbox tooling reaches for when it bind-mounts a host directory into an isolated tree. **unbindable (`MS_UNBINDABLE`)** — private, and additionally the mount may not be used as the *source* of a bind mount. Its purpose is to prevent the pathological growth you get when you recursively bind-mount a tree that contains a shared mount of itself; marking such points unbindable stops the recursion. Each type has a recursive form applied with the `r` variants, which set the type on the mount and everything under it. ## Reading the current state The authoritative source is `/proc/self/mountinfo`. Each line carries optional fields before the separator: ```bash grep ' / ' /proc/self/mountinfo # 25 1 259:2 / / rw,relatime shared:1 - ext4 /dev/nvme0n1p2 rw ``` `shared:1` means this mount is in peer group 1. A line carrying `master:1` and no `shared:` tag is a **slave** of peer group 1. A line carrying both is a slave that is itself shared with further peers. A line with neither is private. `unbindable` appears as its own tag. The friendlier view is `findmnt -o TARGET,PROPAGATION`, which prints the type per mount point directly. ## The failure modes you get asked about **"My container-style sandbox leaked a mount onto the host."** The tree was left shared. Because the copy inherits propagation, a mount made inside a shared subtree propagates straight back out. Marking the tree `rslave` or `rprivate` after creating the namespace is the fix. **"A filesystem mounted on the host after startup is invisible inside."** The tree was made fully private. Nothing propagates in. If the sandbox is supposed to see later host mounts under a particular path, that path needs to be slave rather than private. **"Unmounting on the host did not free the device."** A propagated copy of the mount still exists in another namespace, holding a reference. Until every peer's copy is gone, the filesystem stays busy — which is also why an ordinary `umount` on the host can appear to succeed while the block device is still in use. ## Setting it deliberately ```bash # make a subtree strictly one-way: host -> namespace mount --make-rslave /mnt/hostdata # fully isolate a subtree in both directions mount --make-rprivate /srv/sandbox ``` The order matters: propagation is set on mounts as they exist now, so a sandbox typically creates its mount namespace, immediately re-marks its whole tree, and only then performs the mounts it wants to keep to itself. Doing it the other way round means the first mounts have already propagated. ## The interview framing The question underneath all of this is whether you understand that a mount namespace is a *copy with a relationship*, not a snapshot. Someone who says "a new mount namespace is completely isolated" has not hit the leak, and someone who says "it is a frozen copy" has not hit the invisible-disk case. Both symptoms come from the same mechanism seen from opposite ends.
- A sandbox made a bind mount inside its own mount namespace and it appeared on the host. What happened?The mount point was still `shared`, inherited from the parent namespace when the table was copied. Shared propagation is bidirectional, so the new mount propagated back into the peer group. Re-marking the tree with `mount --make-rslave` or `--make-rprivate` immediately after creating the namespace prevents it.
- How do you tell a slave mount from a private one by reading /proc/self/mountinfo?By the optional fields. A slave carries a `master:N` tag naming the peer group it receives from; a private mount carries no propagation tag at all. If a line shows both `shared:N` and `master:M`, it is a slave that is itself shared with its own peers.
- Why does unmounting a filesystem on the host sometimes leave the device busy even though the umount succeeded?Because a propagated copy of that mount still exists in another mount namespace and holds a reference to the superblock. The kernel only tears the filesystem down when the last mount of it goes away, so you have to find and unmount the peer — or let the namespace holding it exit.
saying these in an interview costs you the question
- Says a new mount namespace is fully isolated by default
- Thinks propagation is a namespace property rather than per-mount
- Believes slave propagation works in both directions
- Assumes unmounting on the host always frees the device
- Confuses unbindable with read-only