skip to content

What role do AppArmor and SELinux play in confining containers, how do they differ from Linux capabilities and syscall filtering, and how do you tell when one of them is causing a failure?

level: seniorimportance: nice to knowfreq 30%

answer

  1. MAC = which objects, not which syscalls
  2. AppArmor path-based, docker-default profile
  3. SELinux label-based: container_t + MCS category
  4. bind mounts: :z shared, :Z private
  5. denials land in host dmesg / ausearch -m AVC

basics

~20 s

They are mandatory access control systems that restrict which objects a process may touch, regardless of Unix permissions. AppArmor is path-based (Docker applies a docker-default profile); SELinux is label-based, giving each container a distinct MCS category. Denials appear in the host's kernel log, not as application errors.

solid answer

~50 s

Capabilities control *which privileged operations* a process may perform; seccomp controls *which syscalls* it may issue; **MAC (mandatory access control)** controls *which objects* it may act on — files, mounts, sockets, other processes — enforced by the kernel's LSM hooks and unalterable by the file owner. **AppArmor** (Debian/Ubuntu/SUSE) uses path-based rules. Docker generates and loads a `docker-default` profile: no writing to `/proc/sys` and most of `/proc`, no mounting, no ptrace of processes outside the container, restricted raw access. Override with `--security-opt apparmor=my-profile` or, for debugging only, `apparmor=unconfined`. **SELinux** (RHEL/Fedora) uses labels. Container processes run as `container_t` and may only touch files labelled `container_file_t`, and each container gets a unique **MCS** category pair so container A cannot read container B's files even though both are `container_t`. Bind mounts need relabelling: `:z` for a shared label, `:Z` for a private one. Diagnosis is distinctive: the container reports a plain `Permission denied` while the *host* logs the denial — `dmesg`/`journalctl` for AppArmor `DENIED`, `ausearch -m AVC` for SELinux.

code

bash · 13 lines
bash
# AppArmor: named profile, or unconfined for diagnosis only
docker run --security-opt apparmor=my-nginx-profile nginx
docker run --security-opt apparmor=unconfined nginx        # TEST ONLY

# SELinux: relabel a bind mount for exclusive container access
docker run -v /srv/data:/data:Z myapp:1.4.2

# where the denials actually appear (on the HOST)
dmesg | grep -i 'apparmor="DENIED"'
ausearch -m AVC -ts recent

# what confinement is this container running under?
docker inspect --format '{{.AppArmorProfile}} {{.HostConfig.SecurityOpt}}' ctr

go deeper

for a junior

Know that these are host-level security systems that can block a container's file access even when Unix permissions allow it, and that Docker applies a default profile.

for a middle

Distinguish path-based AppArmor from label-based SELinux, know the :z/:Z mount suffixes, and know denials are logged on the host.

for a senior

Triage a permission failure across DAC, capabilities, seccomp and MAC in a deterministic order, read AVC/DENIED records, and adjust labels or profiles rather than disabling confinement.

for a principal

Decide the fleet posture — which distro/LSM, how custom profiles are built into node images and version-controlled, permissive-mode rollout for new policy, and an explicit register of unconfined workloads.

## Where MAC sits in the stack A containerised process passes several independent gates: 1. **Discretionary access control (DAC)** — the classic uid/gid and file mode checks. "Discretionary" because the file's owner can change them. 2. **Capabilities** — is this process permitted to perform this *privileged operation class*? 3. **seccomp** — is this *syscall number* allowed at all? 4. **Mandatory access control (MAC)** — does policy allow *this subject* to act on *this object*? Set by an administrator, and not overridable by the file owner or by root. MAC is the only one of the four that reasons about the object being touched. seccomp cannot (it cannot dereference pointers to read a path); capabilities are coarse (`CAP_DAC_OVERRIDE` is all-or-nothing). So MAC is what expresses "this container may read its own volume and nothing else on the host". ## AppArmor AppArmor policies are path-based and attached to executables or applied by the runtime. Docker generates a profile named `docker-default` and applies it to every container on systems where AppArmor is active. Its restrictions include: - Deny writes to `/proc/sys/**` (blocking the `core_pattern` escalation trick) and to most `/sys` paths. - Deny `mount` and `umount` operations. - Deny `ptrace` and signals across the container boundary. - Deny raw access to kernel interfaces such as `/proc/kcore`, `/proc/sysrq-trigger`. Custom profiles are written in AppArmor's own language, loaded on the *host* with `apparmor_parser -r`, and selected per container with `--security-opt apparmor=profile-name`. That host-side load step is the operational catch: the profile must exist on every node before a workload referencing it starts, otherwise the container fails to start. Because the rules are path-based, an AppArmor profile is readable and easy to reason about — and brittle when paths change, e.g. after a bind-mount is moved. ## SELinux SELinux labels every process and every object with a context of the form `user:role:type:level`. The enforcement rule is type-based: policy states which *source types* may perform which operations on which *target types*. For containers the key types are `container_t` (the process) and `container_file_t` (files it may access). A host file labelled `etc_t` or `admin_home_t` is simply not accessible to `container_t`, no matter what the file mode says or whether the process is root — which is precisely why SELinux blunts many escape and host-file-read scenarios. On top of type enforcement, containers use **MCS** (multi-category security): each container is assigned a unique category pair such as `s0:c123,c456`, and its files are labelled with the same pair. Two containers both running as `container_t` therefore still cannot read each other's files. This is the mechanism that turns SELinux from "containers vs host" into "container vs container" isolation. **Bind mounts and relabelling.** A host directory has whatever label it had; a container cannot read it. Docker's volume suffixes solve this: - `-v /host/data:/data:z` relabels the content `container_file_t` with a *shared* category, so several containers can use it. - `-v /host/data:/data:Z` relabels it with the container's *private* category — exclusive access. Use `:Z` unless sharing is intended, and be careful: relabelling a directory like `/home` or `/var` recursively can break the host badly. `--security-opt label=disable` turns SELinux confinement off for one container; treat that like `seccomp=unconfined`. ## Diagnosing MAC denials The signature is a mismatch between symptom and evidence. Inside the container you see `Permission denied` on a file that appears world-readable with the correct uid — DAC clearly permits it. The denial is logged on the **host**: - AppArmor: `dmesg | grep -i apparmor` or `journalctl -k`, lines containing `apparmor="DENIED"` with `operation=`, `profile=docker-default`, `name=<path>`. - SELinux: `ausearch -m AVC -ts recent` or `journalctl -t setroubleshoot`; `audit2allow` can suggest a policy delta (review it, do not paste it blindly). A quick triage: run the workload once with `--security-opt apparmor=unconfined` or `--security-opt label=disable` **in a test environment**. If the failure disappears, MAC is the layer; if not, look at capabilities (EPERM on a privileged operation) or seccomp (EPERM/SIGSYS on a syscall). Because all three can surface as `EPERM`, having a deterministic order to test them is what makes this tractable under incident pressure. SELinux also offers `permissive` mode — policy evaluated and denials logged but not enforced — which is the right way to gather a complete denial set before writing policy, and much better than disabling it. ## Practical posture Keep the runtime default profile on (`docker-default` / `container_t`) for everything. Write custom profiles only where a workload genuinely needs a narrower or slightly wider policy, keep them in version control, and distribute them with the node image rather than by hand. Never disable MAC as a fix for a startup failure — the denial message names the exact object, which is usually enough to correct a mount label or a path. And record which workloads run unconfined, because that list is a standing risk register, not a configuration detail.

  • A container gets 'Permission denied' reading a bind-mounted host directory that is world-readable and owned correctly. What do you check?
    On an SELinux host this is the classic missing-label case: the host directory keeps its original type, and `container_t` has no rule permitting access to it. Confirm with `ls -Z` on the host and `ausearch -m AVC`, then remount with `:Z` for exclusive access or `:z` if several containers share it. On an AppArmor host, look for an `apparmor="DENIED"` line in `dmesg` naming the path, and adjust the profile rather than running unconfined.
  • Why is AppArmor's path-based model both easier and more fragile than SELinux's label-based model?
    Path rules are readable — an engineer can see exactly which directories are permitted — so profiles are quicker to write and review. But identity is tied to the name rather than the object, so bind mounts, symlinks, and moved paths can change what a rule covers. SELinux labels travel with the object regardless of where it is mounted, which is more robust but far harder to reason about and requires relabelling discipline on every mount.

Capabilities decide whether you are allowed to use a crowbar; seccomp decides whether you may pick up a crowbar at all; MAC decides which doors in the building the crowbar will ever be allowed near.

saying these in an interview costs you the question

  • Describing AppArmor or SELinux as a syscall filter, confusing them with seccomp
  • Expecting MAC denials to appear in the container's own logs rather than the host kernel log
  • Disabling SELinux (or using label=disable) as the standard fix for a mount permission error
  • Relabelling broad host directories with :Z, damaging host services that rely on the original labels
  • Assuming a MAC profile is portable — forgetting a custom AppArmor profile must be loaded on every node first

context