skip to content

Hardening & Confinement

Hardening a container against the host it shares: dropped capabilities, seccomp and AppArmor confinement, rootless daemons, secrets kept out of layers, and the blast radius of a mounted docker.sock. Probed because a default docker run is over-privileged.

part ofDockeroverview, primer and where to startread it →
on this pageshow

questions

page 1 of 2

What does the `--privileged` flag actually grant a Docker container, and why is running with it considered equivalent to giving away root on the host?

level: juniorimportance: must knowfreq 60%

answer

  1. all caps + seccomp unconfined + AppArmor off
  2. cgroup device controller wide open
  3. /sys read-write
  4. mount /dev/sda1 → host root
  5. fix: --cap-add / --device / targeted profile

basics

~20 s

--privileged grants all Linux capabilities, disables the default seccomp and AppArmor/SELinux confinement, lifts the cgroup device restrictions so all host devices are usable, and mounts /sys writable. With that, a process can mount the host disk or load kernel modules — effectively host root.

solid answer

~50 s

`--privileged` is not "a bit more permission"; it removes essentially every isolation layer except namespaces: - **All capabilities** granted, including `CAP_SYS_ADMIN` and `CAP_SYS_MODULE`. - **Seccomp set to unconfined**, so blocked syscalls like `mount`, `keyctl`, and `bpf` become available. - **AppArmor/SELinux confinement dropped** to unconfined. - **The cgroup device controller allows all devices**, and `/dev` is populated from the host — so the raw host disk is readable and writable. - **`/sys` is mounted read-write.** Escape is then trivial and well documented: `mount /dev/sda1 /mnt` and edit anything on the host filesystem, or `insmod` a kernel module. Note filesystem and PID namespaces are still applied, so it is not literally "no isolation" — but every mechanism that would *stop* you crossing them is gone. The fix is to grant only the specific thing needed: `--cap-add` for one capability, `--device` for one device, a `--mount` for one path, or a targeted seccomp profile.

code

bash · 13 lines
bash
# DANGEROUS: this is host root
docker run --privileged -it debian bash
#   inside: fdisk -l ; mount /dev/sda1 /mnt ; ls /mnt/root

# scoped alternatives
docker run --cap-drop ALL --cap-add NET_BIND_SERVICE nginx        # bind :80
docker run --device /dev/ttyUSB0 my-serial-app                    # one device
docker run --cap-drop ALL --cap-add NET_RAW --cap-add NET_ADMIN \
           nicolaka/netshoot tcpdump -i eth0                      # packet capture

# audit what is running privileged
docker ps -q | xargs -r docker inspect \
  --format '{{.Name}} privileged={{.HostConfig.Privileged}}'

go deeper

for a junior

Know that --privileged removes the container's protections and is effectively host root; avoid it and ask what specific permission is needed.

for a middle

Enumerate the four things it turns off — capabilities, seccomp, MAC, device cgroup — and give the scoped --cap-add/--device alternatives.

for a senior

Demonstrate a concrete escape path, show how to diagnose the actual denial with dmesg/strace, and describe auditing and admission policy to keep it out of production.

for a principal

Treat privileged workloads as a named risk acceptance with an owner and an exception list, and design the platform so legitimate node agents are the only holders.

## What container isolation is made of A container is not one mechanism. It is a stack: 1. **Namespaces** — separate views of PIDs, mounts, network, users, IPC, UTS. 2. **Capabilities** — root's powers split into ~40 discrete privileges; Docker grants a small default subset. 3. **cgroups** — resource limits, and importantly the *device* controller, which decides which device nodes may be opened. 4. **seccomp** — a syscall filter; Docker applies a default profile blocking a few dozen dangerous syscalls. 5. **MAC (AppArmor or SELinux)** — a mandatory access control profile restricting file, mount and ptrace operations regardless of Unix permissions. 6. **Filesystem shape** — masked and read-only paths under `/proc` and `/sys`. `--privileged` turns off items 2 through 6. Only namespaces remain, and namespaces alone are a *view*, not a wall — with the right capabilities you can reach around them. ## Concretely, what changes **All capabilities.** The default set is fourteen; privileged gives the full bounding set. `CAP_SYS_ADMIN` alone is often called "the new root": it permits `mount`, namespace manipulation, and much more. `CAP_SYS_MODULE` permits loading kernel modules — code running in the kernel, outside every container boundary by construction. `CAP_SYS_RAWIO` permits raw port and memory access. **Seccomp unconfined.** The default profile denies syscalls that have no business in a normal workload — `mount`, `umount2`, `reboot`, `kexec_load`, `keyctl`, `add_key`, `bpf`, `perf_event_open` under some configs. Privileged removes the filter entirely. **MAC unconfined.** The `docker-default` AppArmor profile denies writes to `/proc/sys`, mounting, and ptrace of processes outside the container. SELinux gives each container a distinct MCS category so it cannot touch another container's labelled files. Privileged disables both. **Devices.** The cgroup device controller normally allows a tiny whitelist (`/dev/null`, `/dev/zero`, `/dev/urandom`, ttys). Privileged allows everything, and the host's `/dev` entries are available. So `/dev/sda`, `/dev/nvme0n1`, `/dev/mem` are all reachable. **`/sys` writable.** Kernel tunables and subsystem controls become writable. ## Why that is host root Three well-known escapes, each a couple of commands: - **Mount the host disk.** `mount /dev/sda1 /mnt` then write an SSH key into `/mnt/root/.ssh/authorized_keys` or drop a file into the host's cron directory. The mount namespace does not stop you because you have `CAP_SYS_ADMIN` and the device is permitted. - **Load a kernel module.** With `CAP_SYS_MODULE`, `insmod evil.ko` executes attacker code in ring 0. - **Abuse a kernel handler.** Classic variants write to `/proc/sys/kernel/core_pattern` (which pipes crash handling to a program run by the host kernel) or, on cgroup v1, to a `release_agent` file. Both cause the host to execute a chosen binary as root. Because of this, "privileged container" and "root on the node" should be treated as the same statement in any threat model. If an attacker gets code execution inside a privileged container — through an application vulnerability, a poisoned dependency, a build step — they own the machine and every other container on it. ## What people actually needed instead `--privileged` usually appears because someone hit a permission error and reached for the biggest hammer. The disciplined alternatives: - Need to bind port 80 as non-root? `--cap-add NET_BIND_SERVICE`, or simply listen on a high port and map it. - Need one device (a GPU, a serial port, a USB dongle)? `--device /dev/ttyUSB0`. - Need to run `tcpdump`? `--cap-add NET_ADMIN` and/or `NET_RAW`, nothing else. - Need to `mount` a FUSE filesystem? `--device /dev/fuse` plus `--cap-add SYS_ADMIN` — still bad, but scoped, and better with a custom seccomp profile. - Need to debug with `strace`/`gdb`? `--cap-add SYS_PTRACE` and `--security-opt seccomp=unconfined` on a *throwaway* debug container, never in the deployed spec. The method is always the same: reproduce the failure, identify the exact denied operation (`dmesg` for AppArmor/SELinux denials, `strace` for `EPERM` syscalls), grant that one thing, and re-test. ## Detecting and preventing it `docker inspect --format '{{.HostConfig.Privileged}}' <ctr>` reports it per container. In an orchestrator, policy admission can reject privileged workloads outright, and that should be the default posture with a narrow, reviewed exception list — typically only node-level agents such as storage or CNI plugins. Treat every exception as a standing risk acceptance with an owner, not a checkbox.

  • Is a privileged container still isolated by namespaces?
    Yes, the namespaces are still created — the container has its own PID, mount and network views unless you also pass `--pid=host` or `--network=host`. But namespaces only change what a process *sees*; with all capabilities, no seccomp filter and full device access, the process can mount the host filesystem or load a kernel module and step outside that view. Isolation that you can dismantle from the inside is not a security boundary.
  • Someone says they need `--privileged` to run Docker inside their CI container. What do you propose?
    First check whether they need a Docker *daemon* at all — most CI image builds can use a rootless, daemonless builder such as BuildKit in rootless mode, Buildah, or Kaniko, none of which need privileged. If a real nested daemon is required, isolate it: a dedicated node pool, a short-lived VM or a sandboxed runtime, treated as a build-farm boundary. Sharing the host's Docker socket is not a safer substitute — that is a different route to the same host-root outcome.

Namespaces are the walls of the room; capabilities, seccomp and AppArmor are the locks. --privileged leaves the walls up and removes every lock, then hands you the keys to the building's plant room.

saying these in an interview costs you the question

  • Describing `--privileged` as just "adds all capabilities" and omitting seccomp, MAC and device access
  • Claiming a privileged container is still safely isolated because namespaces remain
  • Believing running as a non-root user inside the container neutralises `--privileged`
  • Using `--privileged` as a first response to any permission error instead of finding the denied operation
  • Treating sharing the Docker socket as the safe alternative to privileged

context

open as a page

What does it mean to 'scan a container image for vulnerabilities', and how does a tool like Trivy or Grype actually find CVEs in an image?

level: juniorimportance: must knowfreq 55%

basics

~20 s

The scanner inventories everything installed in the image (OS packages per layer plus app dependencies), determines each component's name and version, then matches those against vulnerability databases (like the NVD and distro advisories). Matches are reported as CVEs with severities. It is metadata matching, not runtime analysis.

open as a page

A teammate asks to be added to the `docker` group on a shared Linux build server so they can run builds without sudo. Why do security reviewers treat that as equivalent to granting passwordless root, and what would you propose instead?

level: juniorimportance: must knowfreq 62%

basics

~20 s

The Docker daemon runs as root and docker group members can talk to its socket. One container that bind-mounts the host filesystem gives them full root. Offer rootless Docker, Podman, or a narrow sudo rule instead.

open as a page

A Dockerfile receives a private registry token through `ARG NPM_TOKEN` and CI passes it with `docker build --build-arg NPM_TOKEN=...`. Explain why that token can still be recovered from the published image, and how someone would extract it.

level: juniorimportance: must knowfreq 64%

basics

~20 s

Build arguments are recorded in the image's build history, and any ENV keeps the value in image config. Anyone who pulls the image can read them with docker history or docker inspect, or by unpacking the saved image — and deleting a file later does not remove earlier layers.

open as a page

Container hardening guides say every image should end with a USER instruction instead of leaving the process as root. What actually changes on the host when you do that, and what does it not protect you from?

level: juniorimportance: must knowfreq 72%

basics

~20 s

By default the container process runs as UID 0, the same UID 0 the host kernel calls root. A USER instruction runs it as an unprivileged UID, so escapes, bind mounts and host files get no root rights. It is defence in depth, not isolation.

open as a page

Why is `RUN apt-get upgrade` inside a Dockerfile treated as an antipattern for patching CVEs?

level: middleimportance: must knowfreq 58%

basics

~20 s

Upgrading OS packages inside a Dockerfile makes the image depend on when it was built, its cached instruction stops re-running so it patches nothing, and the superseded files stay in the layers below. Rebuild on a refreshed base instead.

open as a page

Explain Linux capabilities in the context of containers: what does Docker grant by default, and how do `--cap-drop ALL` and `--cap-add` change what a containerised process can do?

level: middleimportance: must knowfreq 55%

basics

~20 s

Capabilities split root's power into ~40 separate privileges. Docker grants a default subset of 14 (CHOWN, SETUID, NET_RAW, NET_BIND_SERVICE and others) rather than all. --cap-drop ALL removes them all; --cap-add NET_BIND_SERVICE then adds back only what the workload needs.

open as a page

Why does bind-mounting the host root (-v /:/host) into a Docker container hand over the whole host?

level: middleimportance: must knowfreq 55%

basics

~20 s

Bind-mounting the host root gives the container's root user read-write access to every host file. It can append an SSH key, edit /etc/passwd or sudoers, or drop a cron/systemd unit -- arbitrary code execution as root on the host, with no kernel exploit and no --privileged needed.

open as a page

What does `docker run --read-only` actually make read-only, and why do most containers then need `--tmpfs`?

level: middleimportance: must knowfreq 58%

basics

~20 s

--read-only mounts the container's root filesystem read-only, so every write to a path that came from the image fails with EROFS. Mounts are unaffected: volumes, bind mounts and tmpfs stay writable, so scratch paths such as /tmp and /run need an explicit --tmpfs.

open as a page

Explain what Docker's rootless mode changes about how the daemon and its containers run, including the role of `/etc/subuid` and the `newuidmap` helper, and what ownership a host file gets when a rootless container writes it as its own root user.

level: middleimportance: must knowfreq 44%

basics

~20 s

Rootless mode runs the daemon as an ordinary user inside a user namespace. Ranges in /etc/subuid and /etc/subgid, applied by the setuid helper newuidmap, map container uid 0 to a high unprivileged host uid — which is what host files end up owned by.

open as a page

How does BuildKit's `RUN --mount=type=secret` differ from passing a value with `--build-arg`, what exactly does the build step see, and what still leaks if you use it carelessly?

level: middleimportance: must knowfreq 46%

basics

~20 s

A secret mount exposes the value as a temporary in-memory file, mounted only for that one RUN instruction. It is not written to any layer and its value never appears in the image history. It still leaks if the command copies it into the filesystem or prints it.

open as a page

Why is giving a container access to /var/run/docker.sock considered equivalent to giving it root on the host, and how would an attacker turn that access into a full host takeover?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The daemon runs as root and its API can create a container that bind-mounts the host's / and runs privileged. From inside that container an attacker chroots into the host filesystem or writes to it, gaining full root. The socket is a root-equivalent control plane, not a limited API.

open as a page

You're wiring image scanning into a CI pipeline. How do you decide which findings should fail the build, and how do you handle CVEs that have no fix yet or that you've assessed as not applicable?

level: seniorimportance: must knowfreq 45%

basics

~20 s

Gate on severity thresholds plus fixability: typically fail on CRITICAL/HIGH that have a fix available, and don't block on unfixable ones (report them). Handle noise with a reviewed, time-boxed allow-list (e.g. .trivyignore / VEX) that documents why each CVE is ignored and expires, so exceptions get revisited rather than becoming permanent.

open as a page

For delivering a database password to a running container, compare an environment variable set with `docker run -e DB_PASSWORD=...` against a file mounted into the container. What leakage paths does each have, and which would you choose?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Environment values show up in docker inspect, are inherited by every child process, are readable in /proc/<pid>/environ, and get captured by crash and error reporters. A file can be permission-scoped, stays out of inspect output, and can be rotated in place. Prefer a mounted file, ideally on tmpfs.

open as a page

What does `docker build --pull` do, and why can a nightly rebuild still ship an outdated base image?

level: juniorimportance: should knowfreq 42%

basics

~20 s

docker build --pull makes the builder re-resolve every FROM reference against the registry instead of reusing the base image already in the host's local image store. Without it a rebuild can inherit a months-old base, and its CVEs, forever.

open as a page

What is the Unix socket at /var/run/docker.sock, and what does a container gain when you bind-mount that file into it?

level: juniorimportance: should knowfreq 45%

basics

~20 s

It is the Unix socket the root Docker daemon listens on for its REST API. Bind-mounting it lets a container send API calls to the host daemon (run, build, exec, mount, inspect), so it can control every container on that host.

open as a page

What does the Docker run option `--security-opt no-new-privileges` do at the kernel level, and which attack does it stop?

level: middleimportance: should knowfreq 35%

basics

~20 s

It sets the kernel's no_new_privs bit on the container process, which is inherited by all children and cannot be unset. With it, executing a setuid binary or a file with file capabilities grants no extra privilege, so a compromised unprivileged process cannot escalate that way.

open as a page

Some setups expose the Docker Engine API over a TCP port instead of the local Unix socket. What is the danger of binding it to tcp://0.0.0.0:2375, and how do you expose it remotely without opening the host to anyone?

level: middleimportance: should knowfreq 30%

basics

~20 s

Plain TCP on port 2375 is unauthenticated and unencrypted; anyone who can reach it gets root-equivalent daemon control, and internet scanners actively hunt for it. Expose remotely only over port 2376 with TLS mutual authentication (verify client certs), or better, tunnel over SSH (DOCKER_HOST=ssh://).

open as a page

What do Docker's --pid=host, --net=host and --ipc=host flags each expose to a container, and why are they escape risks?

level: middleimportance: should knowfreq 50%

basics

~20 s

Each drops one isolation boundary: --pid=host reveals and lets the container signal every host process, --net=host puts it directly on the host's network stack, and --ipc=host shares host shared memory. All three erode container-host isolation and need no --privileged to be dangerous.

open as a page

How do `docker run --pids-limit` and `--ulimit nproc` differ as defences against a fork bomb in a container?

level: middleimportance: should knowfreq 38%

basics

~20 s

--pids-limit sets the container cgroup's pids.max, capping every process and thread inside that one container; further clone() calls fail with EAGAIN. --ulimit nproc sets a per-UID kernel limit counted across the whole host, so it is the weaker and more surprising control.

open as a page

What is an SBOM (Software Bill of Materials), and how does SBOM-based scanning differ from pointing a scanner directly at a running or stored image?

level: middleimportance: should knowfreq 38%

basics

~20 s

An SBOM is a machine-readable inventory of every component and version in an image (formats: SPDX, CycloneDX). You generate it once at build time; then you can re-scan that list against fresh CVE feeds anytime without the image, and detect newly-disclosed CVEs in components you already shipped. Direct scanning re-inventories the image each run.

open as a page

An image now ends with `USER 10001`, and the application fails to start: it binds TCP port 80 and writes to /var/run and its own cache directory. What are your options to keep the process unprivileged?

level: middleimportance: should knowfreq 46%

basics

~20 s

Move the listener to a high port and publish 80 on the host, or grant CAP_NET_BIND_SERVICE. For writes, chown the needed directories to the UID at build time, redirect state to /tmp, and mount tmpfs or a named volume for anything else.

open as a page

A colleague enables the Docker daemon's --userns-remap option. Explain what user-namespace remapping does to the UIDs a container sees versus the UIDs the host kernel sees, and where the mapping ranges come from.

level: middleimportance: should knowfreq 44%

basics

~20 s

The daemon puts containers in their own user namespace. Inside, the process still sees UID 0; the kernel maps that to an unprivileged host UID from a subordinate range in /etc/subuid and /etc/subgid, typically belonging to a user named dockremap.

open as a page

A Docker host keeps running the old image after you pushed a patched rebuild to the same tag — why?

level: seniorimportance: should knowfreq 46%

basics

~20 s

A tag is a mutable pointer, and the host already holds an image under that name, so docker run, docker compose up and docker restart never ask the registry. Only a pull plus a container recreate lands the rebuild.

open as a page

How does Docker's default seccomp profile work, what kinds of syscalls does it block, and how would you handle an application that fails because of it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

seccomp is a kernel syscall filter. Docker applies a default JSON profile that allows most syscalls and returns EPERM for a few dozen dangerous ones (mount, reboot, kexec_load, keyctl, add_key, and namespace-creating clone flags). If an app breaks, identify the denied syscall and write a narrow custom profile — never use seccomp=unconfined in production.

open as a page

In CI, containerized build agents often need to build and run Docker images. Compare the two common approaches: Docker-in-Docker (DinD) versus bind-mounting the host's Docker socket into the agent. What are the operational and security tradeoffs?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Socket-mount reuses the host daemon: fast and cache-sharing, but the job gets host root and can see sibling containers, so it's unsafe for untrusted/shared runners. DinD runs a private daemon per job (needs privileged): stronger isolation and clean caches, but privileged is its own risk and caches are cold. Rootless builders (BuildKit/Kaniko/Buildah) avoid both.

open as a page

A monitoring agent (say cAdvisor or a Traefik reverse proxy) legitimately needs to read container metadata from the Docker daemon, but you don't want to hand it root-equivalent control. How does a Docker socket proxy solve this, and what would you allow-list?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Put a small proxy (e.g. tecnativa/docker-socket-proxy) in front of the real socket. It exposes the Engine API over a network port but allow-lists only the endpoints the client needs, usually read-only GETs like /containers and /events, and denies container create, exec, and image mutations. The agent connects to the proxy, never to the raw socket.

open as a page

How can a privileged Docker container with CAP_SYS_ADMIN break out to the host using the cgroup v1 release_agent?

level: seniorimportance: should knowfreq 40%

basics

~20 s

With CAP_SYS_ADMIN (which --privileged grants) on a cgroup v1 host, the container mounts a cgroup hierarchy, sets its release_agent to a script on a host-visible path, and empties a child cgroup. The kernel then runs that script as root in the host namespace -- a full breakout, needing no CVE.

open as a page

After adding `--read-only`, a Node.js queue worker dies with EROFS. How do you find every path it writes?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Run the container once without --read-only, exercise it fully, then read docker diff to list every file it changed in its writable layer. Classify each path: sized tmpfs for scratch, a named volume for durable state, and move one-time startup writes into the image.

open as a page

Your team scans images in CI and the last build passed clean, yet weeks later the same deployed image shows new critical CVEs. Why does this happen, and what base-image patch-and-rebuild strategy keeps images current?

level: seniorimportance: should knowfreq 40%

basics

~20 s

New CVEs are disclosed continuously against packages already baked into your image, and a passing build only reflects the feed on build day. Fix it by rebuilding regularly (scheduled/nightly), pinning base images by digest but bumping them on a cadence, using minimal/distroless bases to shrink the surface, and re-scanning deployed images so new CVEs trigger a rebuild.

open as a page

showing 1–30 of 45