skip to content

Hardening & Confinement

Hardening a container against the host it shares: dropped capabilities, seccomp and AppArmor confinement, rootless daemons, secrets kept out of layers, and the blast radius of a mounted docker.sock. Probed because a default docker run is over-privileged.

part ofDockeroverview, primer and where to startread it →
on this pageshow

questions

page 2 of 2

After moving a CI runner to rootless Docker, a job can no longer bind port 80, `--memory` limits appear to be ignored, and throughput on published ports has dropped noticeably. Walk through the known limitations of a rootless daemon and how you would address each of these.

level: seniorimportance: should knowfreq 38%

basics

~20 s

Rootless daemons cannot bind ports below 1024, need cgroup v2 with systemd delegation for resource limits, and route traffic through a userspace network stack that costs throughput. Fix with the unprivileged-port sysctl, Delegate=yes, and a suitable port driver or host proxy.

open as a page

You discover that an image already pushed to a shared registry contains a live cloud credential inside one of its layers. What is your response, and why is deleting the file in a follow-up layer, or squashing the image, not a fix on its own?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Rotate the credential first — treat it as compromised. Layers are immutable and content-addressed, so a later deletion only hides the file, and squashing produces a new digest while the old one stays pullable from the registry, mirrors, caches and node image stores until it is deleted and garbage-collected.

open as a page

Which container-escape scenarios does daemon-level user-namespace remapping actually mitigate, and which Docker features stop working once it is enabled?

level: seniorimportance: should knowfreq 30%

basics

~20 s

It blunts abuses of in-container root against the host: writing root-owned bind-mounted files, host device access, host-privileged kernel operations. It does not stop kernel exploits, a mounted daemon socket, or container-to-container access, and it is incompatible with host PID/network sharing and privileged mode.

open as a page

A service was changed to run as UID 10001 inside the container, and the Docker daemon has user-namespace remapping enabled. Writes to its bind-mounted data directory now fail with permission denied. How do you reason about the ownership and fix it?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Work out the host-side UID: remap offset plus the in-container UID (165536 + 10001 = 175537), then chown the host directory to it. Named volumes avoid this because Docker seeds their ownership from the image.

open as a page

A platform team proposes rebuilding every Docker image in the estate weekly against refreshed bases — what does that cost, and how do you make it safe?

level: principalimportance: should knowfreq 40%

basics

~20 s

A rebuild produces new bytes, so unless application inputs are pinned the base refresh smuggles in unreviewed upgrades. Make the base the only floating input, gate on each service's tests, stagger the rollout, and remember rebuilding is not deploying.

open as a page

You are setting a container runtime-confinement baseline for an organisation. Which restrictions do you make mandatory by default, how do you roll it out without breaking existing workloads, and how do you handle the workloads that genuinely need more privilege?

level: principalimportance: should knowfreq 30%

basics

~20 s

Baseline: drop all capabilities, no-new-privileges, runtime default seccomp, MAC profile on, read-only root filesystem, no privileged and no host namespaces. Roll out in audit/permissive mode first, fix the failures found, then enforce. Exceptions get a named owner, a scope, and an expiry.

open as a page

A GPU inference service genuinely needs host device access; how do you contain the blast radius rather than reaching for --privileged?

level: principalimportance: should knowfreq 38%

basics

~20 s

Grant the minimum, not the superset: expose only the specific device and the few capabilities the workload needs, never blanket --privileged. Then assume breakout is possible and shrink what a breakout reaches -- a dedicated low-trust host or node pool, no other tenants or secrets, a minimal cloud identity, network segmentation, and a patched runtime.

open as a page

How do you choose `--pids-limit` and file-descriptor ulimit defaults for a whole container fleet?

level: principalimportance: should knowfreq 28%

basics

~20 s

Treat them as host-resource protection, not per-service tuning. Measure peak task and descriptor use across the estate, define two or three tiers rather than per-service numbers, set the tier as the platform default, and make a higher tier an explicit, reviewed choice.

open as a page

You are choosing the container runtime standard for a fleet of shared Linux build hosts and edge devices. Compare a rootful Docker daemon, rootless Docker, and Podman's daemonless model, and explain how you would decide.

level: principalimportance: should knowfreq 26%

basics

~20 s

Rootful Docker centralizes root in one long-lived daemon and socket. Rootless Docker keeps the daemon but confines it to a user. Podman drops the daemon entirely: containers are children of the invoking user, so identity and audit follow the human. Decide on isolation, auditability and feature needs.

open as a page

How does pinning a Dockerfile's `FROM` to an image digest change a multi-platform build?

level: middleimportance: nice to knowfreq 22%

basics

~20 s

A multi-platform tag resolves to an index listing one manifest per architecture. Pinning the index digest keeps every platform buildable; pinning one architecture's manifest digest hard-codes that architecture, and a build requested for another platform finds no matching entry.

open as a page

Which `docker run` flag grants a container one host device without `--privileged`, and what do its rwm bits mean?

level: middleimportance: nice to knowfreq 24%

basics

~20 s

--device /dev/ttyUSB0:/dev/ttyUSB0:rwm creates that one node inside the container and adds a matching allow rule to its device cgroup. The trailing letters mean read, write and mknod; drop letters to narrow the grant, for example :r for read-only access.

open as a page

Trivy, Grype, and Docker Scout are all common image scanners. At a high level, how are they similar, and what practical factors would guide picking one (or more) for a team?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

All three do the same core job: inventory image components and match versions against CVE feeds, output findings, and can produce/consume SBOMs. Choice comes down to coverage breadth, integration (Trivy also scans IaC/secrets/misconfig; Scout integrates with Docker Hub/Desktop and shows remediation; Grype pairs with Syft), false-positive behaviour, licensing/cost, and where it runs (CLI/CI/registry).

open as a page

What does mounting `tmpfs` inside a container buy you for secret material, and how do Docker Compose and Swarm `secrets:` entries actually deliver a value into a container?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

A tmpfs mount is memory-backed: nothing written there reaches the container's writable layer or the host disk, and it vanishes when the container stops. Compose and Swarm secrets mount each value read-only at /run/secrets/<name> — Compose from a local file, Swarm from the encrypted raft store over mutual TLS.

open as a page

What role do AppArmor and SELinux play in confining containers, how do they differ from Linux capabilities and syscall filtering, and how do you tell when one of them is causing a failure?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

They are mandatory access control systems that restrict which objects a process may touch, regardless of Unix permissions. AppArmor is path-based (Docker applies a docker-default profile); SELinux is label-based, giving each container a distinct MCS category. Denials appear in the host's kernel log, not as application errors.

open as a page

What kind of vulnerability were the runc escapes CVE-2019-5736 and CVE-2024-21626, and what did they let an attacker do?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Both are escapes in runc, the low-level binary Docker uses to start containers. CVE-2019-5736 let a malicious container overwrite the host runc binary via /proc/self/exe and gain root on the next exec. CVE-2024-21626 (Leaky Vessels) used a leaked host file descriptor plus a crafted WORKDIR to break out to the host filesystem, at build or run time.

open as a page

showing 31–45 of 45