You are setting a container runtime-confinement baseline for an organisation. Which restrictions do you make mandatory by default, how do you roll it out without breaking existing workloads, and how do you handle the workloads that genuinely need more privilege?
answer
- cap-drop ALL + no-new-privileges + seccomp default + MAC on + read-only rootfs
- hard no: privileged, host namespaces, docker socket
- audit/permissive first, enforce after fixing classes
- fix in shared base images, not per team
- exceptions: owner + narrow grant + node pool + expiry
basics
~20 sBaseline: drop all capabilities, no-new-privileges, runtime default seccomp, MAC profile on, read-only root filesystem, no privileged and no host namespaces. Roll out in audit/permissive mode first, fix the failures found, then enforce. Exceptions get a named owner, a scope, and an expiry.
solid answer
~50 s**Mandatory defaults** - `--cap-drop ALL`, additions only by exception. - `--security-opt no-new-privileges`. - Runtime default seccomp profile — explicitly named, never unconfined. - AppArmor `docker-default` or SELinux `container_t` left on. - `--read-only` root filesystem with explicit tmpfs for scratch. - No `--privileged`, no host PID/network/IPC namespaces, no `/var/run/docker.sock` mount. **Rollout**: never flip enforcement first. Inventory current settings across the fleet, run the policy engine in audit mode to see who would break, publish the report per team, fix the top classes centrally (base images that `chown` at startup, entrypoints using `sudo`, apps binding port 80), then enforce for new workloads, then for existing ones on a dated deadline. **Exceptions**: a named owner, the narrowest possible grant (one capability or one device, never privileged), a compensating control such as a dedicated node pool, an expiry date, and a review. Keep the exception list small, visible and dated — it is a risk register, not configuration.
code
bash · 14 linesdocker run -d \
--cap-drop ALL \
--security-opt no-new-privileges \
--security-opt seccomp=/etc/docker/seccomp/default.json \
--read-only --tmpfs /tmp:rw,noexec,nosuid,size=64m \
--pids-limit 256 --memory 512m \
-p 8080:8080 \
registry.example.com/prod/myapp@sha256:ab12cd34...
# audit the fleet for baseline violations
docker ps -q | xargs -r docker inspect --format \
'{{.Name}} priv={{.HostConfig.Privileged}} caps_add={{.HostConfig.CapAdd}} \
net={{.HostConfig.NetworkMode}} pid={{.HostConfig.PidMode}} \
secopt={{.HostConfig.SecurityOpt}} ro={{.HostConfig.ReadonlyRootfs}}'go deeper
Recall the baseline flags — drop all capabilities, no-new-privileges, default seccomp, read-only root filesystem, never privileged.
Explain what each control removes from an attacker and what commonly breaks when you enable them.
Run the rollout: inventory, audit mode, fix failure classes in shared base images, enforce for new then existing, and keep a sanctioned debug path.
Own the tradeoffs and the governance — one LSM per fleet, exceptions as a dated risk register with compensating node pools, drift measured continuously, and a clear statement of where a shared kernel stops being sufficient.
## What the baseline is trying to achieve The threat model is not "someone attacks the container from outside". It is: **an attacker already has code execution inside a container** — via an application vulnerability, a dependency, or a poisoned build step. Every control in the baseline answers one question: how much can that attacker do next, and how fast can they reach the node or the neighbouring workloads? The controls are complementary, and the layering matters: | Control | What it removes from the attacker | |---|---| | `cap-drop ALL` | privileged operations (raw sockets, mount, chown, permission bypass) | | `no-new-privileges` | escalation to container-root via setuid binaries | | default seccomp | the escape-relevant syscall surface (mount, keyctl, module loading) | | AppArmor/SELinux | reach to host objects and to other containers' files | | read-only rootfs | dropping tools and persisting a foothold | | no privileged / no host namespaces / no docker socket | the direct routes to host root | None alone is sufficient; together they turn "RCE in a container" from a node compromise into a contained incident. ## Choosing the defaults **Drop all capabilities.** The default fourteen include `NET_RAW` (spoofing against neighbours) and `DAC_OVERRIDE` (reading a mounted secret whose only protection is its file mode). Most services need none. Make zero the default and require justification for each addition. **no-new-privileges** is nearly free and closes the in-container escalation path. Combine it with images that carry no setuid binaries. **Seccomp: default, explicitly stated.** An implicit default is invisible to reviewers and silently disappears under `--privileged`. Requiring the profile to be *named* in the spec makes "unconfined" a visible, reviewable choice. Custom profiles live in version control beside the workload spec. **MAC on.** Pick one LSM per fleet — AppArmor or SELinux, not a mixture — because profile distribution, tooling and debugging skills all differ. Custom profiles ship in the node image, not by hand. **Read-only root filesystem** with explicit `tmpfs` mounts for the few writable paths. This is the control most likely to expose sloppy images, and also the one that most reduces attacker persistence. **Hard nos:** privileged, host PID/network/IPC namespaces, and the Docker socket. Each is a direct path to node root, and each has a legitimate but rare use served better by a dedicated, isolated pool. ## Rolling it out 1. **Inventory first.** You cannot set policy for a fleet you have not measured. Enumerate every workload's current capabilities, seccomp setting, privileged flag, host-namespace usage and socket mounts. The distribution is usually surprising: a small number of legacy workloads account for nearly all violations. 2. **Audit mode before enforcement.** Run the admission policy in warn/audit mode and let it report for a full release cycle. On SELinux hosts, `permissive` gives the equivalent signal at the LSM layer. The output is a per-team list of what would break — concrete, dated, and specific. 3. **Fix the classes centrally.** Failures cluster into a handful of causes: entrypoints doing `chown -R` (fix ownership at build time), `sudo`-based entrypoints (start as the right user), binding port 80 (listen high and map), scratch writes to `/tmp` (add a tmpfs), and images with setuid binaries (strip them in the base). Fixing them in shared base images resolves most of the fleet at once — far cheaper than dozens of independent migrations. 4. **Enforce for new first.** New workloads are cheap to make compliant; existing ones get a dated deadline with the audit report as the work list. 5. **Make it the default, not an option.** Ship a compliant workload template so the secure configuration is what people get by copying an existing service, and keep the policy engine as the backstop rather than the primary mechanism. ## Handling genuine exceptions Some workloads really do need more: storage and network node agents, monitoring agents that need host PID access, build farms. The discipline: - **Narrowest possible grant.** One capability, one `--device`, one host path — never `--privileged` because it is easier. Ask what specific operation failed and grant exactly that. - **Compensating controls.** Put the exception on a dedicated node pool so a compromise does not reach general workloads; restrict who can deploy to it; increase logging around it. - **Named owner and expiry.** An exception with no owner and no date becomes permanent. A dated review forces the question "is this still needed?" at least annually. - **Visibility.** Exceptions belong in a register the security team and the platform team both read, and ideally in a dashboard counting them. The metric that matters is not "is policy enabled" but "how many exceptions exist and are they trending down". ## The tradeoffs to say out loud - **Friction versus breach cost.** A strict baseline creates real work for teams whose images were written casually. The honest answer is that the cost is front-loaded into image fixes and mostly disappears afterwards, whereas the breach cost is unbounded. - **Debuggability.** Hardened containers are harder to troubleshoot — no shell, read-only, no ptrace. Provide a sanctioned path: ephemeral debug containers with elevated settings, short-lived and audited, so engineers do not weaken production specs to get a shell. - **False confidence.** A compliant baseline still shares a kernel. Workloads with genuinely hostile input — untrusted code execution, multi-tenant builds — need stronger isolation such as a sandboxed runtime or a VM boundary, not a stricter flag set. - **Drift.** Baselines decay. Enforcement must be automated in admission and reported continuously, because an unmeasured policy is an aspiration. The defensible summary: make the hardened configuration the default and the path of least resistance, roll it out by measuring before enforcing, and treat every exception as a dated, owned risk rather than a setting.
- How do you keep a hardened baseline from making production undebuggable?Provide a sanctioned debug path rather than letting teams weaken the production spec. Ephemeral debug containers that attach to the target's namespaces with elevated settings — extra capabilities, an unconfined seccomp profile, a shell and tooling — give engineers what they need for one session, and they disappear afterwards. Make those sessions audited and time-boxed so the elevated configuration never becomes the deployed one.
- Which single control would you enforce first if you could only have one, and why?Banning privileged containers, host namespaces and Docker-socket mounts, because each is a direct, one-step path from container compromise to node root, and the other controls are largely irrelevant while any of them is present. It is also the easiest to inventory and usually affects the fewest workloads. Capability dropping is the highest-value second step since it shrinks what a compromised process can ask the kernel for.
- When is this baseline not enough?When the workload runs genuinely untrusted code — multi-tenant build farms, customer-supplied functions, CI running arbitrary pull requests. All the controls still share one kernel, so a kernel vulnerability defeats them together. Those workloads need a stronger isolation boundary: a sandboxed runtime with a user-space kernel, a lightweight VM per workload, or physically separate node pools per tenant.
It is building-code enforcement: measure the existing stock, publish what fails inspection, fix the common defects in the standard blueprint, then enforce on new construction first — with variances that are signed, scoped and dated rather than permanent.
saying these in an interview costs you the question
- Enforcing policy fleet-wide before measuring what would break
- Granting `--privileged` as the standard exception because it is quicker than finding the specific capability
- Treating a hardened baseline as sufficient isolation for untrusted multi-tenant code sharing one kernel
- Letting exceptions accumulate without an owner, scope or expiry date
- Leaving the seccomp and MAC profiles implicit so reviewers cannot see whether a workload is confined