How does Docker's default seccomp profile work, what kinds of syscalls does it block, and how would you handle an application that fails because of it?
answer
- seccomp-bpf filter on every syscall entry
- default: allow ~300, ERRNO the rest
- blocks mount/keyctl/kexec/module/CLONE_NEW*
- capability-conditional rules
- diagnose with SCMP_ACT_LOG, never ship unconfined
basics
~20 sseccomp is a kernel syscall filter. Docker applies a default JSON profile that allows most syscalls and returns EPERM for a few dozen dangerous ones (mount, reboot, kexec_load, keyctl, add_key, and namespace-creating clone flags). If an app breaks, identify the denied syscall and write a narrow custom profile — never use seccomp=unconfined in production.
solid answer
~50 s**seccomp-bpf** attaches a BPF filter to a process so the kernel evaluates every syscall against it. Docker generates a default profile from a JSON document: a `defaultAction` of `SCMP_ACT_ERRNO` plus an allowlist of the roughly 300 syscalls normal workloads use. The remaining ~40-50 are denied — `mount`, `umount2`, `reboot`, `kexec_load`, `init_module`, `keyctl`, `add_key`, `ptrace` in some versions, `bpf`, `perf_event_open`, `clone`/`unshare` with `CLONE_NEW*` flags. Some rules are conditional on capabilities, so an entry is allowed only if the container holds the matching capability. When an app fails, the symptom is `EPERM` (or a SIGSYS crash under `SCMP_ACT_KILL`) from an ordinary-looking call. Diagnose by running the workload once with `--security-opt seccomp=unconfined` **in a test environment** to confirm seccomp is the cause, then `strace -f -c` to find the syscall, then copy the default profile and add that one syscall. Never ship `seccomp=unconfined`, and remember `--privileged` disables the profile implicitly.
code
bash · 13 lines# 1. confirm seccomp is the cause (TEST ONLY)
docker run --rm --security-opt seccomp=unconfined myapp:1.4.2
# 2. find the offending syscall
docker run --rm --security-opt seccomp=unconfined --cap-add SYS_PTRACE \
myapp:1.4.2 strace -f -c /usr/local/bin/app
# 3. run with a reviewed custom profile
docker run --rm --security-opt seccomp=./seccomp-myapp.json \
--cap-drop ALL --security-opt no-new-privileges myapp:1.4.2
# audit what confinement a running container actually has
docker inspect --format '{{.HostConfig.SecurityOpt}} priv={{.HostConfig.Privileged}}' ctrgo deeper
Know that seccomp filters syscalls, that Docker applies a default profile, and that turning it off is not the fix.
Describe the allowlist structure, name representative blocked syscalls, and know that unconfined is test-only.
Diagnose a real failure end to end — confirm the layer, find the syscall with strace or an SCMP_ACT_LOG profile, ship a minimal custom profile — and explain capability-conditional rules and new-syscall drift.
Own the policy: runtime-default required by admission, custom profiles version-controlled and reviewed, engine-version drift managed, and privileged workloads accounted for as unconfined by definition.
## What seccomp is Secure Computing Mode with BPF ("seccomp-bpf") lets a process install a filter program that the kernel runs on every syscall entry. The filter sees the syscall number, the architecture, and the raw register arguments, and returns an action: `ALLOW`, `ERRNO` (fail with a chosen errno, usually `EPERM`), `KILL` (terminate the process/thread), `TRAP` (deliver SIGSYS), `LOG`, or `TRACE`. Crucially the filter cannot dereference pointers, so it can inspect scalar arguments and flags but not, say, the path string passed to `open`. Filters are inherited across `fork` and `exec` and cannot be removed. An unprivileged process may install one only if `no_new_privs` is set, which is why container runtimes set that bit whenever they apply a profile. The reason this matters for containers is attack surface. A modern kernel exposes 300-400 syscalls, and the container boundary is enforced *by* the kernel. Every syscall reachable from inside is a potential path to a kernel bug that breaks out. Historically, a large share of container escapes involved syscalls a normal application never issues. Filtering them is cheap surface reduction. ## Docker's default profile Docker ships a JSON profile (the moby `default.json`) applied to every container unless overridden. Structure: - `defaultAction: SCMP_ACT_ERRNO` — anything not matched is refused with `EPERM`. - `syscalls`: a list of groups, each with `names`, an `action` (usually `SCMP_ACT_ALLOW`), and optional `args` conditions and `includes.caps`. Categories of what it blocks: - **Namespace and mount manipulation** — `mount`, `umount2`, `pivot_root`, `unshare`/`clone` with `CLONE_NEWUSER` and friends. These are the building blocks of escape attempts. - **Kernel modification** — `init_module`, `finit_module`, `delete_module`, `kexec_load`, `reboot`. - **Kernel keyring** — `keyctl`, `add_key`, `request_key`. The keyring is *not* namespaced, so it has been a cross-container leakage vector. - **Tracing and performance** — `perf_event_open`, `process_vm_readv/writev`, `bpf`; `ptrace` was blocked historically and is allowed again on modern kernels. - **Misc dangerous** — `swapon`, `settimeofday`/`clock_settime`, `acct`, `nfsservctl`, obsolete syscalls. The capability-conditional rules are elegant: `mount` may be permitted only when the container holds `CAP_SYS_ADMIN`, so the profile automatically tightens for the common case and relaxes for the container that was deliberately given the capability. This is why the layers compose — dropping capabilities also narrows what seccomp will let through. ## Diagnosing a failure Symptoms are unhelpfully generic: `EPERM` from a call that "should" work, a runtime aborting at startup, or `Bad system call (core dumped)` when the action is `KILL`. Method: 1. **Confirm the cause.** In a *non-production* environment, run once with `--security-opt seccomp=unconfined`. If the failure disappears, seccomp is responsible. (If it persists, suspect capabilities or AppArmor/SELinux instead — check `dmesg` for MAC denials.) 2. **Find the syscall.** `docker run --security-opt seccomp=unconfined --cap-add SYS_PTRACE ... strace -f -c yourapp` shows the call and its error. A `LOG`-action profile is the more surgical variant: take the default profile, change `defaultAction` to `SCMP_ACT_LOG`, and read the kernel audit log for exactly which syscalls would have been denied. 3. **Write a minimal profile.** Copy the upstream default, add the specific syscall names to the allowlist, and run with `--security-opt seccomp=/path/profile.json`. Keep the profile in version control next to the deployment spec and review changes like code. 4. **Consider fixing the app.** Many hits are avoidable: a runtime probing `perf_event_open` for optional profiling, a library trying `keyctl`, a package manager running at container start. Often the right fix is to stop doing the thing at runtime. ## Historical examples worth knowing - Java and Node have at times used `membarrier`, `statx`, `clone3`, or `faccessat2` before the profile allowed them; a container that works on one host kernel and fails on a newer one with `EPERM` is the signature. The general shape — a **new syscall added by a newer kernel or glibc, not yet in the profile** — is the most common real-world seccomp breakage, and the reason `clone3` caused widespread fallout when glibc started using it. - Debuggers (`gdb`, `strace`) inside a container need `CAP_SYS_PTRACE` and, on older versions, a profile permitting `ptrace`. ## What not to do `--security-opt seccomp=unconfined` in production is a real reduction in security for a convenience gain, and it is invisible in normal monitoring — nothing logs that a container is unconfined until someone audits it. Similarly, `--privileged` implies unconfined seccomp, so "we only use privileged for one sidecar" also means "that sidecar has the full syscall surface". Audit with `docker inspect --format '{{.HostConfig.SecurityOpt}}'` and, in an orchestrator, with admission policy that requires the runtime default profile or a named custom one. Making the profile explicit rather than implicit is worth doing on its own: a spec that names its seccomp profile is one a reviewer can reason about.
- An image runs fine on one host and fails at startup with EPERM on another, with no configuration difference. How does seccomp explain that?The seccomp profile is an allowlist of syscall names, so a newer kernel or a newer glibc in the image can start using a syscall the profile does not list — `clone3`, `faccessat2` and `statx` all caused this in practice. The container engine version on the host determines which default profile is applied, so two hosts with different engine versions enforce different allowlists. The fix is to update the container engine, or ship a custom profile that permits the specific syscall.
- Why can a seccomp filter not simply block `open` on `/etc/shadow`?A seccomp-bpf filter runs in a restricted BPF context and cannot dereference user-space pointers, so it sees the pointer argument but not the path string it points to. Allowing it to follow the pointer would open a time-of-check-to-time-of-use race, since another thread could rewrite the memory after the check. Path-based restrictions therefore belong to AppArmor or SELinux, which evaluate inside the kernel's LSM hooks after the arguments are safely copied.
- When is `--security-opt seccomp=unconfined` defensible?During diagnosis in a throwaway environment, to confirm whether seccomp is the cause of a failure, and occasionally for a debug container that needs tracing syscalls. It should never appear in a deployed workload spec, because it silently restores the full syscall surface that container escapes have historically relied on. If a workload genuinely needs an extra syscall, add that syscall to a reviewed custom profile instead.
seccomp is a bouncer with a guest list at the kernel's door: it checks the name of every request going in, but it cannot read the message inside the envelope.
saying these in an interview costs you the question
- Calling seccomp a capability or a MAC system rather than a syscall filter
- Believing the default profile blocks most syscalls, when it allows the large majority
- Expecting seccomp to filter by file path or argument contents behind a pointer
- Shipping `seccomp=unconfined` as the fix for an application failure
- Forgetting that `--privileged` disables the seccomp profile as a side effect