skip to content

Beyond choosing a non-root UID, which Kubernetes container securityContext fields would you set to harden a workload — specifically capabilities.drop, readOnlyRootFilesystem, allowPrivilegeEscalation and seccompProfile — and what does each one actually prevent?

level: middleimportance: must knowfreq 58%

answer

  1. drop ALL, add back at most NET_BIND_SERVICE
  2. readOnlyRootFilesystem + emptyDir for /tmp
  3. allowPrivilegeEscalation:false = no_new_privs, kills setuid
  4. seccompProfile: RuntimeDefault blocks unused syscalls
  5. sidecars and init containers need the same block

basics

~20 s

Drop all Linux capabilities (add back only what you need), set readOnlyRootFilesystem so the image cannot be modified at runtime, set allowPrivilegeEscalation: false so setuid binaries cannot gain rights, and set seccompProfile RuntimeDefault to block rarely-used syscalls. Together they shrink the kernel attack surface.

solid answer

~50 s

Four container-level fields, all cheap: - **`capabilities.drop: ["ALL"]`** — Linux splits root's powers into capabilities (`NET_RAW`, `SYS_ADMIN`, `CHOWN`, …). Runtimes grant a default set even to non-root containers. Drop everything, then `add` only the one you truly need, most often `NET_BIND_SERVICE`. - **`readOnlyRootFilesystem: true`** — mounts the image's filesystem read-only, so an attacker cannot drop a binary, rewrite config, or persist. Anything that must be writable becomes an explicit `emptyDir` mount (`/tmp`, cache dirs), which also documents the app's true write set. - **`allowPrivilegeEscalation: false`** — sets the kernel's `no_new_privs` bit: no `setuid`/`setgid` binary or file capability can raise the process's privileges after exec. This is what stops the classic `sudo`/`ping`-style escalation inside the container. - **`seccompProfile.type: RuntimeDefault`** — applies the runtime's syscall filter, blocking hundreds of syscalls the workload never uses and which historically host kernel CVEs. Also keep `privileged: false` and avoid host namespaces. Set them at pod level where possible and let admission enforce them.

code

yaml · 24 lines
yaml
spec:
  securityContext:
    runAsNonRoot: true
    runAsUser: 10001
    seccompProfile:
      type: RuntimeDefault
  containers:
    - name: api
      image: registry.example.com/api:2.1.0
      securityContext:
        allowPrivilegeEscalation: false
        readOnlyRootFilesystem: true
        capabilities:
          drop: ["ALL"]
      volumeMounts:
        - name: tmp
          mountPath: /tmp
        - name: cache
          mountPath: /var/cache/app
  volumes:
    - name: tmp
      emptyDir: {}
    - name: cache
      emptyDir: {}

go deeper

for a junior

Name the four fields and one sentence each on what they stop; know that emptyDir is how you keep /tmp writable under a read-only root.

for a middle

Explain the mechanism — default capability sets, no_new_privs, the runtime seccomp profile — and the concrete breakages each causes plus their fixes.

for a senior

Show how you roll these across an existing fleet: audit first, fix writable paths, enforce via admission, and debug Operation-not-permitted failures back to a specific capability or syscall.

for a principal

Argue the policy: hardened defaults injected centrally, an exception process with named owners and expiry for the workloads that need SYS_ADMIN or host access, and how much residual risk remains against a kernel bug.

## Why these four A container is a process on a shared kernel. Hardening means shrinking what that process can ask the kernel to do, and what it can change about itself. Each of these fields closes one class. ## capabilities.drop Linux long ago split root's omnipotence into ~40 **capabilities**. `CAP_NET_RAW` allows raw sockets (packet spoofing, ARP tricks). `CAP_SYS_ADMIN` is a grab-bag that includes mounting — the single most abused capability. `CAP_CHOWN` and `CAP_DAC_OVERRIDE` bypass file ownership and permission checks. `CAP_SETUID` lets a process change UID. Container runtimes grant a **default set** (historically about fourteen capabilities) to every container, and those apply to the file capabilities a binary can pick up even when the process is non-root. The hardened posture is: ```yaml capabilities: drop: ["ALL"] add: ["NET_BIND_SERVICE"] # only if you must bind <1024 ``` In practice most application containers need nothing added — the better fix for port 80 is to listen on 8080 and map it in the Service. Reach for `add` only with a named justification; `SYS_ADMIN`, `NET_ADMIN`, `SYS_PTRACE` and `SYS_MODULE` should be treated as near-equivalent to privileged. ## readOnlyRootFilesystem Set to `true`, the container's root filesystem layer is mounted read-only. The attacker who lands remote code execution can still run what is already in the image but cannot install tooling, overwrite a config file, or drop a webshell that survives. It also catches sloppy applications that write logs or state into the image layer instead of a volume — data that would be lost on restart anyway. Making it work is a small amount of archaeology: find every path the process writes and mount an `emptyDir` there. `/tmp` is nearly universal; JVMs want a temp dir, nginx wants `/var/cache/nginx` and `/var/run`, many frameworks want a cache directory. Set `TMPDIR`/`java.io.tmpdir` to your mounted path if needed. The end state is a manifest that explicitly lists the app's writable surface, which is valuable documentation in itself. ## allowPrivilegeEscalation Setting `false` sets the kernel's `no_new_privs` flag on the process. From that point, `execve` can never grant more privilege than the caller had: setuid bits are ignored, file capabilities are not picked up, and LSM transitions cannot raise privilege. This is what neuters an image that happens to ship `sudo`, a setuid `ping`, or a setuid helper with a bug. Note the interactions: `privileged: true` or adding `CAP_SYS_ADMIN` implies the ability to escalate, so the API server rejects `allowPrivilegeEscalation: false` combined with `privileged: true`. Also note that `runAsNonRoot` without `no_new_privs` is weaker than it looks — a non-root process that can exec a setuid-root binary is one step from root. ## seccompProfile **seccomp** filters syscalls. `seccompProfile: { type: RuntimeDefault }` applies the container runtime's curated profile, which blocks on the order of 40–60 syscalls almost no application uses — `keyctl`, `ptrace` in some profiles, `add_key`, `mount`, kernel module loading, obscure legacy calls. Several well-known container escapes were built on syscalls this default blocks. The cost is essentially zero for normal workloads; the exceptions are profilers, debuggers and some JVM/Node native tooling that wants `ptrace` or `perf_event_open`, which need a `Localhost` profile pointing at a custom JSON filter on the node. Historically, seccomp was off unless annotated; the `seccompProfile` field is the modern, GA way to say it, and Kubernetes 1.27+ can also default it cluster-wide via the kubelet's `SeccompDefault` feature. ## Putting it together The four fields plus `runAsNonRoot`, no host namespaces (`hostNetwork`/`hostPID`/`hostIPC`), no `hostPath` volumes and no privileged mode are almost exactly the **restricted** Pod Security Standard. Writing them by hand per workload is fine; enforcing them is the job of Pod Security Admission at the namespace level, so that a manifest missing them is rejected rather than merely reviewed. Container-level settings override pod-level ones, so a single unhardened sidecar can quietly reopen the hole — check every container in the pod, including init and ephemeral containers. ## Debugging when it breaks `Permission denied` on write → missing writable mount under a read-only root. `Operation not permitted` on a syscall → a dropped capability or the seccomp filter; `strace` under a permissive profile or run the audit-mode profile to identify which. A process that used to become root and now cannot → `no_new_privs` doing its job; redesign so the privileged step happens in the image build or an init container with a narrow capability instead.

  • Your app must bind port 80. How do you satisfy that without running as root or adding broad capabilities?
    Preferred: change the app to listen on 8080 and let the Service map port 80 to targetPort 8080 — no capability needed at all. If the binary genuinely cannot be reconfigured, drop ALL capabilities and add back only NET_BIND_SERVICE, which permits binding low ports and nothing else. Never solve it by running as root or by setting privileged.
  • An application fails with 'Operation not permitted' after you set seccompProfile: RuntimeDefault. How do you find the offending syscall?
    Run the workload temporarily with an Unconfined or audit-logging seccomp profile and trace it — strace, or the runtime's seccomp audit logs on the node — to identify the blocked syscall. Then decide: either the behaviour is unnecessary (a profiler or debug agent that should not be in production), or you author a Localhost profile that is RuntimeDefault plus that specific syscall, deployed to the node's seccomp directory.
  • Why is allowPrivilegeEscalation: false valuable even when the container already runs as a non-root UID?
    Non-root only sets the starting identity. Without no_new_privs, an unprivileged process that can execute a setuid-root binary present in the image — sudo, a buggy helper, a setuid ping — inherits root on exec. Setting allowPrivilegeEscalation: false makes the kernel ignore setuid bits and file capabilities for that process tree, closing the cheapest escalation path inside the container.

saying these in an interview costs you the question

  • Adding CAP_SYS_ADMIN 'to make it work' and calling that hardened
  • Thinking readOnlyRootFilesystem also makes mounted volumes read-only
  • Believing dropping capabilities is unnecessary because the container is non-root
  • Assuming seccomp RuntimeDefault is on by default in every cluster
  • Hardening the main container but leaving sidecars and init containers untouched

context