In a Kubernetes Pod spec, what do the securityContext fields runAsNonRoot and runAsUser do, and why does it matter whether a container process runs as UID 0?
answer
- runAsUser = assign UID; runAsNonRoot = refuse UID 0
- UID 0 in container = UID 0 in kernel
- kubelet check → CreateContainerConfigError
- named USER without numeric UID also fails the check
- fsGroup fixes volume write permissions
basics
~20 srunAsUser sets the numeric UID the container process runs as. runAsNonRoot: true makes the kubelet refuse to start the container if it would run as UID 0. Root inside the container is real root on the host kernel, so any container escape starts privileged.
solid answer
~50 s`runAsUser: <uid>` tells the kubelet to start the container's main process under that numeric UID, overriding the image's `USER`. `runAsNonRoot: true` is a guard rather than a setting: it does not choose a UID, it makes the kubelet **refuse to start** the container if the resolved user is UID 0, failing the pod with `CreateContainerConfigError`. `runAsGroup` and `fsGroup` do the analogous thing for the primary group and for volume ownership. It matters because there is one kernel. UID 0 in the container is UID 0 on the host unless user namespaces remap it, so a container breakout, a hostPath mount, or a writable device starts from root privileges instead of an unprivileged account. Running non-root also blocks the easy in-container mischief: installing packages, writing to `/etc`, binding ports below 1024. Set it at pod level so it applies to every container, and make sure the image's files are readable by that UID.
code
yaml · 21 linesapiVersion: v1
kind: Pod
metadata:
name: web
spec:
securityContext:
runAsNonRoot: true
runAsUser: 10001
runAsGroup: 10001
fsGroup: 10001
containers:
- name: app
image: registry.example.com/web:1.4.2
ports:
- containerPort: 8080
volumeMounts:
- name: scratch
mountPath: /app/tmp
volumes:
- name: scratch
emptyDir: {}go deeper
Know that runAsUser sets the numeric UID, runAsNonRoot blocks UID 0, and that container root is host root at the kernel level.
Add the failure modes — named USER not resolvable, port-below-1024 binds, volume permissions solved with fsGroup — and where pod-level versus container-level context applies.
Frame it as one field in a hardening set enforced at two layers (admission rejects, kubelet backstops), and describe migrating existing images to non-root without downtime.
Discuss it as a fleet-wide default: base images built for arbitrary UIDs, defaults injected by policy, and the tradeoff between mandating non-root everywhere and carving exceptions for agents that genuinely need host access.
## The setting Every container in a Kubernetes pod eventually becomes an ordinary Linux process on some node. `securityContext` is the block that decides the identity and privileges of that process. It exists in two places: on the **pod** (`spec.securityContext`), where it applies to all containers and to volume ownership, and on each **container** (`spec.containers[].securityContext`), where it overrides the pod-level value for that container only. - `runAsUser: 10001` — start the process with numeric UID 10001. This overrides whatever `USER` the container image declared. - `runAsGroup: 10001` — the primary GID. - `runAsNonRoot: true` — a *validation*, not an assignment. At container-creation time the kubelet resolves the effective user; if it is UID 0 it refuses to start the container and the pod reports `CreateContainerConfigError` with a message about running as root. If the image only specifies a user by *name* (e.g. `USER app`) and no `runAsUser` is given, the kubelet cannot resolve it to a number before start and also fails the check — which is why hardened images set a numeric `USER`. - `fsGroup: 10001` — a supplemental GID applied to mounted volumes; the kubelet chgrp's and setgid's the volume so a non-root process can write to it. ## Why UID 0 is the thing to avoid Containers are not virtual machines. All containers on a node share the host kernel; isolation comes from namespaces (what you can see) and cgroups (what you can consume), plus capabilities, seccomp and LSMs (what you can ask the kernel to do). Unless user namespaces are enabled to remap IDs, **UID 0 inside the container is UID 0 in the kernel's eyes**. So: - Any kernel vulnerability or misconfigured mount that lets the process reach host resources is exercised as root. - A `hostPath` volume mounted by a root process can read or rewrite host files — including kubelet config or the container runtime socket. - Root retains a large default capability set (`CHOWN`, `SETUID`, `DAC_OVERRIDE`, `NET_RAW`, …), which is the raw material for most privilege escalation chains. Running as an unprivileged UID does not by itself prevent escape, but it removes the cheap paths and makes the remaining ones need a second bug. ## Practical consequences Most breakages when you flip a workload to non-root are file permissions: - The image writes to a directory owned by root (e.g. `/var/run`, `/app/tmp`). Fix in the Dockerfile by `chown`ing to the runtime UID, or mount an `emptyDir` and use `fsGroup`. - The app binds port 80. An unprivileged process cannot bind ports below 1024 without `CAP_NET_BIND_SERVICE`; the clean fix is to listen on 8080 and let the Service map 80 → 8080. - The image expects a matching entry in `/etc/passwd` for the UID. Most software tolerates a missing entry; some (git, some JVM tooling) does not, so add a passwd entry at build time or use an image built for arbitrary UIDs. ## How it relates to the wider hardening set `runAsNonRoot` is one field among several that the Pod Security Standards' *restricted* level requires, alongside dropping all Linux capabilities, forbidding privilege escalation, and setting a seccomp profile. It is the first one to adopt because it is the easiest to reason about and catches the largest class of mistakes. Note the two-layer nature: `runAsNonRoot` is enforced by the **kubelet at start time** on a per-container basis, whereas Pod Security Admission rejects the pod earlier, at the API server, before it is ever scheduled. Both firing on the same field is normal — admission gives the fast, clear rejection, the kubelet is the backstop for anything that got past it. ## Verifying `kubectl exec <pod> -- id` shows the effective UID/GID at runtime. In a review, grep manifests for pods with no `securityContext` at all — an unspecified `runAsUser` means "whatever the image says", and a large share of public images still say root.
- A pod with runAsNonRoot: true and no runAsUser fails to start even though the image has a USER line. Why?The image declared the user by name (for example `USER app`) rather than by numeric UID. The kubelet performs the non-root check before the container starts and cannot resolve a username to a UID at that point, so it treats the identity as unverifiable and refuses. The fix is either a numeric `USER 10001` in the Dockerfile or an explicit `runAsUser` in the pod spec.
- Does running as non-root make a container escape impossible?No. It removes the easiest paths and forces an attacker to find a second bug, but kernel vulnerabilities, over-permissive capabilities, mounted host paths and mounted service-account tokens can all still be abused by an unprivileged UID. Non-root is one layer, combined with dropping capabilities, blocking privilege escalation, a seccomp profile and a read-only root filesystem.
- What is the difference between setting the user in the Dockerfile and setting runAsUser in the pod spec?The Dockerfile `USER` is the image default and travels with the image, which is good hygiene but can be overridden by whoever runs it. `runAsUser` in the pod spec is the cluster-side statement and is what admission control and policy can inspect and enforce. Do both: a numeric `USER` in the image, and an explicit securityContext in the manifest so the intent is visible and auditable.
runAsUser is choosing which employee badge the process wears; runAsNonRoot is the door guard that turns away anyone wearing the master badge.
saying these in an interview costs you the question
- Believing container root is somehow a different, sandboxed root than host root
- Thinking runAsNonRoot picks a UID for you rather than just rejecting UID 0
- Setting the securityContext only on one container and assuming siblings inherit it
- Claiming non-root alone makes the workload secure
- Fixing permission errors by reverting to root instead of chowning image paths or using fsGroup