skip to content

An image now ends with `USER 10001`, and the application fails to start: it binds TCP port 80 and writes to /var/run and its own cache directory. What are your options to keep the process unprivileged?

level: middleimportance: should knowfreq 46%

answer

  1. listen 8080, publish -p 80:8080
  2. CAP_NET_BIND_SERVICE via setcap on the binary
  3. net.ipv4.ip_unprivileged_port_start is namespaced
  4. COPY --chown, not a later RUN chown -R
  5. --read-only + --tmpfs /tmp proves it

basics

~20 s

Move the listener to a high port and publish 80 on the host, or grant CAP_NET_BIND_SERVICE. For writes, chown the needed directories to the UID at build time, redirect state to /tmp, and mount tmpfs or a named volume for anything else.

solid answer

~40 s

Two separate problems. **Port 80.** The cleanest fix is to listen on 8080 and publish it: `-p 80:8080`. The host-side port is opened by the daemon, so nothing in the container needs privilege. If the port number is not configurable, either grant `CAP_NET_BIND_SERVICE` (a file capability set with `setcap` on the binary, or `--cap-add`) or lower the container's `net.ipv4.ip_unprivileged_port_start` sysctl to 0. **Writable paths.** A non-root process cannot write into root-owned directories. Fix ownership at build time — `RUN mkdir -p /app/cache && chown 10001:10001 /app/cache`, and `COPY --chown=10001:10001`. Point anything the app insists on writing under `/var/run` at a path you control, or mount tmpfs there. Set `HOME` so libraries do not write to `/`. Then you can add `--read-only --tmpfs /tmp` and know exactly which paths are mutable.

code

dockerfile · 9 lines
dockerfile
FROM nginx:1.27-alpine
RUN adduser -D -u 10001 web \
 && sed -i 's/listen  *80;/listen 8080;/' /etc/nginx/conf.d/default.conf \
 && mkdir -p /var/cache/nginx /tmp/nginx \
 && chown -R 10001:10001 /var/cache/nginx /tmp/nginx \
 && sed -i 's#/var/run/nginx.pid#/tmp/nginx/nginx.pid#' /etc/nginx/nginx.conf
ENV HOME=/tmp
USER 10001
EXPOSE 8080

go deeper

for a junior

Know the standard move: listen on a high port and publish 80 on the host, and chown the app's writable directories to the runtime UID.

for a middle

Explain both capability routes for the port and the build-time ownership techniques (COPY --chown, tmpfs, named volume seeding), and why chmod 777 is not one of them.

for a senior

Turn it into a verifiable property: --read-only plus --tmpfs plus --cap-drop ALL in CI so a regression to root or a stray writable path fails the build.

for a principal

Define the fleet convention — fixed UID, high ports everywhere, writable paths declared explicitly — so services do not each invent their own privilege workaround.

## Why an unprivileged process fails these two ways Running as UID 10001 removes two things root had implicitly: the capability to bind privileged ports, and `CAP_DAC_OVERRIDE`, which let root ignore file permission bits. Both show up as start-up failures, usually as permission denied on `bind()` and on `open()`. ## Ports below 1024 The kernel reserves TCP/UDP ports under 1024 for `CAP_NET_BIND_SERVICE`. Options, best first: 1. **Listen high, publish low.** Configure the app for 8080 and run `-p 80:8080`. The daemon, running as root on the host, opens host port 80 and forwards; the container process only ever binds 8080. This is the standard answer and keeps the image capability-free. 2. **File capability on the binary.** `setcap cap_net_bind_service=+ep /usr/sbin/nginx` at build time grants exactly that one capability to that one executable when it execs — much tighter than `--cap-add` at run time, which raises it for the whole container. The binary must sit on a filesystem supporting extended attributes, and the capability must be in the container's bounding set to take effect. 3. **Lower the sysctl.** `--sysctl net.ipv4.ip_unprivileged_port_start=0` applies inside the container's network namespace, so it does not affect the host, but it removes the guard for every process in the container. Deeper capability mechanics belong to the capabilities and seccomp topic; here you only need to know which single capability is at stake and that the port-remap answer usually avoids it entirely. ## Writable paths Design the image so exactly the directories the app needs are owned by the runtime UID and nothing else is writable: - Create the account and directories in one build step, then chown them: `RUN adduser -D -u 10001 app && mkdir -p /app/cache && chown 10001:10001 /app/cache`. - Copy artifacts with ownership rather than fixing it afterwards: `COPY --chown=10001:10001 target/app.jar /app/`. A later `RUN chown -R` duplicates the whole tree in a new layer and bloats the image. - Set `ENV HOME=/app` so libraries that write dotfiles land somewhere owned by the user instead of failing on `/` or `/root`. - For paths you cannot move, such as a hard-coded `/var/run/app.pid`, mount a writable tmpfs over that directory (`--tmpfs /var/run:rw,size=8m`) or symlink the file to a path under `/tmp` at build time. - Prefer tmpfs for scratch and a named volume for data that must survive. A named volume mounted over a path the image already chowned inherits that ownership, so it works without host-side setup. Once every write is accounted for, add `--read-only` to the run command. That converts "probably nothing writes to the image filesystem" into an enforced property, and any missed path fails loudly on first run instead of quietly persisting state into the container layer. ## Verifying Run the container and check `id`, then exercise the start-up path. `docker run --rm --read-only --tmpfs /tmp myapp` is a good smoke test in CI: if the app starts under it as a non-root user, the image is genuinely unprivileged rather than accidentally-still-root. Watch for entrypoint scripts inherited from a base image that chown or switch user — they will either fail or quietly re-add root.

  • When would you use setcap on the binary instead of --cap-add CAP_NET_BIND_SERVICE at run time?
    setcap grants the capability only to that executable when it execs, so other processes in the container — including a shell an attacker spawns — do not get it. --cap-add raises it for the container's whole bounding set. Prefer the file capability when the image is under your control and the filesystem supports extended attributes.
  • Why is a final `RUN chown -R 10001 /app` a poor fix in a Dockerfile?
    chown rewrites metadata on every file, and the layer records a fresh copy of all of them, so the image can nearly double in size for a large tree. Use COPY --chown on the copy step, or create and chown only the specific directories that need to be writable.

saying these in an interview costs you the question

  • chmod 777 on application directories to make the non-root user work
  • Adding --privileged so the app can bind port 80
  • Believing publishing -p 80:8080 requires privilege inside the container
  • Running a final RUN chown -R over the whole app tree
  • Assuming the container's ip_unprivileged_port_start change affects the host

context