How would you run a Docker container whose filesystem is immutable except for the few paths that genuinely need to be written, and what does marking an individual mount read-only actually enforce?
answer
- --read-only = root fs; mounts are exempt
- add back: tmpfs /tmp + /run, volume for state, :ro binds for config
- ro constrains this container only, not the host
- bind-recursive=readonly for nested submounts
- docker inspect .Mounts shows RW true/false
basics
~20 sRun with --read-only to make the root filesystem immutable, then add back exactly what is needed: sized tmpfs mounts for /tmp and /run, a named volume for real data, and :ro bind mounts for config and certs. A read-only mount blocks writes from that container only — the host and other containers can still change the data.
solid answer
~50 sTwo independent controls. **`--read-only`** makes the container's root filesystem (image layers plus its writable layer) immutable. Nothing the process does can modify the container's own filesystem, so a compromised process cannot drop a binary, rewrite a config, or leave persistence behind. You then re-open only what is required: `--tmpfs /tmp` and `--tmpfs /run` (sized) for scratch, and a named volume for data that must persist. Mounts are exempt from `--read-only` — that is what makes the pattern usable. **Per-mount `readonly` / `:ro`** applies to one mount. `-v /etc/certs:/etc/certs:ro` or `--mount type=bind,...,readonly` mounts the filesystem read-only inside this container's mount namespace. Two limits to state: it constrains *this container only* — the host or another container with a read-write mount can still change the files underneath it — and historically a read-only bind did not propagate to nested submounts under the source, which recent Docker addresses with `bind-recursive`. Verify, don't assume: `docker inspect` the mounts and try a write in the running container.
code
bash · 11 linesdocker run -d --name api \
--read-only \
--tmpfs /tmp:size=64m,mode=1777 \
--tmpfs /run:size=16m \
--mount type=bind,source=/srv/api/config.yaml,target=/etc/api/config.yaml,readonly \
--mount type=volume,source=apidata,target=/var/lib/api \
myorg/api:1.4
# verify
docker inspect -f '{{json .Mounts}}' api
docker exec api sh -c 'touch /etc/api/x || echo blocked'go deeper
Know that :ro on a mount stops the container writing to it, and that --read-only locks the container's own filesystem.
Be able to assemble the pattern — read-only root plus tmpfs for /tmp and /run plus a volume for state — and explain that mounts are exempt from --read-only.
Own the limitations: scope is this container only, it is not integrity, nested submounts and the bind-recursive option, and how you would discover the writable paths an image actually needs via docker diff.
Position it as a baseline enforced in the deployment definition and in review, alongside non-root users and dropped capabilities, and be clear about the residual risk it does not cover — shared writable data, supply-chain integrity of the image itself.
## Two different knobs Candidates often blur these together, and the interviewer is usually checking whether you can separate them. 1. **`--read-only`** is a container-level flag. It mounts the container's root filesystem read-only, so the image layers *and* the container's thin writable layer become immutable. Explicit mounts you attach (volumes, binds, tmpfs) are not affected — they keep whatever mode they were given. 2. **A read-only mount** is per-mount: `-v src:dst:ro`, `--mount ...,readonly`, or Compose's `read_only: true` on a long-syntax entry. It says "this particular filesystem is not writable from inside this container". You almost always use them together: lock the root, then decide path by path what gets opened and in which mode. ## Building the pattern Start from `--read-only` and let the app tell you what it needs. Typical result: - `--tmpfs /tmp:size=64m,mode=1777` — nearly every runtime, library and shell tool expects a writable `/tmp`. - `--tmpfs /run:size=16m` — PID files, unix sockets, and anything that follows the FHS for runtime state. Nginx also wants its cache directories; Java may want a writable `/tmp` for the perf/hsperfdata files unless disabled. - A **named volume** on the one directory that holds real state (`/var/lib/app`), read-write. This is the data you back up. - **Read-only bind mounts** for injected material: `/etc/app/config.yaml:ro`, `/etc/ssl/certs:ro`. Config the container reads but must never rewrite is exactly the case for `:ro`. The payoff is concrete: an attacker with code execution cannot persist a payload in the image filesystem, cannot rewrite the config to point at their own endpoint, and any tampering they achieve dies with the container. It also surfaces sloppy applications that scribble into their install directory, which is worth knowing before production does. ## What read-only actually enforces — and does not - **Scope is this container's mount namespace.** A read-only bind mount does not make the host directory read-only. The host, and any other container that mounts the same path read-write, can still change the content, and the read-only container will see those changes. `:ro` is a guarantee about *who may write*, not about *what the data will be*. - **It is not integrity.** If you need the config to be unchanged, you need signing or checksums; a read-only mount only stops this container from being the one that changes it. - **Recursion caveat.** A bind mount is a mount of one point, and historically the `ro` flag applied to that mount but not to submounts already present under the source path — a nested mount could remain writable inside a nominally read-only bind. Modern Docker on a recent kernel supports recursive read-only binds; `--mount type=bind,bind-recursive=readonly,...` makes the intent explicit. If you are relying on read-only for a security property over a path with nested mounts, check this rather than assuming. - **Volumes can be read-only too.** `-v appdata:/var/lib/app:ro` is legitimate — for example a sidecar or exporter that reads another container's data without being able to corrupt it. Note the seeding interaction: an *empty* volume mounted read-only never gets populated in a useful way, so seed it from a writer first. - **Root can sometimes remount.** A container process running as root with `CAP_SYS_ADMIN` may be able to remount paths; read-only mounts pair with dropping capabilities and running as a non-root user rather than replacing them. ## Making it stick Do it in the deployment definition, not by convention: `read_only: true` plus `tmpfs:` and `:ro` bind entries in Compose, so the property is reviewable in a diff. Verify with `docker inspect --format '{{json .Mounts}}'` (each mount reports `RW: true/false`) and by shelling into the container and attempting a write — the failure mode you want is a clean `Read-only file system` error, at start time, not a surprise in production. Expect to iterate once: most images need one or two writable paths you did not predict, and finding them is the actual work of adopting the pattern. ## Summarising for the interviewer "`--read-only` for the root, tmpfs for scratch, one named volume for state, `:ro` binds for injected config — and I know `:ro` binds only this container, so the host can still change the data underneath it." That sentence covers the mechanic, the pattern and the limitation.
- You mount a host directory into a container with `:ro`. Can the files still change while the container runs?Yes. The read-only flag applies to that container's view: it cannot write. The host, or another container mounting the same directory read-write, can modify or replace the files at any time, and the read-only container will observe the new content. If you need the content to be stable or trustworthy you need copying, signing or checksums, not `:ro`.
- An image fails to start under `--read-only`. How do you find what it needs writable?Run it without the flag and observe what it writes — `docker diff <container>` lists files added or changed in the writable layer, which is a direct list of candidate paths. Then reintroduce `--read-only` and add a sized tmpfs for the ephemeral ones and a volume for anything that must persist. Repeat until it starts clean; the usual answers are `/tmp`, `/run`, and a cache or log directory.
saying these in an interview costs you the question
- Thinking `--read-only` also makes attached volumes and bind mounts read-only.
- Claiming a `:ro` bind mount prevents the host from modifying the files.
- Treating a read-only mount as an integrity guarantee rather than a write restriction on one container.
- Assuming a read-only bind mount automatically covers nested submounts under the source path.
- Abandoning `--read-only` at the first startup failure instead of adding a tmpfs for `/tmp` and `/run`.