skip to content

What does `docker run --read-only` actually make read-only, and why do most containers then need `--tmpfs`?

level: middleimportance: must knowfreq 58%

answer

  1. Which filesystem does the flag cover
  2. Mounts are not part of it
  3. Where does a normal app still write
  4. RAM-backed scratch, and its ceiling
  5. noexec and nosuid on the scratch mount

basics

~20 s

--read-only mounts the container's root filesystem read-only, so every write to a path that came from the image fails with EROFS. Mounts are unaffected: volumes, bind mounts and tmpfs stay writable, so scratch paths such as /tmp and /run need an explicit --tmpfs.

solid answer

~40 s

`docker run --read-only` mounts the assembled root filesystem read-only, so a process that tries to write anywhere the image supplied gets `EROFS: read-only file system`. It does not touch mounts: named volumes, bind mounts and tmpfs mounts are separate mounts and stay read-write unless you mark them `:ro`. That is the hardening value — an attacker with code execution cannot drop a binary, overwrite a config the runtime will reload, or leave persistence behind. Almost every real program still writes somewhere, so you hand back the exact paths it needs: `--tmpfs /tmp:rw,noexec,nosuid,size=48m` and `--tmpfs /run:rw,noexec,nosuid,size=6m` give RAM-backed scratch that is empty at start and gone at stop, while genuine state goes on a named volume. Always set `size=` — tmpfs pages count against the container's memory.

code

bash · 6 lines
bash
docker run -d --name scorer \
  --read-only \
  --tmpfs /tmp:rw,noexec,nosuid,size=48m \
  --tmpfs /run:rw,noexec,nosuid,size=6m \
  -v scorer-state:/var/lib/scorer \
  fraud-scorer:1.4

go deeper

for a junior

Recall the one-line effect: writes to paths that came from the image fail with a read-only-file-system error, and --tmpfs /tmp is how you give a container scratch space back. Knowing the two flags belong together already puts you ahead.

for a middle

Be ready to explain the mechanics: the root filesystem is mounted read-only while mounts keep their own flags, which paths a typical service still writes, and why tmpfs needs an explicit size plus noexec and nosuid.

for a senior

An interviewer expects you to convert a real service without breaking it: identify the writable paths from evidence, decide tmpfs versus volume versus moving the write to build time, and state plainly what this control does not cover.

for a principal

Own the tradeoff of making immutable root filesystems the default. Talk about which classes of workload cannot comply, the memory cost of RAM-backed scratch across a fleet, and how you keep exceptions visible rather than silently dropping the flag.

## What the flag actually changes A container normally gets a thin writable layer stacked on top of the image's read-only layers by the storage driver (with overlay2, that is the `upperdir`). Everything the process creates or modifies at runtime lands in that layer and is discarded when the container is removed. `docker run --read-only` still gives the container that layer, but Docker mounts the assembled root filesystem read-only, so a write to any path that came from the image fails immediately with the kernel error `EROFS` — surfaced as `Read-only file system` in strace and shell output, and as messages like `EROFS: read-only file system, open '/app/.cache/entry.json'` from a language runtime. The purpose is not disk hygiene, it is blast radius. Once the root filesystem is read-only, an attacker who reaches code execution inside the container cannot drop a payload next to the application, cannot overwrite a JavaScript file or configuration the runtime will load on the next request, cannot install a package, and cannot leave anything behind that survives a restart. A whole family of exploitation steps that assume "write a file, then execute it" turns into an error at the first step. It also makes the container genuinely immutable: the only code that can run is what the image shipped, plus whatever you deliberately made writable. ## What stays writable The flag applies to the root filesystem only. Anything mounted into the container is a separate mount with its own flags: - **Named volumes** — `-v scorer-data:/var/lib/scorer` is writable under `--read-only`. - **Bind mounts** — writable unless you append `:ro`. - **tmpfs mounts** — writable by definition. So `--read-only` is not "nothing in this container can be written"; it is "nothing can be written except the paths I granted". If you want a mounted path read-only too, say so explicitly (`:ro`, or `readonly` in `--mount` syntax). ## Why tmpfs is nearly always required Very few programs write nothing. The paths that break first are boringly predictable: - `/tmp` — upload staging, temp files, unix sockets, JVM `hsperfdata`, extracted jars. - `/run` or `/var/run` — pid files and unix sockets (an nginx-style server writes `nginx.pid` here before it ever serves a request). - Package-manager and language caches — an npm or pip cache directory, `__pycache__`, a compiled-template directory. - `/var/log/...` for anything that logs to a file rather than to stdout. `--tmpfs /tmp:rw,noexec,nosuid,size=48m` mounts a RAM-backed filesystem there. Properties worth knowing: it starts empty on every run, disappears when the container stops, is private to that container (two replicas do not share it), and its pages are charged to the container's memory accounting — which is why an unsized tmpfs is a real hazard. A runaway temp file can consume host RAM or push the container into an OOM kill. Set `size=` on every tmpfs mount, even generously; the number is a ceiling, not a reservation. The `noexec` and `nosuid` options matter as much as the size. A writable `/tmp` is exactly where an attacker would drop a payload; `noexec` means the kernel refuses to execute it from that mount, and `nosuid` means a setuid bit on a file written there is ignored. Adding a tmpfs without those options gives back a good part of what `--read-only` just took away. ## Deciding what each path deserves When a path fails under `--read-only`, there are three good answers and one bad one: 1. **Ephemeral scratch** — mount a sized tmpfs. 2. **State that must survive a restart** — mount a named volume; then you have also documented, in the run command, exactly what the service's durable state is. 3. **Something written once at startup that never changes** — a downloaded model file, a generated asset bundle, a compiled template cache. Move it into the image at build time. That removes the write, speeds up start, and makes every replica byte-identical. 4. **The bad answer** — drop `--read-only`, or bind-mount something writable over a large part of the tree. ## Operating with it `docker inspect -f '{{.HostConfig.ReadonlyRootfs}}' <container>` confirms the setting. In Compose the same controls are `read_only: true` plus a `tmpfs:` list. Finally, keep the claim honest in an interview: a read-only root filesystem restricts *writes*. It does not restrict syscalls, capabilities, devices or the network, and it does not stop an attacker executing a shell, a package manager or an interpreter that the image already contains. It composes with a minimal base image and the rest of the runtime confinement flags; on its own it is one strong control, not a sandbox.

  • Under `--read-only`, can the process still write to a named volume mounted at /var/lib/scorer?
    Yes. `--read-only` applies to the root filesystem only; a named volume, a bind mount or a tmpfs is a separate mount and stays read-write. If you want one of them read-only as well you must say so — `-v scorer-state:/var/lib/scorer:ro`, or `readonly` in `--mount` syntax. That is exactly how you hand back the small set of writable paths a service genuinely needs while everything from the image stays fixed.
  • What goes wrong if you mount a tmpfs without a `size=` option?
    Nothing bounds it except memory. tmpfs is RAM-backed and its pages are charged to the container, so a log file or an upload buffer that grows without limit can drive the container into an OOM kill, or eat host memory if no memory limit is set. Always give every tmpfs mount an explicit ceiling; it is a maximum, not a reservation, so a generous number costs nothing when unused.
  • Why add `noexec` and `nosuid` to a tmpfs you mount at /tmp?
    Because a writable /tmp is the obvious place for an attacker to drop a payload after `--read-only` closed the rest of the filesystem. `noexec` makes the kernel refuse to execute anything from that mount, and `nosuid` makes it ignore setuid bits on files written there. Without them, the one directory you handed back becomes the staging area that the read-only root filesystem was supposed to eliminate.

It is the difference between a rented flat and a hotel room: the walls and fittings are fixed and you cannot alter them, but the desk drawer they hand you is yours to fill, and it is emptied the moment you check out.

saying these in an interview costs you the question

  • Thinks --read-only also makes mounted volumes and bind mounts read-only
  • Claims a read-only root filesystem makes the container immune to compromise
  • Mounts a tmpfs with no size limit and calls the container hardened
  • Says an app that writes to /tmp simply cannot be run read-only
  • Confuses --read-only with appending :ro to a single bind mount
  • Believes the flag changes the image rather than the running container

context