skip to content

Non-Root & Read-Only Root

Hardening one workload at the moment it starts: which account it runs as, whether its filesystem can be written, and which privileges it keeps. Asked because the convenient answer is the unsafe one.

on this pageshow

questions

5

After its root filesystem is made read-only, a report renderer cannot write scratch files - what must be declared?

level: juniorimportance: must knowfreq 62%

answer

  1. writes refused, reads untouched
  2. the exceptions have to be declared
  3. scratch space is where it breaks first
  4. size limit and owning id both matter
  5. point the temp setting at the new path

basics

~20 s

A read-only root filesystem refuses every write except into paths declared writable, so the renderer needs one mounted at its scratch directory and its temporary-file setting pointed there. Declare the size and the owning id too, or it breaks again.

solid answer

~50 s

Marking the root filesystem read-only tells the runtime to refuse writes to everything the image shipped and to the container's own writable layer. Reads are unaffected. Any process that writes at runtime - and a report renderer writes its intermediate files before uploading them - then fails at the first open, usually during start-up. The fix is not to undo the setting but to **declare** the exceptions: mount a writable path at the scratch directory, point the process's temporary-file setting at that path, give it a size limit so a runaway render cannot fill the node, and make sure it is owned by the id the process runs under. Done that way the hardening is a statement of intent: these three paths are writable, everything else is not, and a reviewer can read that off the spec.

code

yaml · 15 lines
yaml
workload:
  account:
    userId: 10001
    groupId: 10001
  rootFilesystem: read-only
  writablePaths:
    - mountPath: /scratch
      backing: ephemeral
      sizeLimit: 512Mi
      owner: 10001
  environment:
    RENDER_TEMP_DIR: /scratch
  privileges:
    dropDefaults: true
    allowPrivilegeGainOnExec: false

go deeper

for a junior

Remember that read-only applies to writes only, and that the fix for a failing scratch write is to declare a writable path and point the process at it.

for a middle

Explain the mechanics: the writable layer is withheld, the failure lands at the first open, and the declaration needs a mount point, a backing, a size limit and the right owning id.

for a senior

Diagnose it from the error rather than reverting: find every path the workload and its libraries write, declare each one narrowly, and bound the scratch area so a runaway job cannot take a node's disk with it.

for a principal

Decide how far the standard goes across teams: read-only by default with declared exceptions, who reviews a new writable path, and what the platform does when a workload cannot be made to fit.

## What a read-only root filesystem actually blocks A container normally sees the image's layers plus a thin writable layer stacked on top, and any write goes into that writable layer. Declaring the root filesystem **read-only** tells the runtime not to give the process that writable layer at all: every write to a path under `/` is refused with a permission-style error. Reads are untouched - the process can still read everything the image shipped. It is a write control, not a read control, and confusing the two is the most common misreading. What that buys you is narrow and worth having: - A compromised process **cannot rewrite the program files or the configuration** the running process reads, so it cannot quietly change the workload's behaviour under you. - It cannot drop tooling somewhere convenient and run it from there. - The set of paths that are writable becomes a **short, declared list** in the spec, which is something a reviewer can actually read and a gate can actually check, instead of an implicit 'everything'. What it does not buy you: it is not a durability control and not a boundary. A path you declare writable is writable, and a compromised process that finds one uses it. ## Why a report renderer is the first thing to break The renderer in this scenario builds a document into scratch space and then uploads it. That is an ordinary shape - anything that decompresses an upload, renders, buffers, or caches on disk does it - and it is exactly the shape that a read-only root breaks. The failure is unglamorous: the process opens its temporary file, gets a refusal, and either exits at start-up or fails on the first real request. Two details make it harder to diagnose than it should be: 1. The write is often not in the code you wrote. A library, a document toolchain or a font cache may pick a temporary directory on its own, and the error surfaces from somewhere you did not expect. 2. The same image was fine yesterday, because the writable layer was quietly absorbing all of it. ## Declaring the exceptions The fix is to enumerate the writable paths rather than to abandon the setting. A declaration has four parts worth getting right: - **Where** it is mounted - the exact directory the process writes to, not its parent. - **What backs it** - an ephemeral scratch area that lives and dies with the container is the right default for a renderer, because nothing here needs to survive replacement. - **How big** it may get. An unbounded scratch area is a way to fill the node's disk, and a runaway render is the classic cause. A declared ceiling turns that into a failed render instead of a degraded host. - **Who owns it**. A freshly mounted path whose ownership does not match the id the process runs under is writable in name only, and you get the same failure a second time. And one thing outside the spec: point the process at the path. Most runtimes and toolchains read a temporary-directory setting from the environment; if that still names the old location, the declaration achieves nothing. ## The two settings work together An unprivileged account and a read-only root filesystem are separate declarations, and each covers what the other does not: | declaration | what it stops | what it leaves open | |---|---|---| | unprivileged account id | reading and writing what that account may not touch | writes to anything the account does own, including its own program directory if ownership allows | | read-only root filesystem | every write outside the declared paths, whichever account is used | reads of everything, and writes into the declared paths | Used together, the answer to 'where can this process write?' becomes a list you can print, and that list is the thing a security review is asking for. ## How teams get this wrong - **Reverting the setting** the first time a start-up fails, and never coming back. The failure is information: it tells you exactly which paths the workload writes. - **Declaring the parent directory** writable to make the error go away, which usually re-opens most of what you just closed. - **Leaving it unbounded**, so the scratch area competes with every other workload on the node for disk. - **Assuming reads are affected**, and going looking for a permission problem that is really an ownership problem from a different change. - **Believing it is pointless** because the writable layer is discarded when the container is replaced anyway. It is discarded - but the attacker is inside the container now, and the whole value is that the changes cannot be made in the first place.

  • Why give the declared scratch path a size limit?
    Because an ephemeral writable path usually draws on the node's own storage, so an unbounded one lets a single runaway render consume the disk every other workload on that node depends on. A declared ceiling converts a host-wide problem into a single failed request with a clear error, and it puts the expected working-set size in the spec where a reviewer can question it.
  • Does a read-only root filesystem stop an attacker who already has code execution in the container?
    It removes a step rather than the attacker. They cannot modify the program files or configuration the process reads, and cannot leave tooling in a convenient place, so the compromise is harder to extend and does not change the workload's behaviour. Everything the process may legitimately read, and anything it may write in the declared paths, is still fully available to them.
  • Should the declared writable path survive the container being replaced?
    For scratch space, no - and saying so explicitly is part of the hardening. Declaring it ephemeral makes the intent visible, keeps the path from accumulating anything across replacements, and means nothing important can quietly come to depend on it. A path that genuinely must outlive the container is a different declaration and should be justified as one.

saying these in an interview costs you the question

  • A read-only root filesystem blocks reads as well as writes.
  • It is pointless because the writable layer is discarded anyway.
  • Just revert the setting when the container fails to start.
  • Declaring the parent directory writable is an equivalent fix.
  • A declared writable path needs no size limit or ownership.
open as a page

A workload runs as the image's default superuser account - what does that actually cost you?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Running as the superuser hands any code-execution bug full in-container privilege: it can rewrite the program files the running process reads, open every mounted credential regardless of ownership, and exploit whatever the boundary fails to close. An unprivileged, explicitly declared numeric id removes that head start.

open as a page

What does dropping a container's default privilege set add when the workload already runs under an unprivileged account?

level: middleimportance: should knowfreq 45%

basics

~20 s

Who the process is and what powers it holds are separate axes. Runtimes hand every container a default slice of superuser powers, held by the process rather than its account. Dropping them all and adding back only what is needed closes that second axis.

open as a page

Switching a container to a declared unprivileged id makes files the image shipped unreadable - why, and what fixes it?

level: middleimportance: should knowfreq 50%

basics

~20 s

The build ran as the superuser, so the files it installed are owned by id 0 with bits that grant access only to that owner. At runtime the new id matches neither owner nor group and falls to the least-permissive bits. Fix ownership at build time.

open as a page

Under an unprivileged account, how can a process still gain privilege when it executes another program?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Images can ship program files marked to run with their owner's identity or extra powers rather than the caller's, so executing one raises privilege even from an unprivileged account. A declaration that forbids any privilege gain at exec blocks that, permanently and for every child.

open as a page