Switching a container to a declared unprivileged id makes files the image shipped unreadable - why, and what fixes it?
answer
- the build ran as someone else
- the superuser was never being checked
- owner, then group, then everyone else
- directories need a traverse bit too
- set ownership at build, not at start
basics
~20 sThe build ran as the superuser, so the files it installed are owned by id 0 with bits that grant access only to that owner. At runtime the new id matches neither owner nor group and falls to the least-permissive bits. Fix ownership at build time.
solid answer
~50 sTwo things changed at once. Before, the process was the superuser, which is not subject to ordinary permission checks, so every file the build installed was readable whatever its bits said. Now the process is some id such as `10001`, and each open is decided by whether that id is the file's owner, in its group, or neither - and for build-installed files it is usually neither, so the most restrictive bits apply. Directories need a traverse bit as well, so a single missing bit high in the path makes everything under it unreachable. The fix belongs in the build: set the ownership and the modes of the paths the process must read to the id it will run as, so the shipped layers are already correct. Fixing it at start-up is worse - it needs a privilege you just dropped, and a read-only root refuses the change anyway.
code
pseudocode · 8 lines# build time, after the application files are copied in
set_owner(path="/app", user=10001, group=10001, recursive=true)
set_mode(path="/app", directories="r-x", files="r--") # no write bit anywhere
# run time: process identity is user 10001, group 10001
open("/app/templates/report.tpl", "read") -> allowed # owner matches, read bit set
open("/app/cache/index.db", "write") -> denied # no write bit, and root is read-only
open("/scratch/index.db", "write") -> allowed # declared writable path, owned by 10001go deeper
Know that the files in an image are owned by the account that built it, so a container told to run as a different id may not be able to read them.
Explain the selection rule - owner, then group, then everyone else - why the superuser never hit it, and why the traverse bit on directories catches people out.
Fix it in the build rather than reverting the id, choose between re-owning and a shared group deliberately, and read each start-up symptom back to ownership, traverse, read-only root or a mount's ownership.
Make the safe shape the default: base images that already ship an unprivileged id and correct ownership, so every team inherits a working starting point instead of rediscovering this failure.
## Two changes arrive together Declaring an unprivileged id changes more than the name in the spec, and the breakage surprises people because both halves move at once. 1. **The process stops being exempt from permission checks.** The superuser is not evaluated against ordinary file permissions in the usual way; whatever the bits say, the open succeeds. An unprivileged id is evaluated on every single open. 2. **The identity no longer matches what the build produced.** Everything the build installed is owned by the account the build ran as - almost always id `0` - so the running id is, for those files, a stranger. The result is a container that worked yesterday and now exits during start-up, complaining about a file that is visibly present in the image. ## The permission arithmetic, briefly Every file carries an owning id, an owning group, and three sets of bits. Exactly one set applies to a given process, chosen in this order: | the process is... | the bits that apply | typical outcome for a build-installed file | |---|---|---| | the file's owner | the owner's set | not our case - the owner is the build account | | a member of the owning group | the group set | only if you deliberately arranged a shared group | | neither | the everyone-else set | the most restrictive set, and often no access at all | Two details cause most of the confusion: - **Directories need a traverse bit**, separate from the bit that lets you list them. A single directory high in the path without it makes everything underneath unreachable, and the error names the file rather than the directory that actually refused you. - **A group is as usable as an owner.** Giving the files a fixed group and running with that group id is a legitimate alternative to re-owning them file by file, and it is the better answer when several ids must read the same tree. ## Fix it in the build, not at start-up The durable fix is to make the shipped layers already correct: after the application files are copied in, set their ownership to the numeric id the workload will run as, and set modes that grant exactly what the process needs - read for data and configuration, traverse for directories on the path, execute for programs, and write for nothing at all, because writable paths are declared separately. The tempting alternative - a start-up step that walks the tree and re-owns it - is worse on three counts: - Changing another account's file ownership requires a privilege that a hardened workload has dropped, so the step fails exactly where the hardening is working. - With a read-only root filesystem, the change cannot be written at all. - It re-runs on every start, adding start-up time proportional to the size of the tree, for a result that never changes. ## The id that has no account A numeric id with no matching entry in the image's account database is normal - the kernel deals only in numbers - but anything that resolves a name behaves oddly. Lookups fail or return an empty result, a home directory may resolve to a path that does not exist, and a toolchain may then try to create a cache there and be refused. None of it is a permissions bug, and chasing it as one wastes an afternoon. The remedies are to give the id an entry in the image, or to point the affected settings at a declared writable path. ## Reading the symptom When a container starts failing right after the id is declared, the symptom tells you which of these it is: 1. **Refused on reading a file the image shipped** - ownership or modes from the build; fix in the build. 2. **Refused on reading a file inside a directory whose contents you can otherwise see** - a missing traverse bit somewhere on the path. 3. **Refused on writing anywhere under the image's paths** - not ownership at all, but the read-only root filesystem doing its job; declare a writable path. 4. **Refused on writing into a freshly mounted path** - the mount's ownership does not match the running id; set it in the declaration. 5. **An empty or nonsensical home directory, or a failed account lookup** - the id has no entry in the account database. ## Why this is worth getting right rather than reverting Every one of these failures is loud, happens at start-up, and happens in the same way on every environment. That makes it a far better failure mode than the alternative, which is a workload that runs as the superuser for two years because someone reverted the id the first time a start failed. The work is one build change and one pass through the paths the workload actually touches, and it is done once per image rather than once per incident.
- Why not re-own the files in a start-up step instead?Because changing another account's ownership needs a privilege a hardened workload has dropped, and a read-only root filesystem refuses the write regardless. It also costs start-up time on every replacement to produce a result that is identical every time. The build is where the answer is computed once and shipped.
- When is a shared group better than re-owning every file to one id?When more than one id must read the same tree - several containers in one group sharing a mounted path, or an image used by workloads that run under different declared ids. Give the files a fixed group with group-read, run the workloads with that group id, and ownership stops being a per-workload decision.
- The file is readable but the open still fails. What is left?Almost always a directory on the path missing its traverse bit: the file's own bits are irrelevant if a parent will not let you walk through it, and the error names the file rather than the directory. Check each directory from the mount point down. Failing that, the path is on a mount whose ownership does not match the running id.
saying these in an interview costs you the question
- A numeric id with no account entry cannot read any file.
- Re-owning the tree at start-up is the standard fix.
- A read-only root filesystem blocks the read as well.
- Permission bits stop mattering once the spec declares an id.
- Only the file's own bits matter, not the directories above it.