Root inside a container cannot load kernel code or set the clock, yet it can change file ownership - what decides the difference?
answer
- authority is a set, granted in part
- the line is blast radius
- one clock, one kernel, one host
- container-scoped powers are kept
- ask for one privilege, not all
basics
~20 sRoot's authority is a set of separately grantable privileges, and a runtime hands a workload a small default subset. Privileges whose effect reaches the shared host are withheld; privileges whose effect ends inside the container's own view are kept.
solid answer
~50 sA machine's root authority is not one switch - it is a collection of individually grantable privileges, and a container runtime starts a workload holding only a default subset of them. The line is drawn by blast radius. Loading kernel code and setting the clock act on things every workload on the host shares, and there is exactly one of each, so granting them to one container would hand it the machine. Changing ownership inside the container's own root filesystem acts on a view that belongs to that container alone, so it is safe to keep. The useful consequence is diagnostic: when a process that reports as root is refused, look at what the operation touches. If it reaches the host, the refusal is deliberate, and the answer is to request that single privilege or to move the work out of the container.
go deeper
Take away the shape: root's authority comes in separate pieces and a container gets some of them. Two refusals worth remembering are setting the clock and loading kernel code.
Explain the line the default draws - shared with the host versus bounded by the container - and use it to predict whether a given operation will be refused before you try it.
Handle the real cases: know the three answers to a reserved low port, and insist on requesting one named privilege rather than restoring the set, so the reason stays reviewable.
Your call is what the default set and the exception path look like across the estate: who may request an added privilege, what evidence the request carries, and which requests mean the work does not belong in a container.
## Root is a set of privileges, not a single flag On a modern host, the authority historically bundled into "being root" is split into many separate privileges that can be held or not held independently: the authority to load kernel code, to set the clock, to change the machine's network configuration, to mount filesystems, to reach into other processes, to bind a listening socket to a port in the reserved low range, and so on. A container runtime uses that split. It starts the workload's first process holding a small **default subset** rather than the whole collection, so a process that reports as root inside the boundary has root's identity in its own view and only a fraction of root's powers. That is why the question's two halves are both true at once: two operations refused, one allowed, and no contradiction between them. (The kernel's own naming for these privileges and how they are tracked is a host-operating-system subject; what matters here is the shape - separately grantable, granted in part, by default.) ## The rule the default follows Run down what is normally withheld and what is normally kept, and one line explains both: | Normally withheld | Why | Normally kept | Why | |---|---|---|---| | Loading or unloading kernel code | One kernel serves every workload on the host | Ownership and mode changes in its own root filesystem | That filesystem view is the container's alone | | Setting the clock | One clock for the whole machine | Signalling its own processes | Its process table contains only its own | | Changing the host's network configuration and packet-filtering rules | Shared with every other workload | Binding ordinary ports in its own network view | That view is private to the container | | Mounting and unmounting filesystems | Reshapes what everything on the host can see | Creating users and groups in its own view | They exist nowhere else | | Writing kernel tunables or rebooting | Machine-wide, and not settable per container | Reading and writing its own files | Confined to its own layers | **A privilege that would change something shared with the host is withheld; a privilege whose effect ends at the boundary is kept.** That is the whole rule, and it predicts refusals you have not personally hit yet. ## The omissions teams actually run into Three show up repeatedly: 1. **Binding a listening socket to a reserved low port.** Historically only a privileged process may take a port below a fixed threshold, and that privilege is commonly not in the default set. The workload starts as root, and the bind is still refused. 2. **Setting the clock.** A workload that wants to correct time drift for itself discovers the clock is not a per-container thing at all. Time synchronisation is the host's job. 3. **Anything that wants to load kernel code or reshape the host's networking.** Agents ported from a host installation hit this immediately, because on a host they had it. The low-port case has three honest answers, and it is worth knowing all three. Listen on an unreserved port inside and let the hop that publishes the service present the low port externally - usually the cleanest, because the port a client sees is a publishing decision, not a property of the process. Or configure the host so the reserved range starts lower, which is a host-wide change with host-wide consequences. Or grant that one privilege to that one workload. ## Asking for one privilege back Platforms expose a way to add a single privilege to a workload's declaration, and that is the right granularity. The cost of one extra privilege is bounded and reviewable; a reviewer can ask what it is for and what a compromise of that workload would then reach. The temptation is the blanket alternative, which restores everything at once and usually disables the confinement layers as well. It always works, and it works precisely because it stops enforcing the thing that would have told you which privilege you were missing. ## What this rule is not It is not an argument that running as root inside is fine because the set is trimmed. The default subset is a floor under the damage, not a posture; the case for running the workload as an ordinary unprivileged account is separate and stands on its own. It is also not a fixed list to memorise. The exact default subset varies between runtimes and between configurations, and asserting a specific one as universal is how people get surprised. Carry the rule - host-wide out, container-scoped in - and check the specific environment when a refusal matters.
- A service must be reachable on a reserved low port. What are the options inside a container?Listen on an unreserved port inside and let the hop that publishes the service present the low port externally - the port a client sees is a publishing decision. Failing that, grant the workload that single privilege, or change the host's reserved-range threshold, which affects everything on the machine. The first option is preferred because it changes nothing about the workload's authority.
- Why is a fixed list of the default privileges a bad thing to memorise?The exact subset differs between runtimes and between configurations, and a platform can trim it further by policy. The rule generalises and the list does not: effects that reach the shared host are withheld, effects bounded by the container are kept. Check the environment when a specific refusal matters.
saying these in an interview costs you the question
- Treats root as one switch that is either on or off
- Assumes the default privilege set is identical across every runtime
- Thinks a refusal on a host-wide operation indicates a broken image
- Restores the whole set instead of requesting the one privilege needed
- Believes a trimmed privilege set makes running as inside-root a good posture
- Expects time synchronisation to be a per-container concern