A workload spec asks to run a container in privileged mode — what does that switch off, and what still applies?
answer
- not a permission, an off switch
- restraints go, views stay
- full privilege set plus host devices
- call allowlist and profile dropped
- host disk is one mount away
basics
~20 sPrivileged mode hands the container the host's full privilege set, reachable host devices and no confinement profile, so the process can reconfigure the machine. The per-resource views and the resource ceilings usually remain, which is why it still looks contained.
solid answer
~50 sPrivileged mode is not a permission, it is the boundary's off switch. The container runs with the host's full privilege set instead of the trimmed default, the host's device nodes become reachable, and the default allowlist of kernel calls and the mandatory access-control profile are dropped. Platforms word it differently and include slightly different things, but every version of it includes enough to mount the host's own disk, change kernel tunables or load code into the kernel. What usually stays is the cosmetic half: the container still gets its own process table, mount table and network view, and its CPU and memory ceilings still apply. So it appears as an ordinary workload in every listing while being one command away from owning the host — which is why a spec asking for it is a request for administrative access to the machine.
go deeper
Recall the one-line version: privileged mode gives the container the host's full privilege set and reachable host devices, which is practically the same as administrative access to the machine.
Explain the asymmetry — the restraints are dropped while the per-resource views and the resource ceilings stay — and why that makes a privileged container indistinguishable from a normal one in ordinary listings.
Show the review move: find the single operation that was actually denied, grant that narrowly, refuse the blanket mode, and state the fleet multiplier when the workload runs on every host.
Own the standard: which classes of workload may ever hold this grant, who may change their specs and images, and how the exception list is kept short enough that it is still read.
## The two halves of a container boundary A container is an ordinary process that the kernel has been asked to treat specially in two different ways, and privileged mode touches only one of them. - **The views.** The process gets its own process table, mount table, network stack, user-id mapping and hostname, so it sees a small private version of the machine instead of the whole one. - **The restraints.** The process runs with a trimmed privilege set rather than the host's full one, with an allowlist of the kernel calls it may make, with a mandatory access-control profile saying which files and devices it may open, and with ceilings on the CPU and memory it may consume. Privileged mode is a single switch that removes most of the second half and leaves the first half in place. That asymmetry is the whole point of the question: **the container keeps looking like a container**. ## What the switch removes | Restraint | Ordinary container | Privileged container | |---|---|---| | Privilege set | trimmed to a small default | the host's full set | | Device reachability | a handful of pseudo-devices | the host's device nodes, disks included | | Kernel-call allowlist | a default allowlist applied | typically dropped | | Access-control profile | a default profile applied | typically dropped | Platforms differ on the exact contents of the mode, and that variation does not change the review: every implementation of it includes enough privilege to reconfigure the machine it runs on. ## What the switch does not remove The views survive, and so do the ceilings. A privileged container still lists only its own processes, still has its own address, still gets killed by the kernel for exceeding its memory ceiling, still shows a container's name in the platform's inventory. Monitoring that reads those views reports it as a normal workload, because from inside the views it *is* one. This is why the mode is dangerous in a way a junior candidate often misses. Nothing visibly changes. The container does not suddenly appear in the host's process listing as something special, and no alert fires because a ceiling was breached. The difference is entirely in what the process is *allowed to attempt*. ## Why full privilege plus devices equals host administrator Any one of these is enough, so a reviewer never has to pick which one an attacker would use: 1. **The host's storage is a reachable device.** With the privilege to mount, the container can mount the host's own filesystem inside itself and write to anything on it — the binaries the host runs, the table of scheduled tasks, the trusted-keys file of the administrative account. 2. **Kernel tunables and loadable code.** Full privilege includes changing kernel settings and, where the route is reachable, loading code into the kernel — which is under every boundary the platform could have drawn. 3. **Stepping into another view.** With enough privilege a process can enter the views belonging to other processes on the machine, which collapses the first half of the boundary as well. The practical statement to make in an interview is: *granting privileged mode and granting an administrative account on that host are the same grant with different paperwork*. ## Reviewing a request for it - Ask what the workload actually needs. "It only needs it to read X" is the standard opening, and the mode has no smaller version — if a specific resource is needed, grant that resource narrowly instead. - A resource ceiling is not a mitigation here. It caps consumption, not privilege, and a privileged container that rewrites the host does not need much memory to do it. - "The image is internal and trusted" moves the question rather than answering it: the grant is then only as strong as the controls on who may change that image and that spec. - **A workload placed on every host multiplies the grant by the fleet.** One copy per machine means one compromise is every machine. ## When it is honest Some workloads genuinely manage the machine — a storage driver, a device manager, a low-level host agent. For those the mode may be the right answer, and the control moves elsewhere: keep that set of workloads small and explicitly named, know who can change their specs and images, and keep them off hosts that do not need them. What is never honest is a general-purpose workload that acquired the mode because something failed once and the mode made the failure go away.
- If the views and the ceilings still apply, in what sense is the container still contained?Only cosmetically. The views keep it tidy — its own process table, its own addresses, its own mount table — and the ceilings keep it from starving the host of CPU and memory. Neither restrains privilege, and privilege is what an escape needs. The container is contained against accidents and not against intent.
- A workload failed with a permission error and adding privileged mode fixed it. What should the review ask?Which single operation was denied. The mode is a blanket grant used as a diagnostic, and the answer is almost always one specific device, one file path or one narrow privilege. Reproduce the failure, identify that one thing, grant it, and remove the mode — otherwise the workload keeps a permanent grant because of one transient error.
- Does the mode change what other containers on the same host can do?Not directly — their own specs are unchanged. Indirectly it changes everything, because a privileged container can reach the host, and from the host it can reach the files, the configuration and the running processes of every other workload on that machine. The blast radius is the host, not the workload.
saying these in an interview costs you the question
- Says privileged mode just means the process runs as the administrative account inside
- Claims a privileged container has no separate process or network view at all
- Treats it as a performance setting that heavy workloads need
- Assumes CPU and memory ceilings limit what a privileged container can damage
- Argues it is acceptable because the image was built by an internal team