A team fixes a failing container by turning on privileged mode - what does that switch off, and what should they ask for instead?
answer
- not one privilege - several layers
- profiles and devices go too
- it works, which hides the cause
- read the denied operation first
- grant one named privilege, then re-run
basics
~20 sPrivileged mode restores the full privilege set at once and typically disables the confinement profiles and exposes host devices, leaving little but the separate views. The disciplined alternative is to read the denied operation and request that single privilege.
solid answer
~50 sThe blanket mode is not one more privilege; it is the removal of several layers together. It gives the workload back every privilege the runtime had trimmed, usually stops applying the syscall allowlist and the mandatory access-control profile, and commonly exposes the host's devices to the container. What is left of the boundary is mostly the separate views of filesystem, process table and network - the layers that were never meant to carry the weight alone. It also works on the first try, which is the real damage: it removes the refusal that would have named the missing privilege. The alternative is a procedure, not a preference - read the denied **operation**, decide whether it is host-wide, and add that one privilege to the workload declaration so a reviewer can see what was granted and why.
code
pseudocode · 21 lineson a refused operation inside a container:
read the denied OPERATION, not the id the process reports
if operation is host-wide (clock, kernel code, host network, mount):
-> the privilege was withheld deliberately
-> add that one privilege to the workload declaration
-> re-run; if a second refusal appears, repeat
else if operation is binding a reserved low port:
-> listen on an unreserved port, publish the low one in front
-> or add that single privilege
else if operation is a read or write on a path from the host:
-> this is ownership / id mapping, not privilege
-> granting a privilege changes nothing here
blanket privileged mode:
-> restores every privilege AND drops the confinement profiles
-> also removes the refusal that names the cause
-> reserve it for software whose job is managing the hostgo deeper
Know that this setting is not a small step: it restores the whole privilege set and usually drops the confinement profiles too. Treat seeing it in a declaration as something to ask about.
Explain the layers it removes together, and why its immediate success is a problem - it also removes the refusal that would have named the one missing privilege.
Show the procedure: reproduce with the boundary intact, classify the denied operation, grant one privilege by name, re-run, and treat a growing list as evidence the workload belongs elsewhere.
The policy question is how your estate distinguishes host-managing software that legitimately needs broad authority from application workloads that acquired a blanket grant once and kept it.
## What the blanket switch actually does Most platforms expose a single setting that means "stop holding this workload back". Reaching for it feels like adding a privilege. It is not: it is switching off several independent protections at the same time. - **The trimmed privilege set is restored in full.** Every host-wide operation the runtime had withheld - loading kernel code, setting the clock, reconfiguring the host's networking, mounting filesystems - becomes available again. - **The confinement profiles usually stop applying.** The syscall allowlist that narrows which kernel entry points the workload may use, and the mandatory access-control profile that constrains what it may touch, are commonly dropped along with it. - **Host devices are typically exposed.** The container gets a device view close to the host's rather than the minimal set a workload normally sees. What remains is mostly the separate views of filesystem, process table and network. Those were designed to be one layer of several, not the only one. ## Why it is so tempting, and why that is the problem The sequence is always the same. A workload fails with a refusal nobody has time to decode. Someone flips the blanket setting. It works immediately. The ticket closes. Notice what was lost in that minute: the refusal was the only artefact that named the missing privilege. Turning off the enforcement also turns off the diagnosis, so nobody ever learns whether this workload needed one narrow privilege or three broad ones. The setting then persists through every later deployment, because removing something that is apparently load-bearing and definitely working is nobody's Tuesday. The second-order cost is reviewability. One added privilege is a line a reviewer can ask about. A blanket grant is a line that answers nothing: it tells the reviewer the workload needed *something*. ## The targeted alternative, as a procedure 1. **Reproduce the failure with the boundary intact** and capture the exact operation that was refused, not just the fact that start-up failed. 2. **Classify the operation.** Host-wide (clock, kernel code, host network configuration, mounting) means a deliberately withheld privilege. A read or write on a path from the host means an ownership or id-mapping problem instead, and no privilege will fix it. 3. **Add the single privilege** the operation needs to the workload declaration, and nothing else. 4. **Re-run and look for the next refusal.** If a second appears, repeat. Two or three named privileges are a normal outcome and are still far better than the blanket grant. 5. **If the list keeps growing**, that is information: the workload's job may genuinely be managing the host, and it should be treated as host software rather than as an ordinary workload. ## When a broad grant is still the honest answer Some workloads exist to manage the machine - an agent that attaches storage, one that configures host networking, one that collects low-level metrics. These genuinely need what the boundary is built to withhold, and pretending otherwise produces a long list of privileges that adds up to the same thing with more paperwork. The honest handling is to stop calling them ordinary workloads: run them from a small, separately reviewed set, keep them off the same declaration path as application workloads, and be explicit that they are part of the host's trusted software. The failure mode to avoid is a fleet where an application service carries a blanket grant because of one afternoon in its history, indistinguishable from the agent that legitimately needs it. ## Two things this is not It is not a claim that the blanket mode is the only way a boundary is weakened - it is simply the one a team turns on themselves, on purpose, in a workload declaration anyone can read. And it is not a substitute for asking whether the workload should have been in a container at all. A workload that needs most of the host's authority is telling you that the fence is not the right tool for it, and the interesting decision is where it runs, not which settings it carries.
- A workload declaration you inherit carries the blanket grant. How do you decide whether it is needed?Remove it in a non-production copy and reproduce the failure with the boundary intact, so the refusal names the operation. Classify what comes back: if it is one or two host-wide operations, grant those by name; if nothing fails at all, the grant was history rather than a requirement, which is the common outcome.
- Why is granting three named privileges better than the blanket grant that covers them?Each name is reviewable - someone can ask what it is for, and reason about what a compromise of that workload would reach. The blanket grant also drops the syscall allowlist and the access-control profile, which the three named privileges leave in place, so the two are not equivalent even when the privileges overlap.
- The list of needed privileges keeps growing on one workload. What is that telling you?That its job is probably managing the host rather than serving traffic. Treat it as host software: a small, separately reviewed set kept off the ordinary workload path, so it is not indistinguishable from an application service that acquired a broad grant by accident.
Instead of getting a key cut for the one door that stopped you, someone props open every door in the building and disables the alarm. It certainly ends the interruption, and nobody ever finds out which door it was.
saying these in an interview costs you the question
- Describes the blanket mode as adding one more privilege
- Believes the confinement profiles still apply underneath it
- Leaves it enabled because removing it might break something
- Cannot say which single operation the refusal was about
- Thinks the separate views alone make the remaining boundary equivalent
- Grants it to every workload in an environment to save time