skip to content

Container Security Posture

What keeps a compromised workload from becoming a compromised host: what it runs as, what the kernel will let it call, and what is allowed to start at all. Every default here is the loose one.

on this pageshow

questions

22

A workload spec asks to run a container in privileged mode — what does that switch off, and what still applies?

level: juniorimportance: must knowfreq 68%

answer

  1. not a permission, an off switch
  2. restraints go, views stay
  3. full privilege set plus host devices
  4. call allowlist and profile dropped
  5. host disk is one mount away

basics

~20 s

Privileged mode hands the container the host's full privilege set, reachable host devices and no confinement profile, so the process can reconfigure the machine. The per-resource views and the resource ceilings usually remain, which is why it still looks contained.

solid answer

~50 s

Privileged mode is not a permission, it is the boundary's off switch. The container runs with the host's full privilege set instead of the trimmed default, the host's device nodes become reachable, and the default allowlist of kernel calls and the mandatory access-control profile are dropped. Platforms word it differently and include slightly different things, but every version of it includes enough to mount the host's own disk, change kernel tunables or load code into the kernel. What usually stays is the cosmetic half: the container still gets its own process table, mount table and network view, and its CPU and memory ceilings still apply. So it appears as an ordinary workload in every listing while being one command away from owning the host — which is why a spec asking for it is a request for administrative access to the machine.

go deeper

for a junior

Recall the one-line version: privileged mode gives the container the host's full privilege set and reachable host devices, which is practically the same as administrative access to the machine.

for a middle

Explain the asymmetry — the restraints are dropped while the per-resource views and the resource ceilings stay — and why that makes a privileged container indistinguishable from a normal one in ordinary listings.

for a senior

Show the review move: find the single operation that was actually denied, grant that narrowly, refuse the blanket mode, and state the fleet multiplier when the workload runs on every host.

for a principal

Own the standard: which classes of workload may ever hold this grant, who may change their specs and images, and how the exception list is kept short enough that it is still read.

## The two halves of a container boundary A container is an ordinary process that the kernel has been asked to treat specially in two different ways, and privileged mode touches only one of them. - **The views.** The process gets its own process table, mount table, network stack, user-id mapping and hostname, so it sees a small private version of the machine instead of the whole one. - **The restraints.** The process runs with a trimmed privilege set rather than the host's full one, with an allowlist of the kernel calls it may make, with a mandatory access-control profile saying which files and devices it may open, and with ceilings on the CPU and memory it may consume. Privileged mode is a single switch that removes most of the second half and leaves the first half in place. That asymmetry is the whole point of the question: **the container keeps looking like a container**. ## What the switch removes | Restraint | Ordinary container | Privileged container | |---|---|---| | Privilege set | trimmed to a small default | the host's full set | | Device reachability | a handful of pseudo-devices | the host's device nodes, disks included | | Kernel-call allowlist | a default allowlist applied | typically dropped | | Access-control profile | a default profile applied | typically dropped | Platforms differ on the exact contents of the mode, and that variation does not change the review: every implementation of it includes enough privilege to reconfigure the machine it runs on. ## What the switch does not remove The views survive, and so do the ceilings. A privileged container still lists only its own processes, still has its own address, still gets killed by the kernel for exceeding its memory ceiling, still shows a container's name in the platform's inventory. Monitoring that reads those views reports it as a normal workload, because from inside the views it *is* one. This is why the mode is dangerous in a way a junior candidate often misses. Nothing visibly changes. The container does not suddenly appear in the host's process listing as something special, and no alert fires because a ceiling was breached. The difference is entirely in what the process is *allowed to attempt*. ## Why full privilege plus devices equals host administrator Any one of these is enough, so a reviewer never has to pick which one an attacker would use: 1. **The host's storage is a reachable device.** With the privilege to mount, the container can mount the host's own filesystem inside itself and write to anything on it — the binaries the host runs, the table of scheduled tasks, the trusted-keys file of the administrative account. 2. **Kernel tunables and loadable code.** Full privilege includes changing kernel settings and, where the route is reachable, loading code into the kernel — which is under every boundary the platform could have drawn. 3. **Stepping into another view.** With enough privilege a process can enter the views belonging to other processes on the machine, which collapses the first half of the boundary as well. The practical statement to make in an interview is: *granting privileged mode and granting an administrative account on that host are the same grant with different paperwork*. ## Reviewing a request for it - Ask what the workload actually needs. "It only needs it to read X" is the standard opening, and the mode has no smaller version — if a specific resource is needed, grant that resource narrowly instead. - A resource ceiling is not a mitigation here. It caps consumption, not privilege, and a privileged container that rewrites the host does not need much memory to do it. - "The image is internal and trusted" moves the question rather than answering it: the grant is then only as strong as the controls on who may change that image and that spec. - **A workload placed on every host multiplies the grant by the fleet.** One copy per machine means one compromise is every machine. ## When it is honest Some workloads genuinely manage the machine — a storage driver, a device manager, a low-level host agent. For those the mode may be the right answer, and the control moves elsewhere: keep that set of workloads small and explicitly named, know who can change their specs and images, and keep them off hosts that do not need them. What is never honest is a general-purpose workload that acquired the mode because something failed once and the mode made the failure go away.

  • If the views and the ceilings still apply, in what sense is the container still contained?
    Only cosmetically. The views keep it tidy — its own process table, its own addresses, its own mount table — and the ceilings keep it from starving the host of CPU and memory. Neither restrains privilege, and privilege is what an escape needs. The container is contained against accidents and not against intent.
  • A workload failed with a permission error and adding privileged mode fixed it. What should the review ask?
    Which single operation was denied. The mode is a blanket grant used as a diagnostic, and the answer is almost always one specific device, one file path or one narrow privilege. Reproduce the failure, identify that one thing, grant it, and remove the mode — otherwise the workload keeps a permanent grant because of one transient error.
  • Does the mode change what other containers on the same host can do?
    Not directly — their own specs are unchanged. Indirectly it changes everything, because a privileged container can reach the host, and from the host it can reach the files, the configuration and the running processes of every other workload on that machine. The blast radius is the host, not the workload.

saying these in an interview costs you the question

  • Says privileged mode just means the process runs as the administrative account inside
  • Claims a privileged container has no separate process or network view at all
  • Treats it as a performance setting that heavy workloads need
  • Assumes CPU and memory ceilings limit what a privileged container can damage
  • Argues it is acceptable because the image was built by an internal team
open as a page

After its root filesystem is made read-only, a report renderer cannot write scratch files - what must be declared?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A read-only root filesystem refuses every write except into paths declared writable, so the renderer needs one mounted at its scratch directory and its temporary-file setting pointed there. Declare the size and the owning id too, or it breaks again.

open as a page

A workload runs as the image's default superuser account - what does that actually cost you?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Running as the superuser hands any code-execution bug full in-container privilege: it can rewrite the program files the running process reads, open every mounted credential regardless of ownership, and exploit whatever the boundary fails to close. An unprivileged, explicitly declared numeric id removes that head start.

open as a page

Why does mounting the container runtime's control socket into a container hand that container administrative access to the host?

level: middleimportance: must knowfreq 54%

basics

~20 s

Anything that can reach the runtime's control socket can ask it to create a privileged container with the host's filesystem mounted in, and the runtime obeys as a host-level process. No privilege inside the calling container is needed.

open as a page

Where does an admission gate sit when a workload spec is submitted to a cluster, and what can it do to that request?

level: middleimportance: must knowfreq 55%

basics

~20 s

An admission gate runs after a workload spec is submitted and before the platform stores or schedules it. It reads the spec and the image reference the spec names, then admits the request or refuses it outright.

open as a page

A container runtime applies a default syscall allowlist — what happens to a call it does not permit, and why do most workloads never notice?

level: middleimportance: must knowfreq 62%

basics

~20 s

A call outside the allowlist is refused in the kernel before it runs, so the process gets an ordinary error return from deep inside a library. Most workloads only ever make the common calls a shipped default already permits.

open as a page

How does a nightly job prove which workload it is to a data service outside the cluster, with no stored password?

level: middleimportance: must knowfreq 58%

basics

~20 s

The platform mints the job a short-lived token that names the workload itself, addressed to one audience and valid for minutes. The job presents it to the far service, which verifies the issuer's signature and grants only what that workload's grant allows.

open as a page

When the verifier an admission gate calls is unreachable, what does failing open cost and what does failing closed cost?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Failing open admits unverified images exactly when nobody is watching. Failing closed admits nothing new, so rollouts, replacements and scale-ups stop while replicas already running keep serving. Shared clusters usually choose closed and then make it cheap.

open as a page

Why does a credential attached to the node grant far more than one issued to a single workload?

level: seniorimportance: must knowfreq 50%

basics

~20 s

A node credential belongs to the machine, so its grant must cover everything any workload placed there needs, and anything running on that host can reach it. A per-workload token names one workload and carries only that workload's access.

open as a page

A per-node agent's spec asks to share the host's process and network views — what does the boundary stop hiding?

level: middleimportance: should knowfreq 41%

basics

~20 s

Sharing the host's process view exposes every process on the machine — their command lines, and their environment where the account allows it. Sharing the network view puts the container on the host's addresses, reaching services that only expected host-local callers.

open as a page

What does dropping a container's default privilege set add when the workload already runs under an unprivileged account?

level: middleimportance: should knowfreq 45%

basics

~20 s

Who the process is and what powers it holds are separate axes. Runtimes hand every container a default slice of superuser powers, held by the process rather than its account. Dropping them all and adding back only what is needed closes that second axis.

open as a page

Switching a container to a declared unprivileged id makes files the image shipped unreadable - why, and what fixes it?

level: middleimportance: should knowfreq 50%

basics

~20 s

The build ran as the superuser, so the files it installed are owned by id 0 with bits that grant access only to that owner. At runtime the new id matches neither owner nor group and falls to the least-permissive bits. Fix ownership at build time.

open as a page

How does a file-and-device reachability profile differ from a syscall allowlist when both can refuse the same workload?

level: middleimportance: should knowfreq 38%

basics

~20 s

They answer different questions. The allowlist decides which kernel calls may be attempted at all; the reachability profile decides which files and devices a call is allowed to touch. Widening one leaves the other exactly as it was.

open as a page

A per-node log shipper's spec asks for privileged mode, a writable host-root mount, the runtime's control socket and the host's process view — how do you rank and narrow them?

level: seniorimportance: should knowfreq 49%

basics

~20 s

Privileged mode, a writable host-root mount and the control socket are each already administrative access to that machine — one tier, where one grant is the whole loss. Narrow them to a read-only mount of the log directory.

open as a page

A cluster gate admits any image whose reference names the company's own registry host — what does that rule actually enforce?

level: seniorimportance: should knowfreq 42%

basics

~20 s

An approved-source rule enforces where the bytes will be fetched from, not what the bytes are — it is a match on the reference string. Anyone able to publish under that host, or to copy a foreign image there, passes it.

open as a page

Which confinement wall explains a thumbnail worker that fails with a permission error inside its decoding library, on files that are readable?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Most often the syscall allowlist. A refused kernel call is returned to the caller as a plain permission error, so it surfaces as a library failure on an operation that has nothing to do with file permissions.

open as a page

An attacker captures a workload's ten-minute token — what does the short lifetime actually limit, and what does it not?

level: seniorimportance: should knowfreq 36%

basics

~20 s

A short lifetime kills the whole class of credentials found later in an image, a log or a backup, because a captured copy is already dead. It does not limit use inside the window, anything traded for it, or an attacker still executing in the workload and reading each refreshed token.

open as a page

A cluster's admission gate is blocking the fix for a worsening outage, and someone asks you to switch it off — how do you answer?

level: principalimportance: should knowfreq 32%

basics

~20 s

Find what the gate is refusing first: an unapproved source, a missing verification result and a spec asking for privilege are three different problems. Then unblock the narrowest one, time-boxed — not the whole gate.

open as a page

When the platform's default confinement filter blocks a team's workload, what standard would you set for relaxing it?

level: principalimportance: should knowfreq 34%

basics

~20 s

Require the request to name the refused call, not the wall. Then permit that one call in confinement scoped to that workload, with an owner, an expiry and a review — and treat switching the filter off as an exception, not the next rung down.

open as a page

Why can a workload spec with no privileged mode, no host mounts and no shared views still end in host compromise?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

Three escape routes are grants written in the spec; the fourth is a defect in the shared kernel, reached by an ordinary call from inside the container. It needs no grant, appears in no spec, and a spec review cannot see it.

open as a page

Under an unprivileged account, how can a process still gain privilege when it executes another program?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Images can ship program files marked to run with their owner's identity or extra powers rather than the caller's, so executing one raises privilege even from an unprivileged account. A declaration that forbids any privilege gain at exec blocks that, permanently and for every child.

open as a page

Making every outward call depend on a platform token issuer adds a failure domain — how do you design for it?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Design around the one property that makes it survivable: held tokens stay valid to their own expiry, so an issuer outage degrades over one lifetime and hits newly started workloads first. Token lifetime is the lever, traded directly against a stolen token's window.

open as a page