skip to content

Container Isolation Model

What the boundary around a container fences off, what caps it, and what every workload on a host still shares. Treating it as a small machine is where density and security reasoning both break.

on this pageshow

questions

page 1 of 2

When a runtime starts a container, which separate views of the system does it fence off, and what stays shared?

level: juniorimportance: must knowfreq 72%

answer

  1. not one feature, several at once
  2. views hide, the group caps
  3. filesystem, processes, network, identity, channels
  4. first process is number one inside
  5. one shared kernel underneath all of them

basics

~20 s

A container is several fences applied at once: its own root filesystem, process table, network stack, host identity and inter-process channels, plus an accounting group that caps what it consumes. The kernel underneath stays shared.

solid answer

~50 s

There is no single feature that makes a container. A runtime starts an ordinary process and, before running your program, replaces several of the views that process has of the machine: a **root filesystem** assembled from the image, a **process table** of its own in which its first process is number one, a **network** view with its own interfaces and port space, its own **host identity**, its own **inter-process channels**, and optionally its own **user-id range**. Separately, it puts those processes in an **accounting and limiting group** that measures and caps CPU time and memory. The split is the whole point: the views *hide*, the group *caps*, and neither does the other's job. What is not fenced is the kernel itself — one version, one module set, one set of host-wide tunables for every workload on the box.

code

pseudocode · 18 lines
pseudocode
start_container(image, spec):
    child = fork_process()

    # each fence is applied independently, and any may be skipped or joined
    if spec.fence_filesystem: child.root       = assemble_root_from(image)
    if spec.fence_processes:  child.processes  = new_process_table()
    if spec.fence_network:    child.network    = new_network_view()
    if spec.fence_identity:   child.hostname   = spec.instance_name
    if spec.fence_channels:   child.ipc        = new_channel_set()
    if spec.join_processes_of is not none:
        child.processes = processes_of(spec.join_processes_of)   # deliberate hole

    # a separate mechanism, doing a different job: accounting, not hiding
    group = new_accounting_group(cpu_quota = spec.cpu_ceiling,
                                 memory_max = spec.memory_ceiling)
    group.add(child)

    child.run(spec.command)      # first process inside: number one

go deeper

for a junior

Be able to name the fences in plain words — its own filesystem, process list, network and host identity — and add that a separate mechanism caps CPU and memory. Naming four of them confidently is worth more than a vague 'it is isolated'.

for a middle

Explain why the composition is separable: a runtime can omit a fence or let a container join another's, and that flexibility is the reason a boundary can be strong in one dimension and absent in another. Keep hiding and capping in separate sentences.

for a senior

Reason about what the composition is worth in production: which fences your platform actually applies, which are commonly left off, and what a shared kernel means for blast radius when you are deciding whether this boundary is the only one a workload gets.

for a principal

Treat the fence set as a policy you own rather than a default you inherit. The trade-off is density and simplicity against the fact that one shared kernel is one failure domain, and that argument decides which workloads may share a host at all.

## A container is a composition, not a single feature The kernel has no object called a container. What a container runtime does is start an ordinary process and, in the moment before that process begins running your program, change several independent things about the world it will see, then place it in a group that accounts for what it consumes. Each change covers one kind of resource. "A container" is the name we give the bundle when they are applied together. That this is a composition, rather than one switch, is not trivia. It explains three behaviours you will meet in production: a runtime can **leave a fence out**, a second container can be started so that it **joins a fence the first one already has**, and a workload can be perfectly fenced in one dimension while being completely exposed in another. ## The fences a runtime composes - **Root filesystem.** The process sees a root directory assembled from the image's layers rather than the host's. Host paths simply do not exist for it unless someone explicitly mounts them in. - **Process table.** The container gets its own numbering. The first process it starts is number one inside, other containers' processes and the host's are invisible, and a process it cannot see is a process it cannot signal or inspect. - **Network.** Its own interfaces, its own loopback and its own port space, which is why two containers on one host can both listen on the same port number without colliding. - **Host identity.** Its own hostname, normally set per instance by the runtime rather than inherited from the machine. - **Inter-process channels.** Its own shared-memory segments and message queues, so two workloads cannot accidentally attach to each other's. - **User and group ids (often optional).** A runtime *may* also map the ids inside the boundary onto a different range outside it, so the id a process holds inside need not be the id it holds on the host. Many setups do not turn this one on, which is exactly why it is worth asking about. ## Hiding is not capping None of those views limits consumption. A container with a perfectly fenced filesystem, process table and network can still burn every processor on the host. Limiting is the job of a second, entirely separate mechanism: an **accounting and limiting group** that the runtime creates, places the container's processes into, and configures. That group measures usage and enforces ceilings, and the two ceilings behave in opposite ways: - a **CPU ceiling** is a quota of run time per period, so a process that wants more is made to wait — it slows down and stays alive; - a **memory ceiling** is a hard wall with no equivalent of waiting, so when it is reached the kernel ends a process instead. The symmetric mistake is to expect the group to hide things. It does not. It caps consumption without changing a single thing the process can read about the machine, which is the root of a whole family of start-up bugs. ## What stays shared | Fenced per container | Shared by every container on the host | |---|---| | the root filesystem it sees | the kernel itself, its version and its loaded modules | | the process table and its numbering | host-wide tunables the kernel exposes once per machine | | network interfaces, addresses and ports | the hardware, its firmware and its device state | | the host identity it reports | the time of day the host keeps | | inter-process channels | the scheduler and memory subsystem doing the actual work | The right one-line summary of the second column is that every container on a host makes its system calls into **one** kernel. A container does not have its own operating system; it ships a userland — the files, libraries and tools baked into its image — and that userland talks to the same kernel as everything else on the box. ## The composition is optional, fence by fence 1. A runtime can be told to **skip** a fence. Skip the network fence and the workload sees the host's interfaces and port space as its own. 2. A runtime can start a container that **joins** an existing container's view rather than getting a fresh one. This is how two containers are made to look like one machine to each other — a deliberate hole, chosen for a reason. 3. Skipping or joining is an **isolation decision**, not a performance tuning knob. Each one you drop is one less thing standing between that workload and everything else on the host. ## What interviewers are listening for - That you can name several fences rather than saying "it is isolated". - That you keep hiding and capping apart, because almost every follow-up in this area turns on the distinction. - That you know something is still shared, and can say what — the honest answer to "how strong is this boundary?" starts there.

  • Two containers are started so that one joins the other's process-table view. What does each gain and lose?
    They gain mutual visibility: either can list, signal and attach to the other's processes, which is what makes a companion process able to watch or manage the main one. They lose the fence between them — a compromise in either is now a compromise of both. The other fences are unaffected: each still has its own root filesystem unless a mount is shared too.
  • If the fences hide and the accounting group caps, what does neither of them do?
    Neither gives the container a kernel of its own. One kernel version, one module set and one set of host-wide tunables serve every container on the machine, so a kernel-level fault is a host-wide event rather than a per-container one. That is the ceiling on how strong this boundary can be, no matter how many fences you compose.
  • Does a container have its own operating system?
    No. It carries its own userland — the files, libraries, shells and tools its image was built with, which is why one container can carry a completely different set of system tools from its neighbour. But every system call it makes is served by the host's single shared kernel, so it is one operating system with many userlands, not many operating systems.

Less a locked room than a set of one-way mirrors fitted one at a time, with a meter on the door: the mirrors decide what the occupant can see, and the meter alone decides how much it may draw.

saying these in an interview costs you the question

  • Calls a container a lightweight virtual machine that boots its own kernel.
  • Believes one single kernel feature creates the whole boundary.
  • Thinks the fences also limit how much CPU and memory the workload may use.
  • Assumes every fence is always applied, so nothing at all is shared.
  • Says the image ships an operating system rather than a userland.
open as a page

A container passes its CPU ceiling, and later its memory ceiling — what happens to the process each time?

level: juniorimportance: must knowfreq 84%

basics

~20 s

Passing the CPU ceiling only slows a process: it is throttled, waits for the next period's quota, and stays alive. Passing the memory ceiling ends it — there is no memory throttle, so the kernel kills a process inside the boundary.

open as a page

A worker reports as root inside its container but cannot set the host clock - why is root inside not root on the host?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Root inside a container is an id read in that container's own view, and the runtime already trimmed its privilege set, so host-wide powers like setting the clock are missing. Many setups also map that id onto an unprivileged host id.

open as a page

What does a container share with its host that a virtual machine brings its own, and what does that buy at start-up?

level: juniorimportance: must knowfreq 88%

basics

~20 s

A container shares the host's already-running kernel; a virtual machine boots a guest kernel and a whole system of its own from a disk image. Sharing removes the boot, so a container starts a process in milliseconds instead of booting a machine.

open as a page

Why does a containerized service reading the machine's CPU and memory totals see the host's numbers rather than its own ceiling?

level: middleimportance: must knowfreq 60%

basics

~20 s

The fences hide other workloads; they do not rewrite the interfaces that report machine capacity. Those numbers come from the shared kernel's host-wide view, while the container's ceiling lives in an accounting group nobody told the process about.

open as a page

A host's workloads reserve 40% of its memory while their declared ceilings sum to 200% of it — what is that overcommit betting on?

level: middleimportance: must knowfreq 55%

basics

~20 s

It bets that the workloads will not peak at the same time. Placement only subtracts reservations from capacity, so ceilings may sum far past what the host owns. Density is the payoff; coincident peaks are the unpaid bill.

open as a page

What is the difference between a resource reservation the scheduler places against and the ceiling enforced at runtime?

level: middleimportance: must knowfreq 68%

basics

~20 s

A reservation is a placement input: the scheduler subtracts it from a node's free capacity when deciding where a workload fits. A ceiling is enforced on the running process by the kernel's accounting mechanism. Different actors, different moments, different failures.

open as a page

When a container starts, what does the container manager prepare, and what does the low-level runtime actually do with it?

level: middleimportance: must knowfreq 68%

basics

~20 s

Two components split the work. A container manager pulls and unpacks an image into a root filesystem and writes a configuration document beside it; a low-level runtime then applies the fences named in that document and executes the declared command as the first process inside the boundary.

open as a page

A container with strict CPU and memory ceilings triggers a fault in the host's shared kernel — what happens to its neighbours?

level: middleimportance: must knowfreq 58%

basics

~20 s

They stop too. A ceiling bounds how much of a resource a workload consumes; it does not bound the kernel running on that workload's behalf, and one kernel serves every container on the machine, so a panic takes all of them down together.

open as a page

A vendor library needs kernel features newer than the host's shared kernel — can the container image supply them?

level: middleimportance: must knowfreq 65%

basics

~20 s

No. An image carries only userspace: binaries, libraries and files. Every container on a machine executes against the one kernel that machine booted, so a newer kernel version, or a feature from an unloaded module, has to come from the host.

open as a page

A host fits tens of virtual machines of one service but hundreds of containers — which per-instance costs disappeared?

level: middleimportance: must knowfreq 66%

basics

~20 s

Everything the guest model duplicates per instance goes away: a resident kernel, a first process and system services, the memory set aside for the guest up front, and a full disk image. What remains per container is roughly the workload's own working set.

open as a page

A gateway's p99 doubles nightly while it and a co-located batch job stay inside every processor and memory ceiling — what are they fighting over?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Everything the ceilings do not divide: one pool of cached file pages, one storage device and its queue, one network link, and the shared path to memory. Ceilings partition two axes and certify nothing about the rest.

open as a page

An image runs unchanged after its hosts' container stack was replaced — which two published specifications make that possible?

level: juniorimportance: should knowfreq 56%

basics

~20 s

Two agreements carry it: an image specification fixing what an image is — content-addressed layers plus a recorded configuration of command, environment and architecture — and a runtime specification fixing what a runtime is handed and what it must do with it.

open as a page

Root inside a container cannot load kernel code or set the clock, yet it can change file ownership - what decides the difference?

level: middleimportance: should knowfreq 48%

basics

~20 s

Root's authority is a set of separately grantable privileges, and a runtime hands a workload a small default subset. Privileges whose effect reaches the shared host are withheld; privileges whose effect ends inside the container's own view are kept.

open as a page

A rootless worker writes into a bind-mounted host directory, and the new files show a high, unowned id - why?

level: middleimportance: should knowfreq 55%

basics

~20 s

Under id remapping the container's ids sit on a range of unprivileged host ids, so a file created inside as id 0 lands on the host owned by the first id of that range - a number no host account claims, which is why tools print it raw.

open as a page

A workload needs a kernel version or module your hosts do not provide — why can a container not supply it?

level: middleimportance: should knowfreq 46%

basics

~20 s

A container image holds user-space files only; the kernel serving its system calls is the host's. A feature the host kernel does not have is not in the image's power to add, which is the clearest honest signal that this workload wants a machine boundary.

open as a page

A metrics agent moved into its own container can no longer see the application process it watches — which fence did that, and what fixes it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The separate process table view: each container sees only its own processes, so the agent finds nothing to watch. Either start it sharing the target's process view, collect from the host instead, or have the application publish what the agent needs.

open as a page

Why is overcommitting a host's memory a different class of risk from overcommitting its processor time?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Processor time can be shared more thinly, so a shortage slows everything and ends nothing. Memory is occupied space that cannot be split further, so the only way out of a shortage is taking it from someone.

open as a page

When would you set a latency-sensitive replica's reservation equal to its ceiling rather than below it?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Set them equal when the replica's performance must not depend on what else lands beside it: the capacity it is placed against becomes the capacity it may use. You pay for the peak continuously, which is worth it for latency commitments and not for elastic background work.

open as a page

A memory ceiling ends a service seconds after every start, though it is stable in steady state — why?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Most services demand more memory while initialising than while serving: configuration and data are loaded, caches are warmed, pools are opened, often concurrently. A ceiling sized from steady-state observation is below that transient peak, so the wall is hit before serving ever begins.

open as a page

One replica's p99 latency doubles at peak while its average CPU use sits at half its ceiling — why?

level: seniorimportance: should knowfreq 54%

basics

~20 s

A CPU ceiling is a quota per short repeating period, not a long-run average. A burst exhausts one period's quota and the process waits out the rest of that period, so requests stall in slices while the averaged number stays low.

open as a page

A team fixes a failing container by turning on privileged mode - what does that switch off, and what should they ask for instead?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Privileged mode restores the full privilege set at once and typically disables the confinement profiles and exposes host devices, leaving little but the separate views. The disciplined alternative is to read the denied operation and request that single privilege.

open as a page

A host's container management service was restarted, yet every workload kept running and kept its output stream — how?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The manager is not the workload's parent. The low-level runtime that created each container exited at start-up, and a small per-container supervising process sits between: it parents the first process, owns its output pipes and holds its exit status until the manager returns and re-attaches.

open as a page

Two workloads on one host want different values for a host-wide kernel tunable — what can the container boundary do about that?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Nothing, for a value the kernel keeps per machine. Some settings follow a container's own view of a resource and can differ per workload; a single machine-wide value is one number for everyone, so the conflict is settled by placement, not by configuration.

open as a page

As a fleet standard, how far below its ceiling should a workload be allowed to reserve, and how would you know that gap is too wide?

level: principalimportance: should knowfreq 38%

basics

~20 s

There is no single ratio. Set the gap per class of workload, priced by what that class's failure costs, with separate rules for processor time and memory. The guardrail is the victim class's latency, not utilisation.

open as a page

Every host runs one kernel, so an upgrade is a whole-machine change — how do you set a kernel version policy for a shared fleet?

level: principalimportance: should knowfreq 34%

basics

~20 s

Decide two things: how much version fragmentation the estate will carry, and how fast the whole fleet can be turned over. One version is cheapest to certify and patch; each extra host population buys one workload's requirement at a permanent cost.

open as a page

A machine boundary per run costs seconds and density — which customer-template renders still justify it?

level: principalimportance: should knowfreq 42%

basics

~20 s

Runs whose code the customer controls, on hosts that also hold other tenants' data or platform credentials, justify a machine boundary. The deciding questions are whose code it is and what else is reachable through one shared kernel — not how risky the feature feels.

open as a page

Why does a clustered search indexer that derives its node identity from the hostname rejoin as a brand-new member after every restart?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Each container gets its own host-identity fence, and the runtime sets that name per instance, so a replacement carries a different one. Identity meant to outlive an instance must be supplied as configuration, not read back from the hostname.

open as a page

A service's reads were hitting the host's cached file pages until a co-tenant began streaming a large file — why did they get slower?

level: seniorimportance: nice to knowfreq 31%

basics

~20 s

One host keeps one pool of cached file pages. A single pass over a large file fills it with pages nobody will read again, pushing the service's hot pages out, so its reads now reach the storage device.

open as a page

What has to stay the same when a host swaps its low-level runtime for a sandboxed one, and what changes?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Everything above the seam stays: the same image, the same configuration document, the same manager and the same lifecycle operations. What changes is who answers the workload's system calls — and with it start-up time, throughput and which unusual calls and devices still work.

open as a page

showing 1–30 of 31