skip to content

questions

4

When a runtime starts a container, which separate views of the system does it fence off, and what stays shared?

level: juniorimportance: must knowfreq 72%

answer

  1. not one feature, several at once
  2. views hide, the group caps
  3. filesystem, processes, network, identity, channels
  4. first process is number one inside
  5. one shared kernel underneath all of them

basics

~20 s

A container is several fences applied at once: its own root filesystem, process table, network stack, host identity and inter-process channels, plus an accounting group that caps what it consumes. The kernel underneath stays shared.

solid answer

~50 s

There is no single feature that makes a container. A runtime starts an ordinary process and, before running your program, replaces several of the views that process has of the machine: a **root filesystem** assembled from the image, a **process table** of its own in which its first process is number one, a **network** view with its own interfaces and port space, its own **host identity**, its own **inter-process channels**, and optionally its own **user-id range**. Separately, it puts those processes in an **accounting and limiting group** that measures and caps CPU time and memory. The split is the whole point: the views *hide*, the group *caps*, and neither does the other's job. What is not fenced is the kernel itself — one version, one module set, one set of host-wide tunables for every workload on the box.

code

pseudocode · 18 lines
pseudocode
start_container(image, spec):
    child = fork_process()

    # each fence is applied independently, and any may be skipped or joined
    if spec.fence_filesystem: child.root       = assemble_root_from(image)
    if spec.fence_processes:  child.processes  = new_process_table()
    if spec.fence_network:    child.network    = new_network_view()
    if spec.fence_identity:   child.hostname   = spec.instance_name
    if spec.fence_channels:   child.ipc        = new_channel_set()
    if spec.join_processes_of is not none:
        child.processes = processes_of(spec.join_processes_of)   # deliberate hole

    # a separate mechanism, doing a different job: accounting, not hiding
    group = new_accounting_group(cpu_quota = spec.cpu_ceiling,
                                 memory_max = spec.memory_ceiling)
    group.add(child)

    child.run(spec.command)      # first process inside: number one

go deeper

for a junior

Be able to name the fences in plain words — its own filesystem, process list, network and host identity — and add that a separate mechanism caps CPU and memory. Naming four of them confidently is worth more than a vague 'it is isolated'.

for a middle

Explain why the composition is separable: a runtime can omit a fence or let a container join another's, and that flexibility is the reason a boundary can be strong in one dimension and absent in another. Keep hiding and capping in separate sentences.

for a senior

Reason about what the composition is worth in production: which fences your platform actually applies, which are commonly left off, and what a shared kernel means for blast radius when you are deciding whether this boundary is the only one a workload gets.

for a principal

Treat the fence set as a policy you own rather than a default you inherit. The trade-off is density and simplicity against the fact that one shared kernel is one failure domain, and that argument decides which workloads may share a host at all.

## A container is a composition, not a single feature The kernel has no object called a container. What a container runtime does is start an ordinary process and, in the moment before that process begins running your program, change several independent things about the world it will see, then place it in a group that accounts for what it consumes. Each change covers one kind of resource. "A container" is the name we give the bundle when they are applied together. That this is a composition, rather than one switch, is not trivia. It explains three behaviours you will meet in production: a runtime can **leave a fence out**, a second container can be started so that it **joins a fence the first one already has**, and a workload can be perfectly fenced in one dimension while being completely exposed in another. ## The fences a runtime composes - **Root filesystem.** The process sees a root directory assembled from the image's layers rather than the host's. Host paths simply do not exist for it unless someone explicitly mounts them in. - **Process table.** The container gets its own numbering. The first process it starts is number one inside, other containers' processes and the host's are invisible, and a process it cannot see is a process it cannot signal or inspect. - **Network.** Its own interfaces, its own loopback and its own port space, which is why two containers on one host can both listen on the same port number without colliding. - **Host identity.** Its own hostname, normally set per instance by the runtime rather than inherited from the machine. - **Inter-process channels.** Its own shared-memory segments and message queues, so two workloads cannot accidentally attach to each other's. - **User and group ids (often optional).** A runtime *may* also map the ids inside the boundary onto a different range outside it, so the id a process holds inside need not be the id it holds on the host. Many setups do not turn this one on, which is exactly why it is worth asking about. ## Hiding is not capping None of those views limits consumption. A container with a perfectly fenced filesystem, process table and network can still burn every processor on the host. Limiting is the job of a second, entirely separate mechanism: an **accounting and limiting group** that the runtime creates, places the container's processes into, and configures. That group measures usage and enforces ceilings, and the two ceilings behave in opposite ways: - a **CPU ceiling** is a quota of run time per period, so a process that wants more is made to wait — it slows down and stays alive; - a **memory ceiling** is a hard wall with no equivalent of waiting, so when it is reached the kernel ends a process instead. The symmetric mistake is to expect the group to hide things. It does not. It caps consumption without changing a single thing the process can read about the machine, which is the root of a whole family of start-up bugs. ## What stays shared | Fenced per container | Shared by every container on the host | |---|---| | the root filesystem it sees | the kernel itself, its version and its loaded modules | | the process table and its numbering | host-wide tunables the kernel exposes once per machine | | network interfaces, addresses and ports | the hardware, its firmware and its device state | | the host identity it reports | the time of day the host keeps | | inter-process channels | the scheduler and memory subsystem doing the actual work | The right one-line summary of the second column is that every container on a host makes its system calls into **one** kernel. A container does not have its own operating system; it ships a userland — the files, libraries and tools baked into its image — and that userland talks to the same kernel as everything else on the box. ## The composition is optional, fence by fence 1. A runtime can be told to **skip** a fence. Skip the network fence and the workload sees the host's interfaces and port space as its own. 2. A runtime can start a container that **joins** an existing container's view rather than getting a fresh one. This is how two containers are made to look like one machine to each other — a deliberate hole, chosen for a reason. 3. Skipping or joining is an **isolation decision**, not a performance tuning knob. Each one you drop is one less thing standing between that workload and everything else on the host. ## What interviewers are listening for - That you can name several fences rather than saying "it is isolated". - That you keep hiding and capping apart, because almost every follow-up in this area turns on the distinction. - That you know something is still shared, and can say what — the honest answer to "how strong is this boundary?" starts there.

  • Two containers are started so that one joins the other's process-table view. What does each gain and lose?
    They gain mutual visibility: either can list, signal and attach to the other's processes, which is what makes a companion process able to watch or manage the main one. They lose the fence between them — a compromise in either is now a compromise of both. The other fences are unaffected: each still has its own root filesystem unless a mount is shared too.
  • If the fences hide and the accounting group caps, what does neither of them do?
    Neither gives the container a kernel of its own. One kernel version, one module set and one set of host-wide tunables serve every container on the machine, so a kernel-level fault is a host-wide event rather than a per-container one. That is the ceiling on how strong this boundary can be, no matter how many fences you compose.
  • Does a container have its own operating system?
    No. It carries its own userland — the files, libraries, shells and tools its image was built with, which is why one container can carry a completely different set of system tools from its neighbour. But every system call it makes is served by the host's single shared kernel, so it is one operating system with many userlands, not many operating systems.

Less a locked room than a set of one-way mirrors fitted one at a time, with a meter on the door: the mirrors decide what the occupant can see, and the meter alone decides how much it may draw.

saying these in an interview costs you the question

  • Calls a container a lightweight virtual machine that boots its own kernel.
  • Believes one single kernel feature creates the whole boundary.
  • Thinks the fences also limit how much CPU and memory the workload may use.
  • Assumes every fence is always applied, so nothing at all is shared.
  • Says the image ships an operating system rather than a userland.
open as a page

Why does a containerized service reading the machine's CPU and memory totals see the host's numbers rather than its own ceiling?

level: middleimportance: must knowfreq 60%

basics

~20 s

The fences hide other workloads; they do not rewrite the interfaces that report machine capacity. Those numbers come from the shared kernel's host-wide view, while the container's ceiling lives in an accounting group nobody told the process about.

open as a page

A metrics agent moved into its own container can no longer see the application process it watches — which fence did that, and what fixes it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The separate process table view: each container sees only its own processes, so the agent finds nothing to watch. Either start it sharing the target's process view, collect from the host instead, or have the application publish what the agent needs.

open as a page

Why does a clustered search indexer that derives its node identity from the hostname rejoin as a brand-new member after every restart?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Each container gets its own host-identity fence, and the runtime sets that name per instance, so a replacement carries a different one. Identity meant to outlive an instance must be supplied as configuration, not read back from the hostname.

open as a page