skip to content

A host fits tens of virtual machines of one service but hundreds of containers — which per-instance costs disappeared?

level: middleimportance: must knowfreq 66%

answer

  1. marginal cost of the next copy
  2. a guest duplicates a whole system
  3. no per-container kernel or services
  4. read-only pages shared between copies
  5. the workload's footprint still stands

basics

~20 s

Everything the guest model duplicates per instance goes away: a resident kernel, a first process and system services, the memory set aside for the guest up front, and a full disk image. What remains per container is roughly the workload's own working set.

solid answer

~40 s

A guest is a whole system, so every copy pays for a whole system: a kernel resident in memory, a first process, background services, and usually a fixed allocation carved out whether the workload touches it or not — plus a disk image per instance. Containers share one kernel, so the marginal cost of another copy is close to that process's own working set, and copies of the same image share their read-only file pages rather than each holding a private copy. On disk, one set of read-only layers backs every copy instead of N disk images. What does **not** disappear is the workload itself: if one copy genuinely needs a gigabyte of live data, ten copies need ten. Density rises because the per-instance *system* overhead vanished, not because the service got smaller.

go deeper

for a junior

Know the headline and one reason for it: many more containers fit on a host because each one does not carry its own kernel and system services. Do not try to defend a specific ratio you have not measured.

for a middle

Enumerate the duplicated costs precisely — resident kernel, first process and services, memory set aside up front, a disk image each — and explain why copies of one image share their read-only pages on the host.

for a senior

Bring numbers and their assumptions, and then name the limits: the workload's own footprint, CPU, and contention for bandwidth that no ceiling partitions. Density you cannot sustain under load is not density.

for a principal

Treat it as a fleet cost model rather than a fact: what packing ratio your estate targets, what latency you are willing to pay for it, and where you deliberately keep density low to keep a failure domain small.

## What each extra copy is actually paying for The interesting number is not the size of one instance but the **marginal cost of the next one**. The two boundaries answer that very differently. Behind a machine boundary, every copy is a complete system: - a **guest kernel** resident in that guest's memory, one per instance; - a first process and a set of background system services, started and kept running; - usually a **fixed memory allocation**, decided when the guest is defined and held whether the workload touches it or not; - a **disk image** per instance, carrying firmware-to-userland everything, even when ten instances are the same build. Behind a container boundary there is one kernel for the whole host, started once, and no per-container system underneath the workload. The marginal cost of another copy is close to **the process's own working set**: the memory it actually dirties, plus its share of the pages it reads. ## Where the memory goes | per-instance cost | machine boundary | container boundary | |---|---|---| | kernel resident in memory | one per guest | one for the entire host | | first process and system services | one set per guest | none beneath the workload | | memory set aside up front | typically a fixed allocation | consumption up to a ceiling | | read-only program pages | private to each guest | shared between copies of one image | | on-disk artifact | one disk image per instance | one set of read-only layers, shared | The shared-read-only-pages row is the one people forget. Ten containers started from the same image are reading the same files from the same place on the host, so the host's page cache holds one copy of those pages and every container reads it. Each container's *private* memory — what it writes, allocates and dirties — is genuinely its own. ## A worked sense of scale State the assumptions, because this is where hand-waving starts. Take a host with 128 GB of usable memory and a service whose live working set is around 300 MB. - As guests: give each 4 GB because a whole system has to live in there comfortably, and the host fits roughly **30**. The workload is using about 300 MB of that 4 GB; the rest is guest system plus headroom you paid for in advance. - As containers: each costs roughly its 300 MB of private memory plus a shared read-only slice, so the same host holds in the low **hundreds** before memory is the binding constraint. The ratio moved by about an order of magnitude, and every bit of that came from not duplicating a system per instance. ## What density does not remove Three things survive the boundary change, and a candidate who claims otherwise is overselling: 1. **The workload's own footprint.** If a copy needs a gigabyte of live data, ten copies need ten gigabytes on either side of the comparison. 2. **CPU.** Sharing a kernel does not create cores. Hundreds of mostly-idle workers fit; hundreds of busy ones do not. 3. **Everything the boundary never partitioned.** Disk bandwidth, network bandwidth and the kernel's own work are consumed jointly, and packing more copies onto a host makes contention for them more likely, not less. So the honest claim is narrow: **the per-instance system overhead disappeared**, and that is usually the dominant term for small services — which is exactly the shape most services have. ## Why the artifact shrinks too The same fact explains image size. An image needs no kernel, no firmware and no boot loader, and it does not need a full userland unless the workload does. A minimal image can be a few megabytes; a disk image is normally gigabytes. And because images are built in read-only layers, copies that share a base share those bytes on the host and in transit, so pulling the eleventh copy of a build usually transfers only what is unique to it. ## Saying it well in an interview Lead with the mechanism — "a guest duplicates a whole system per instance and a container does not" — then name the four duplicated costs, then volunteer the limit: density is bounded by the workloads themselves, and packing tightly moves your problem from memory footprint to contention for the things no ceiling partitions. That last sentence is what separates a recited benefit from an engineer who has actually packed a host.

  • If ten containers run from the same image, why does the host not hold ten copies of the program in memory?
    They read the same read-only files from the same layers on the host, so the host's page cache holds one copy of those pages and maps it into every container. Only what a container writes or allocates is private to it. That shared-read-only effect is a large part of why density scales the way it does.
  • Where does the density argument stop being true?
    When the workload, not the boundary, is the binding constraint: large live data, real CPU demand, or heavy disk and network use. Packing many copies onto one host also concentrates contention for resources that no ceiling divides, so the win in footprint can be paid back in latency.

saying these in an interview costs you the question

  • Says containers use less memory than the same code in a guest would.
  • Claims density is unlimited because containers have no per-instance cost.
  • Attributes density to image compression rather than to unduplicated system overhead.
  • Thinks the whole image is loaded into memory when a container starts.
  • Ignores that CPU and disk bandwidth still bound how much you can pack.