skip to content

Why does a typical container start in tens of milliseconds and consume roughly the memory of its own processes, while a comparable virtual machine takes tens of seconds and reserves hundreds of megabytes? Walk through where the time and the memory actually go.

level: middleimportance: must knowfreq 66%

answer

  1. VM: firmware → kernel → udev → init → app
  2. Container: overlay mount → isolate → exec
  3. Guest kernel + double page cache = the RAM floor
  4. Shared layers = shared pages across containers
  5. Cold pull and JVM warm-up beat the runtime's ms

basics

~20 s

A VM must emulate firmware, boot a kernel, probe devices and run an init system, and its guest kernel plus OS services permanently occupy RAM. A container skips all of that: the runtime sets up isolation and execs your process against the already-running host kernel.

solid answer

~50 s

**VM startup** is a real machine boot: firmware, bootloader, kernel decompress and init, device probing, udev, init system, then dozens of system services before your app runs. That is seconds. **Memory** is committed to the guest kernel, page tables, its page cache and its daemons — commonly 150–500 MB before the workload exists — and it is largely non-shareable across guests. **Container startup** is process creation. The runtime mounts an overlay filesystem from cached image layers, creates the isolated views, applies limits and `exec`s the entrypoint. No kernel boot, no device probe, no init. Milliseconds, once layers are local. **Memory** is only your process's resident pages; the kernel, its page cache and shared libraries are shared with the host, and identical layers are shared across containers. Consequences: containers give per-request or per-job scaling, cheap scale-to-zero and high density; VMs suit long-lived, coarse-grained instances. Caveats: pulling a cold image dominates start time, and a JVM or Python app's own warm-up may dwarf the runtime's milliseconds.

code

bash · 7 lines
bash
docker pull alpine:3.20
time docker run --rm alpine:3.20 true

# separate the pull from the start
time docker create --name t alpine:3.20 true
time docker start -a t
docker rm t

go deeper

for a junior

Say the VM boots a whole operating system while the container just starts a process on the running host kernel, and give the rough numbers.

for a middle

Enumerate the boot stages and the memory floor — guest kernel, second page cache, OS daemons — and explain layer sharing across containers.

for a senior

Add the caveats you have hit in production: cold pull cost, application warm-up dominating, limits being enforcement not reservation, and density-driven kernel contention.

for a principal

Discuss how start-up cost shapes architecture and cost — scale-to-zero, per-job sandboxes, node warm pools, image distribution strategy — and where you would still choose VM-shaped instances.

## Where a VM's seconds go Booting a VM repeats everything a physical computer does: 1. **Firmware** (BIOS/UEFI) initializes and hands off to a bootloader — hundreds of milliseconds, sometimes more. 2. **Kernel load and decompress**, then early init: memory maps, scheduler, page tables. 3. **Device probing** — the guest discovers virtual disks, NICs, consoles; udev enumerates and loads modules. 4. **Root filesystem mount**, possibly initramfs pivot. 5. **Init system** (systemd) brings up dozens of units: logging, networking, DNS, ssh, cron, monitoring agents. 6. Only then does your application start, plus its own warm-up. Stripped-down images can shorten this, but the shape is fixed: you are booting an operating system. ## Where a container's milliseconds go `docker run` performs no boot. Roughly: 1. **Assemble the root filesystem** — union-mount the already-unpacked image layers (overlayfs) into one read-only stack plus a thin writable layer. Cheap because layers are on disk and shared. 2. **Create the isolated views and limits** — the runtime asks the kernel for a fresh process/mount/network/user view and attaches resource limits. 3. **Pivot root, drop capabilities, apply the seccomp profile.** 4. **`exec` the entrypoint.** That is a handful of syscalls against a kernel that is already running, already has its drivers loaded and already has the machine's page cache warm. Tens of milliseconds is normal; container *creation* can be under 100 ms even for large images already present locally. ## Where the memory goes A VM's footprint has an irreducible floor: - **Guest kernel text and data**, its own page tables and slab caches. - **A second page cache** — the guest caches file data that the host may also cache, so the same bytes can be resident twice. - **Guest OS services**: journald, networking, ssh, agents. Call it 150–500 MB per instance before the workload. Ballooning and page-sharing recover some of it, unevenly. A container's footprint is its processes' resident set, and much of that is *shared*: - One kernel and one page cache for the whole host. - Identical image layers are the same files on disk, so ten containers from one image share the executable and library pages in memory. - No per-instance init or daemons if you follow one-process-per-container. That is why a host that seats 20 VMs may seat several hundred containers. ## The honest caveats - **Cold image pull dominates.** "Milliseconds" assumes the layers are local. A 900 MB image pulled over a slow link turns start-up into minutes. This is why teams obsess over image size, layer reuse, registry locality and pull-through caches, and why lazy-loading snapshotters exist. - **The app's own warm-up is often the real number.** A JVM service with Spring context initialization, JIT warm-up and connection-pool creation may need 5–20 seconds. The runtime's 30 ms is noise. Do not promise instant scale-out on top of a slow application. - **Density is not free.** Hundreds of containers means hundreds of processes contending for one kernel's locks, one network stack, one page cache and one set of file descriptors. Noisy-neighbour effects are real even with limits, because not everything is accounted per-container. - **A memory limit is not a reservation.** A container limit is enforced by the kernel's resource controller against actual usage; exceeding it gets the process OOM-killed rather than the guest swapping. VM RAM is generally assigned up front. - **Cheap start-up changes architecture.** Millisecond starts make per-job containers, ephemeral CI runners, scale-to-zero services and per-request sandboxes viable — patterns you would not attempt if every instance cost 30 seconds and 400 MB. ## The one-sentence answer A VM pays to *boot and maintain an operating system* per instance; a container pays only to *start a process* on an operating system that is already up, and shares the kernel, its caches and its image pages with every neighbour.

  • If containers start so fast, why does our autoscaler still take a minute to add capacity?
    Because runtime start-up is usually the smallest term. Node provisioning or scheduling, pulling a cold image, and the application's own initialization — framework wiring, JIT warm-up, cache and pool priming, readiness probes — dominate. Fix it by shrinking and pre-pulling images, warming nodes, and attacking application start-up, not the container runtime.
  • Ten containers run from the same image on one host. How much memory does the image cost?
    Roughly once, not ten times, for the read-only parts. The layers are the same files, so executable and library pages are shared in the host page cache and mapped by all ten. Each container adds only its private anonymous memory — heap, stacks — plus whatever it writes into its own writable layer.

saying these in an interview costs you the question

  • Claiming containers use no memory or have zero overhead
  • Believing each container has its own page cache or its own kernel memory
  • Quoting millisecond starts while ignoring cold image pulls and app warm-up
  • Treating a container memory limit as a reserved allocation like VM RAM
  • Assuming high density carries no noisy-neighbour or shared-kernel contention cost

context