skip to content

questions

5

A colleague says "a Docker container is just a lightweight virtual machine." Explain what a container actually virtualizes compared with a hardware virtual machine running under a hypervisor, and why the distinction matters in practice.

level: juniorimportance: must knowfreq 88%

answer

  1. VM = own kernel; container = shared kernel
  2. Image ships userland, no kernel
  3. docker run = clone+exec, not boot
  4. Hypervisor boundary vs kernel boundary
  5. Docker Desktop = containers inside a Linux VM

basics

~20 s

A virtual machine virtualizes hardware: the hypervisor gives each guest its own kernel and virtual devices. A container virtualizes the operating system: processes share the host kernel and are isolated by kernel features. A container is a process, not a machine.

solid answer

~50 s

A **VM** is produced by a hypervisor that emulates or paravirtualizes hardware — CPU, memory, disk, NIC. Each guest boots its **own kernel**, its own init, its own drivers. Isolation is enforced below the OS by the hypervisor plus CPU virtualization extensions. A **container** is just one or more ordinary host processes. There is no guest kernel and no boot: the runtime starts the process with a restricted view of the system — its own PID/mount/network/user views, capped CPU and memory, and a root filesystem pointing at the image's unpacked layers. Every container on the box makes syscalls into the *same* host kernel. Why it matters: containers start in milliseconds and cost only what their processes cost, so density is much higher; but an image ships **userland only**, so you cannot run a different kernel or a different OS family, and a kernel bug is a host-wide problem rather than a per-guest one.

code

bash · 8 lines
bash
uname -r
# 6.8.0-45-generic

docker run --rm ubuntu:22.04 uname -r
# 6.8.0-45-generic

docker run -d --name web nginx:alpine
ps -e -o pid,comm | grep nginx   # visible as ordinary host processes

go deeper

for a junior

State the one-line difference cleanly — own kernel via hypervisor versus shared host kernel — and give two consequences: fast startup and small images.

for a middle

Add mechanism: images carry userland only, the runtime clones a process with restricted views and limits, and kernel features come from the host, so kernel-sensitive apps can break across hosts.

for a senior

Frame it as choosing an isolation boundary and note the operational fallout: density and start-up wins versus shared-kernel attack surface, non-namespaced sysctls, and shared page cache.

for a principal

Talk about when each boundary is the right platform primitive — tenancy model, compliance, kernel-version fleet management — and when you deliberately pay for a VM or micro-VM under the container.

## Two different layers to cut at "Virtualization" hides two very different mechanisms, and this question is really asking *which layer* you slice. **Hardware virtualization (the VM).** A hypervisor — KVM, ESXi, Hyper-V, Xen — presents each guest with what looks like a physical computer: virtual CPUs backed by hardware extensions (Intel VT-x / AMD-V), a block device, a NIC, firmware. The guest then boots a full operating system on that fake hardware: bootloader, kernel, device drivers, init system, system daemons, then finally your application. Nothing above the hypervisor is shared between guests. Guest A can run Linux 5.15, guest B Windows Server, guest C FreeBSD, all on one host. **Operating-system virtualization (the container).** There is no fake hardware and no second kernel. A container is a normal Linux process — you can see it in `ps` on the host — that the kernel has been asked to show a *partial view* of the machine. The kernel gives it an isolated process table, mount table, network stack, hostname, IPC space and UID mapping, plus resource limits, and points its root directory at the unpacked image. `docker run` does not boot anything; it `clone()`s and `exec()`s. ## What is actually inside an image A container image contains **userland only**: libc, binaries, config, your app. It contains no kernel. That is why an image of a few megabytes can "be Alpine" while a VM disk image of the same distro is hundreds of megabytes — the VM must carry a kernel and everything needed to boot it. Prove it to yourself: run `uname -r` inside an Ubuntu container on a CentOS host and you get the *host's* kernel version, not Ubuntu's. The consequences follow directly: - **No mixed kernels.** You cannot run a Windows container on a Linux kernel, or a Linux container on a Windows or macOS kernel. Docker Desktop hides this by running a small Linux VM (WSL 2 on Windows, a hypervisor VM on macOS) and putting your containers inside it — which is why "Docker on my Mac" is really containers-in-a-VM. - **Kernel features leak through.** If your workload needs a kernel module, a specific cgroup v2 feature, io_uring, or a newer seccomp behaviour, the *host's* kernel decides, not the image. Kernel-version-sensitive apps are a classic "works on my laptop" trap. - **The kernel is shared attack surface.** Every container's syscalls hit the same kernel; a kernel privilege-escalation bug is potentially a whole-host compromise, whereas in a VM the attacker must also break the hypervisor. ## Where each cost lives | | VM | Container | |---|---|---| | Isolation boundary | Hypervisor + CPU virt extensions | Host kernel | | Boots a kernel | Yes | No | | Start time | Seconds to tens of seconds | Milliseconds | | Fixed overhead per instance | Guest kernel + OS services, typically 100s of MB | Only the process's own memory | | Density on one host | Tens | Hundreds to thousands | | Can run a different OS/kernel | Yes | No | | Image size | Whole disk with kernel | Userland layers only | ## Why the sloppy phrasing is dangerous Calling a container a lightweight VM leads to real mistakes. People treat images as machines and stuff sshd, cron, syslog and an init system in, instead of running one process that logs to stdout. They assume a `--memory` limit is as hard as a VM's RAM allocation and are surprised when the container is OOM-killed by the *host* kernel and the host's page cache is shared. They assume tenant isolation is equivalent and put mutually hostile workloads side by side. And they debug "the container's kernel" that does not exist — kernel tuning, `sysctl` values that are not namespaced, and kernel logs all belong to the host. ## The honest summary for an interview A VM virtualizes *hardware* so a guest OS can run unchanged; a container virtualizes the *OS interface* so a process believes it has the machine to itself. The trade is fidelity and isolation strength (VM) against startup speed, density and image size (container). Everything else — why images are small, why start-up is instant, why kernel exploits are scarier, why Docker Desktop needs a VM — falls out of that one difference.

  • If containers share the host kernel, how can I run an Ubuntu image on a CentOS host?
    Because the image supplies only userland — libc, binaries, package layout — and those talk to the kernel through the stable Linux syscall ABI. Any kernel new enough for the distro's libc will run it. It breaks only when the app genuinely needs a kernel feature, module or version the host does not provide.
  • Why does Docker Desktop on macOS or Windows use a virtual machine?
    Linux containers need Linux kernel features (namespaces, cgroups, overlayfs) and macOS and Windows kernels do not provide them. Docker Desktop therefore runs a lightweight Linux VM — a hypervisor VM on macOS, WSL 2 on Windows — and your containers are processes inside that VM, which is why file sharing and networking across the boundary are slower.

VMs are separate houses, each with its own foundation, plumbing and boiler. Containers are apartments in one building: lockable and metered, but sharing a single set of pipes — the kernel. Renovating a flat is fast; a burst main affects everyone.

saying these in an interview costs you the question

  • Saying the image contains a kernel or that each container boots its own OS
  • Claiming containers give the same isolation strength as VMs
  • Believing a Linux container can run directly on the macOS or Windows kernel
  • Describing Docker as a hypervisor
  • Assuming kernel tuning or sysctl inside a container affects only that container

context

open as a page

Why does a typical container start in tens of milliseconds and consume roughly the memory of its own processes, while a comparable virtual machine takes tens of seconds and reserves hundreds of megabytes? Walk through where the time and the memory actually go.

level: middleimportance: must knowfreq 66%

basics

~20 s

A VM must emulate firmware, boot a kernel, probe devices and run an init system, and its guest kernel plus OS services permanently occupy RAM. A container skips all of that: the runtime sets up isolation and execs your process against the already-running host kernel.

open as a page

If a process inside a Linux container exploits a kernel vulnerability, how does the blast radius compare with the same exploit fired inside a hardware virtual machine — and which container-level controls actually shrink it?

level: seniorimportance: must knowfreq 56%

basics

~20 s

In a container the kernel is the isolation boundary, so a kernel privilege-escalation bug means host compromise and every co-tenant container with it. In a VM the attacker owns only that guest and must also break the hypervisor. Shrink it by cutting syscall and capability reach.

open as a page

Explain what micro-VM and sandboxed runtimes such as Firecracker, Kata Containers and gVisor give you that a standard container started by runc does not, and what each one costs.

level: seniorimportance: should knowfreq 38%

basics

~20 s

They add a second isolation boundary under the container. Firecracker and Kata run each workload in a stripped-down VM with its own kernel and a tiny device model; gVisor keeps one host kernel but services most syscalls in a user-space kernel. Cost: some start-up, memory and feature loss.

open as a page

You are designing a platform that will execute code submitted by untrusted customers. How would you decide between plain containers, micro-VMs, and full virtual machines as the isolation boundary, and what would you put around whichever you pick?

level: principalimportance: should knowfreq 40%

basics

~20 s

Decide from the threat model, not preference: untrusted code needs a boundary below the shared kernel, so micro-VMs are the default; plain containers only for trusted code; full VMs when you need coarse, long-lived, compliance-visible separation. Then add tenant-separated nodes, network egress control and hard resource caps.

open as a page