skip to content

You are designing a platform that will execute code submitted by untrusted customers. How would you decide between plain containers, micro-VMs, and full virtual machines as the isolation boundary, and what would you put around whichever you pick?

level: principalimportance: should knowfreq 40%

answer

  1. Threat model first: who writes the code
  2. Untrusted → boundary below the shared kernel
  3. Micro-VM = container economics, VM-ish worst case
  4. Per-workload runtime class, not cluster-wide
  5. Node separation, no node credentials, deny egress, hard caps

basics

~20 s

Decide from the threat model, not preference: untrusted code needs a boundary below the shared kernel, so micro-VMs are the default; plain containers only for trusted code; full VMs when you need coarse, long-lived, compliance-visible separation. Then add tenant-separated nodes, network egress control and hard resource caps.

solid answer

~60 s

Start from **who writes the code and what a breach costs**. Untrusted, customer-supplied code plus a shared kernel means one kernel bug compromises every tenant on that node — usually unacceptable. My default is **micro-VMs per execution**: each job gets its own kernel and a tiny device model, boot in ~100–200 ms, and the container ecosystem stays intact via a sandboxed runtime class. **Plain containers** are for first-party trusted services, or for a second layer *inside* a micro-VM. **Full VMs** are for long-lived, coarse-grained tenants, for hardware access, or where an auditor needs a boundary they already accept. Axes I weigh: blast radius per tenant, start-up latency budget and whether the workload is per-request or long-lived, density and therefore unit cost, syscall compatibility needs, whether the instance type supports nested virtualization, and the operational cost of a second kernel to patch. Whatever I choose I add: per-tenant node pools, no shared credentials on the node, default-deny egress, hard CPU/memory/PID/disk caps with timeouts, no host mounts or runtime sockets, non-root plus dropped capabilities inside, immutable short-lived instances, and a rapid host-kernel patch path.

go deeper

for a junior

Recognise that untrusted code should not rely on a shared kernel as its only boundary, and name stronger options.

for a middle

Compare the three boundaries on isolation strength, start-up and overhead, and list concrete hardening around whichever is chosen.

for a senior

Make a defensible recommendation with per-workload boundaries, node-level tenant separation, credential and egress control, and a patching plan for both kernels.

for a principal

Own the framing: an explicit blast-radius budget, cost per tenant, latency budget, per-workload boundary policy, validation by adversarial testing, and the organisational commitments — patch SLAs, incident containment — that make the design real.

## Start with the threat model, not the technology The decision follows from three questions: 1. **Who authors the code?** First-party engineers, semi-trusted partners, or anonymous internet users. 2. **What does one compromise cost?** Cross-tenant data exposure, credential theft on the node, lateral movement into the control plane, regulatory consequences. 3. **What is the latency and density budget?** A per-request sandbox has a millisecond budget; a nightly batch job has minutes. Unit economics depend on how many tenants fit on a node. For untrusted customer code the answer to (1) and (2) already rules out the shared kernel as the only boundary: the Linux syscall surface is far too large to bet cross-tenant confidentiality on a single bug not existing. ## What each boundary is genuinely for **Plain containers (runc).** The boundary is the host kernel. Correct choice for first-party services, and as an inner layer inside a stronger boundary. Cheapest, densest, fastest, most compatible. Not, on its own, a tenancy boundary for hostile code. **Micro-VMs / sandboxed runtimes.** Each workload gets its own kernel behind a minimal device model, or its syscalls are served by a user-space kernel. Start-up in the low hundreds of milliseconds, memory overhead in the tens of megabytes, and an attacker needs a two-bug chain to reach the host. This is the modern default for multi-tenant code execution and it is why serverless platforms are built on it. Costs: a second kernel to patch, hardware-virtualization requirements (nested virt is not universally available), reduced host integration, and per-instance overhead that shows up in density and margin. **Full VMs.** Coarse and heavy, but they buy things the middle ground does not: long-lived tenant environments, device and GPU passthrough, an isolation story auditors and customers already recognise, and independent lifecycle and kernel choice per tenant. Right when tenants are few, long-lived and large, or when contracts demand dedicated infrastructure. ## The decision, stated as rules I would defend - **Untrusted code, short-lived, high volume** → micro-VM per execution, containers inside it, aggressive reuse only within a single tenant. - **Untrusted code, long-lived, few tenants** → dedicated VMs or dedicated node pools per tenant. - **Trusted first-party services** → hardened containers on shared nodes; spending micro-VM overhead here buys little. - **Mixed sensitivity** → make the boundary a per-workload attribute (a runtime class), not a cluster-wide decision, so you pay only where risk lives. - **Nested virtualization unavailable** → either move to instance types that support it, or accept a user-space-kernel sandbox with its compatibility and I/O costs, and narrow the workload accordingly. ## The controls that matter regardless of choice The boundary is necessary, not sufficient. Around it: - **Tenant separation at the node level.** Never co-schedule hostile tenants on one host if the data would matter; sandbox escapes are rarer than misconfiguration, but node-level separation makes escape non-catastrophic. - **No credentials on the node worth stealing.** No broad instance role, no shared registry secret, workload identity scoped per job with short TTLs. Assume the sandbox can be broken and ask what the attacker gains. - **Default-deny egress**, DNS through a controlled resolver, no access to the metadata endpoint, no route to the control plane. - **Hard resource limits and wall-clock timeouts**: CPU, memory, PIDs, disk quota, open files, network bandwidth. Denial of service and cryptomining are far more common than escapes. - **Immutable, single-use instances.** One job per sandbox, destroyed afterwards; no state carried between tenants. - **Inner hardening still on**: non-root, all capabilities dropped, read-only root filesystem, seccomp, no host mounts, no runtime socket. - **Fast patching of both kernels** — host and guest — with node rotation as the delivery mechanism, plus an emergency path measured in hours, not weeks. - **Detection**: syscall and process auditing, egress anomaly alerts, per-tenant quota breach alerts. You need to know when someone is probing. ## How I would present the trade-off to stakeholders I would express it as blast radius per dollar: plain containers give maximum density and the worst worst-case; VMs the reverse; micro-VMs sit close to container economics with most of the VM's worst-case reduction. I would set an explicit blast-radius budget — for example, no single compromise may expose more than one customer's data — and let that budget, not aesthetics, pick the boundary. Then I would validate empirically: measure start-up latency and per-sandbox memory under real load, run escape and abuse exercises against the platform, and confirm that the assumed containment actually holds before onboarding customers.

  • Micro-VMs raise per-execution cost. How do you justify that to a business worried about margin?
    Price the alternative: with a shared kernel, a single kernel bug is a cross-tenant data breach, and its expected cost — remediation, notification, contractual and regulatory exposure, churn — dwarfs the compute delta. I would quantify the overhead precisely (start-up latency, memory per sandbox, achievable density) and set a blast-radius budget the business agrees to, so the spend is a stated risk decision rather than an engineering preference.
  • Where do plain containers still belong in such a platform?
    Two places. First, all first-party control-plane and service workloads, which are trusted and benefit from density and speed. Second, as the inner layer inside each micro-VM: the sandbox provides tenant isolation while the container still runs non-root with dropped capabilities and seccomp, so a bug in the guest userland does not immediately own the guest kernel.
  • How would you validate that the isolation you designed actually holds?
    Test it rather than assume it: run known escape techniques and abuse patterns against a staging copy, verify that a compromised sandbox reaches no credentials, no metadata endpoint and no other tenant, and confirm resource caps and timeouts stop fork bombs, disk fillers and long-running miners. Add continuous checks so a configuration drift — a re-enabled host mount, an unfiltered egress rule — is detected rather than discovered by a customer.

saying these in an interview costs you the question

  • Choosing the boundary by fashion rather than from a stated threat model
  • Treating hardened containers as sufficient tenancy isolation for hostile code
  • Applying the heaviest boundary uniformly instead of per workload class
  • Securing the sandbox while leaving node credentials, metadata access or open egress reachable
  • Ignoring abuse and denial-of-service — the common failure — in favour of escape scenarios only
  • Forgetting the guest kernel becomes a second patching obligation

context