What has to stay the same when a host swaps its low-level runtime for a sandboxed one, and what changes?
answer
- the slot has a written contract
- same bundle, same lifecycle operations
- the image never learns which runtime ran it
- what changes is who answers the system calls
- compatibility gaps and start-up cost are the price
basics
~20 sEverything above the seam stays: the same image, the same configuration document, the same manager and the same lifecycle operations. What changes is who answers the workload's system calls — and with it start-up time, throughput and which unusual calls and devices still work.
solid answer
~50 sThe slot a low-level runtime occupies is defined by a written interface: it is handed an unpacked root filesystem plus a configuration document, and must support create, start, state, signal and delete on it. Anything honouring that fits, so a host can be pointed at a **sandboxed runtime** without rebuilding an image, editing a workload spec or changing the manager — the artifact never learns which runtime ran it. What changes is the execution path. A sandboxed runtime puts something between the workload and the host kernel, so its system calls are answered by an intercepting layer rather than going straight to the shared kernel. The price is paid in start-up latency, per-workload memory, I/O throughput, and compatibility: an unusual system call, a direct device or an exotic filesystem behaviour may simply be unimplemented, so a workload can pass on one runtime and fail on the other.
go deeper
The point to recall is that the runtime is a replaceable part. Because it is handed a standard bundle and asked for standard operations, changing it does not require touching the image at all.
Explain the seam concretely — bundle in, lifecycle operations out — and name what stays fixed across a swap: image, configuration document, manager, output and exit-status handling, and the layer above the node.
Show the cost side with the failure you would plan for: start-up latency, per-workload memory and throughput all move, and a compatibility gap around unusual system calls or devices shows up late, so validation has to use the real workload under real traffic.
Treat runtime selection as a per-workload-class policy rather than a host-wide bet: keep the default for the bulk of work, route untrusted or multi-tenant work to the stronger boundary, and require evidence from real traffic before either decision becomes a standard other teams inherit.
This question is really about what an interface buys you. Because the seam between the container manager and the low-level runtime is written down, the runtime is a replaceable part — and a fleet can therefore change what is doing the isolating for some workloads without changing anything the developers of those workloads touch. ## What the slot actually requires A low-level runtime is handed two things and asked for a small set of operations: - **The bundle** — an unpacked root filesystem directory plus a configuration document naming the process to run, the mounts, the per-resource views, the ceilings, the user and the privilege set. - **The lifecycle operations** — create the container, start it, report its state, deliver a signal to it, delete it. Anything that can do those things is a candidate for the slot. That is the whole reason a swap is even possible. ## What stays the same 1. **The image.** Not rebuilt, not re-tagged, not re-pushed. Its digest is unchanged and the choice of runtime is recorded nowhere inside it. 2. **The configuration document.** The same declared command, mounts and ceilings are handed over. 3. **The container manager**, and everything it does: pulling, verifying, unpacking, assembling the document. 4. **The output-stream and exit-status contract.** A supervising process still holds the pipes and the status the same way. 5. **Everything above the node.** How workloads are declared, placed and scaled is untouched, because the swap happens entirely below that layer. ## What changes What changes is the execution path of your process's system calls. Instead of the workload calling straight into the shared host kernel, a sandboxed runtime interposes: one family implements the system-call surface in a separate user-space layer that handles the calls itself, another starts the workload behind a very small machine-style monitor. Either way the host kernel is no longer the first thing a hostile workload's syscalls reach, which shrinks the surface a kernel flaw is exploitable through. The cost comes in four shapes, and all four are measurable: | Dimension | What to expect after the swap | |---|---| | Start-up latency | Higher — there is more to construct before the first instruction runs | | Memory per workload | Higher — the interposing layer has its own footprint per container | | I/O and network throughput | Lower — calls travel through an extra layer, and the hit is worst for syscall-heavy and I/O-heavy work | | Compatibility | Narrower — unusual system calls, direct device access and some filesystem semantics may be unimplemented | ## Why hosts usually run more than one In practice a node keeps several runtimes installed and maps a named class of workload to one of them, because the choice is a property of **how a workload is run**, not of the artifact. Untrusted or multi-tenant work gets the sandboxed runtime; everything else keeps the default one and its performance. The same image can legitimately run under both on the same host on the same day, which is the clearest demonstration that the artifact and the fencing decision are genuinely separate concerns. ## The failure you should predict The compatibility gap is where a swap goes wrong, and it goes wrong *late* — a trivial workload passes every smoke test and the real one fails under load. The workloads at risk are the ones that reach past the ordinary system-call surface: direct device access, unusual kernel interfaces, very high-throughput storage or network work, or anything that quietly depends on a host behaviour the interposing layer has not reproduced. The correct validation is to run the actual workload, with its real traffic shape, under the candidate runtime — not a minimal test container, which proves only that the slot was filled. A second, subtler point: swapping the runtime does not change what is *in* the image, and it does not grant a workload anything it did not have. A container that needed an elevated privilege still needs it, and asking a sandbox for that privilege may simply be refused. The swap changes who enforces the boundary and how strong it is; it does not change what the workload asked for.
- How does one host end up running two different low-level runtimes at the same time?The manager maps a named runtime class to a binary and a default set of options, and the request that starts a workload names a class. So untrusted or multi-tenant work is started under the sandboxed runtime while everything else keeps the default, on the same host, from byte-identical images. The selection lives with how the workload is run, never inside the artifact.
- What kind of workload most often fails after such a swap?One that reaches past the ordinary system-call surface: direct device access, unusual kernel interfaces, very high-throughput storage or network work, or a dependency on a host behaviour the interposing layer has not reproduced. Validate with the real workload under real traffic — a minimal test container proves only that the runtime slot was filled correctly.
saying these in an interview costs you the question
- Thinks the image must be rebuilt to run under a different runtime
- Assumes a sandboxed runtime is a drop-in with no compatibility gaps
- Says the choice of runtime is recorded inside the image
- Believes swapping the runtime also changes the layer above the node
- Expects identical start-up latency and I/O throughput after the swap