A machine boundary per run costs seconds and density — which customer-template renders still justify it?
answer
- whose code, and what else is here
- trust between your own workloads
- one kernel between them and everything
- long runs make strong boundaries cheap
- write it down as a rule
basics
~20 sRuns whose code the customer controls, on hosts that also hold other tenants' data or platform credentials, justify a machine boundary. The deciding questions are whose code it is and what else is reachable through one shared kernel — not how risky the feature feels.
solid answer
~50 sDraw the line on two facts, not on instinct. First, **whose code is it**: a template language your team wrote, restricted and reviewed, is your own workload, and a shared-kernel container between your own workloads is a perfectly good boundary. Arbitrary customer-supplied code is somebody else's program running on your kernel. Second, **what else is on that host**: other tenants' data, platform credentials and anything with reach into your control path all sit behind that same one kernel. Where both point the wrong way, the shared kernel is a thin thing to bet a tenant's data on, and a machine boundary — or a hardware-virtualized sandbox as the middle option — is worth its cost. Price that cost honestly: seconds of start-up unless you keep guests warm, an order of magnitude less density, slower I/O, and a second kernel to patch.
code
pseudocode · 14 lineschoose_boundary(run):
if run.code was authored and reviewed by us:
return SHARED_KERNEL_CONTAINER
# from here the code is the customer's
if run needs a kernel feature the fleet does not offer:
return MACHINE_BOUNDARY_PER_TENANT
if host also holds other tenants' data or platform credentials:
if run is short and starts are frequent:
return HARDWARE_VIRTUALIZED_SANDBOX
return MACHINE_BOUNDARY_PER_TENANT
return SHARED_KERNEL_CONTAINER on a pool that holds nothing elsego deeper
Take away one rule: a shared kernel is a fine boundary between workloads that already trust each other, and a thin one between you and somebody else's code.
Be able to name what the machine boundary actually costs — start-up, density, I/O, a second kernel to patch — instead of treating stronger isolation as free.
Argue a specific case end to end: what runs on the host, how long a run lasts, what a warm pool buys, and what you still do inside the boundary regardless of which one you picked.
Turn it into a platform standard with a stated cost and stated evidence for revisiting it, so teams stop re-litigating per feature and the estate has one defensible line.
## The question is whose code, not how scary it feels A renderer that executes customer-supplied templates is the case where this decision stops being theoretical, because the code being run is not yours. Two facts decide it, and neither is a feeling: 1. **Whose code runs?** A constrained template language your team designed, implemented and reviewed is *your* workload, however untrusted the data flowing through it. Arbitrary customer-supplied code — an escape hatch that lets a template call out, an uploaded extension, a formula language rich enough to be a program — is *their* program, executing against your kernel. 2. **What else is reachable from that host?** Other tenants' rendered output, cached source data, the credentials the renderer itself holds, and any path back into the platform's own control surface. Everything on that host sits behind the one shared kernel. When the answer to (1) is "theirs" and the answer to (2) is "quite a lot", the shared kernel is the only thing between someone else's program and your tenants. That is thin, because a single flaw reachable through the kernel's call surface is enough, and you do not get to know in advance whether one exists. ## What to ask about each run - **Is the code attacker-chosen or merely attacker-influenced?** Influenced data in your own interpreter is a very different risk from arbitrary execution. - **What is co-located?** A host that renders and nothing else is a smaller prize than one that also holds secrets or other tenants' data. - **How long does a run last, and how often?** A boundary that costs a second of start-up is invisible on a 30-second render and ruinous on a 50-millisecond one. Long runs make strong boundaries cheap. - **What does the run need from the kernel?** Unusual kernel requirements push you to a machine boundary for the reason in their own right. - **What does the contract or regime demand?** Some obligations say plainly that one customer's code may not share a kernel with another's data. That answer is given to you. - **What is the blast radius if you are wrong?** One tenant's renders, or every tenant's data, or the platform. ## What the machine boundary costs | cost | effect | mitigation | |---|---|---| | start-up | seconds per run instead of milliseconds | keep a warm pool of guests, pay for idle | | density | tens per host rather than hundreds | more hosts, higher cost per render | | I/O | a virtualized device path, slower for heavy reads and writes | keep the working set small, stream results out | | operations | a second kernel per guest to patch and size | standardise one guest build for everything | | latency variance | pool exhaustion turns into queueing | size the pool against peak, shed load explicitly | And note what it does *not* buy: a machine boundary does not excuse you from the ordinary controls inside it. The render should still run unprivileged, with a read-only root and the narrowest privilege set that works, so that the strong boundary is the last line rather than the only one. ## The middle option Between the two sits a **hardware-virtualized sandbox**: each run gets its own stripped-down guest kernel, but the guest carries only what the workload needs, so start-up is a fraction of a conventional boot. That buys most of the boundary at a fraction of the latency, and it costs density, I/O throughput and some compatibility, because a minimal guest kernel does not necessarily do everything a workload expects. For short, frequent, untrusted runs it is very often the honest answer, and a platform decision worth making once rather than per team. ## Making it a standard rather than a case-by-case argument The useful output of this reasoning is a rule the platform applies, not a debate each team re-runs: - our own reviewed code, co-located with our own workloads: shared-kernel container; - customer code, short and frequent: sandboxed guest kernel on a pool that holds nothing else; - customer code that is long-running, holds sensitive data, or carries its own kernel requirements: a machine boundary per tenant; - in every case: unprivileged, minimal privilege set, and no platform credentials on the render hosts. Write down what evidence would move each line — a measured start-up budget, a compliance obligation, a change in what the template language can reach — so that the rule is revisited on facts rather than after an incident.
- The team says ceilings on CPU and memory make the shared kernel safe enough. What is wrong with that?Ceilings bound consumption, not reach. They stop one run starving the host, which is a real and separate problem, but they do nothing about a flaw reachable through the shared kernel's call surface. Capacity control and isolation strength are different axes, and the second one is what the boundary choice is about.
- How would you make a machine boundary per run affordable for short renders?Keep a pool of pre-started guests and hand each run a fresh one, discarding it afterwards, so the boot happens off the request path. You pay for idle capacity and must size the pool against peak, with explicit load shedding when it empties. A minimal guest also cuts the boot itself to a fraction of a conventional one.
- Does choosing a machine boundary mean the container-level controls no longer matter?No. Run unprivileged, with a read-only root and the narrowest privilege set that works, inside the guest as well. Defence in depth means the strong boundary is what remains after the cheap controls have already made an escape harder, not a substitute for them.
saying these in an interview costs you the question
- Decides by how risky the feature feels rather than by whose code executes.
- Says resource ceilings make a shared kernel safe for untrusted code.
- Claims a machine boundary is free once the images are already built.
- Treats the machine boundary as a reason to skip unprivileged, least-privilege settings.
- Puts platform credentials on the same hosts that execute customer code.