What must a sandbox provide before an agent may run model-written code?
answer
- Assume the code is hostile
- A subprocess is not a boundary
- Strip ambient credentials and mounts
- Block the instance metadata endpoint
- Default-deny egress, ephemeral lifetime
basics
~20 sIsolation strong enough to assume the code is hostile: a hardened runtime or microVM rather than a shared process, no host mounts, no ambient cloud credentials or metadata access, default-deny network egress with a narrow allowlist, resource caps, and a fresh instance destroyed after each session.
solid answer
~50 sDesign the sandbox as though an attacker wrote the code, because in the hijack case they did. Start with a real isolation boundary — a microVM or a hardened container runtime such as gVisor or Firecracker-style isolation, not a subprocess on your API host, because a plain container shares the host kernel and a shared process shares everything. Give it no host filesystem mounts and an ephemeral writable layer, so nothing persists between runs. Strip ambient authority: no environment secrets, no mounted service-account token, and block the cloud instance-metadata endpoint, which is the classic path from arbitrary code execution to real cloud credentials. Default-deny egress with an explicit allowlist and controlled DNS, so the code cannot reach your internal network or an arbitrary host. Add CPU, memory, wall-clock and process caps so a runaway loop cannot starve neighbours. One sandbox per session, torn down afterwards.
code
yaml · 10 linesapiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: agent-sandbox-default-deny-egress
namespace: agent-sandbox
spec:
podSelector: {}
policyTypes:
- Egress
egress: []go deeper
Know that code an agent writes must not run next to your application, and that a sandbox means an isolated environment with no secrets, no access to your files, and no free internet access.
Explain each property and what it stops: isolation boundary versus a subprocess, ambient credentials and the metadata endpoint, default-deny egress with an allowlist, resource caps and ephemeral lifetime.
Show operational judgment — which allowlist entries you would actually grant, how denied egress feeds alerting, and how the sandbox stays independent of tool scoping so that one failing does not collapse both.
Own the platform decision: one hardened sandbox service for every agent team versus per-team implementations, the cost and cold-start tradeoffs of microVMs, and how egress policy is governed as new workloads request exceptions.
## Why the sandbox is a separate wall Giving an agent the ability to write and run code is enormously useful — it turns a fixed tool list into a general capability. It is also the point where a hijacked agent gets to execute arbitrary instructions rather than pick from a menu. Scoping cannot help you here: the whole point of a code tool is that its argument is unbounded. So the containment must come from the environment the code runs in, and the design assumption must be that the code is hostile. The sandbox is deliberately independent of the other walls. If the tool scope holds, the sandbox is redundant; if the scope is bypassed, the sandbox is what remains. Independence is the property that makes layering worth anything. ## The isolation boundary Running model-written code as a subprocess of your application is not sandboxing — it shares the filesystem, the environment variables, the network position and the process credentials of the service that spawned it. A stock container is better but still shares the host kernel, so a kernel exploit or a misconfigured mount escapes it. The practical bar in 2026 is a microVM (Firecracker-style) or a hardened container runtime that interposes on syscalls (gVisor-style), with a seccomp profile narrowing the syscall surface further, non-root execution, and a read-only root filesystem plus a small ephemeral writable layer. The exact technology matters less than the property: a compromise of the guest should not yield the host. ## Removing ambient authority Most real damage from code execution comes not from the code itself but from what the environment hands it for free: - **Environment secrets.** API keys and database URLs in the process environment are readable by anything running there. The sandbox should carry none. - **Mounted credentials.** A service-account token file mounted into the container is a credential the code can read and reuse. - **The instance metadata endpoint.** On major clouds, code running on an instance can request short-lived credentials for that instance's role from a link-local address (169.254.169.254). Blocking it is one of the highest-value single controls for any sandbox. - **Host mounts.** A bind mount for convenience — source code, a shared cache — is a path in and a path out. Copy data in explicitly instead. After this, the sandbox's power equals the data you deliberately put in it, which is the property you want. ## Default-deny egress Network policy is where most sandboxes are quietly wide open. The default in container platforms is usually unrestricted outbound traffic, which means model-written code can reach your internal services, other tenants' endpoints and any host on the internet. Default-deny with an explicit allowlist inverts that: nothing leaves unless a rule names it, DNS is served by a resolver you control so the allowlist cannot be evaded by name, and requests go through a proxy that logs what was attempted. The allowlist should be as small as the workload permits — a package index if the sandbox installs dependencies, one or two APIs the task genuinely needs. Every entry is a channel, and channels are where data leaves. ## Resource and lifetime limits Caps on CPU shares, memory, wall-clock time, process count and disk turn a runaway or deliberately abusive workload into a killed job rather than an outage for everything sharing the host. Pair them with a strict lifetime: one sandbox per session or per task, destroyed afterwards, so nothing an attacker plants survives to the next run. Persistent sandboxes accumulate state, and accumulated state is a foothold. ## What good failure looks like When the sandbox does its job, a hijacked agent's code run ends as a killed process with a denied network attempt in the proxy log and no credential ever readable. That is also the signal to alert on: denied egress from a sandbox is nearly always either a broken task or an attack, and both deserve a look. ## Common shortcuts and what they cost Teams frequently ship a plain container on the app host with default networking, because it is a two-line change and the demo works. The cost is that arbitrary code now runs inside the trust boundary of the service, with its environment, its network position and its instance role. The second common shortcut is a long-lived shared sandbox reused across users, which turns any successful plant into a cross-user compromise. ## Answering this in an interview Say the assumption out loud — treat the code as hostile — then list the properties in order of leverage: real isolation boundary, no ambient credentials or metadata access, no host mounts, default-deny egress, resource caps, ephemeral lifetime. Naming the metadata endpoint and default-deny egress specifically is what distinguishes an answer that has run this in production.
- Why is blocking the cloud instance metadata endpoint singled out so often?Because it is the shortest path from arbitrary code execution to real cloud credentials. Code inside the sandbox can request short-lived credentials for the host instance's role from a link-local address and then act with that role's permissions, entirely bypassing your tool scoping. Blocking that address, or running the sandbox where no instance role is attached, closes the escalation in one rule.
- The task needs to install packages, so egress cannot be fully denied. What do you do?Keep default-deny and allowlist exactly one destination: an internal package mirror you control, reached through a logging proxy. That gives the workload what it needs without opening arbitrary outbound traffic, and the mirror gives you a review point for what can be installed. Better still, pre-build the image with the dependencies so the sandbox needs no egress at all.
- Is a fresh sandbox per session really necessary, or is cleanup enough?Fresh instances are strictly safer and usually cheaper to reason about. Cleanup scripts have to enumerate everything an attacker could have planted — cron entries, shell profiles, cached packages, background processes — and any gap persists into the next session, potentially another user's. Destroying and recreating the instance makes persistence structurally impossible rather than a matter of completeness.
saying these in an interview costs you the question
- Running model-written code as a subprocess of the API service
- Assuming a stock container is an adequate security boundary
- Leaving default outbound network access enabled in the sandbox
- Mounting the repo or a service-account token in for convenience
- Reusing one long-lived sandbox across sessions and users