A product feature genuinely requires running commands or programs that users supply — a build runner, a plugin hook, a conversion pipeline fed by user files. You cannot allow-list the command. How do you build this safely?
answer
- you cannot filter your way there — relocate the trust boundary
- no ambient credentials; no credential endpoint reachable
- container = shared kernel; hostile code wants a VM
- default-deny egress; process caps; hard timeout; ephemeral per job
- output is untrusted input: logs, archive entries, manifests
basics
~20 sStop trying to make the input safe and relocate the trust boundary: execution happens inside a disposable, unprivileged, isolated environment with no ambient credentials, no default egress and hard resource limits. Everything it produces — output, logs, filenames — is then untrusted input to you.
solid answer
~60 sOnce arbitrary execution is the *feature*, input filtering is meaningless; the control is what the execution can reach. Build it as an isolation problem. **Identity**: a dedicated principal with no shared secrets, no credential-endpoint reachability, and none of the orchestrator's tokens in its environment. **Boundary sized to the threat**: a container shares a kernel, so genuinely hostile code warrants a virtual machine or a user-space kernel. **Filesystem**: read-only root, a scratch mount with no execute or setuid, no host paths or management sockets mounted. **Network**: default-deny egress with an explicit allow-list. **Resources**: CPU, memory, process-count limits and a hard wall-clock timeout. **Lifecycle**: one ephemeral instance per job, destroyed afterwards — never reused across tenants, or one job poisons the next through caches. Then the parts that get missed: everything coming back is attacker-controlled — logs rendered into a UI, artefact names, exit metadata — and the sandbox itself is now your security boundary, so its escape surface (kernel version, mounted devices, shared namespaces) is your surface, with a patching story and a written threat model.
code
text · 13 linesIN (what the job receives)
- job spec built by orchestrator: argv array, code-owned structure
- inputs mounted read-only; scratch mounted noexec,nosuid
- env: minimal, no orchestrator tokens, no cloud creds
- net: default-deny + explicit allow-list; credential endpoint blocked
- limits: cpu, mem, pids, io, hard wall-clock timeout
- lifetime: fresh instance, destroyed after, never reused across tenants
OUT (what the job returns — ALL untrusted)
- stdout/stderr -> size-bounded, escaped on render
- artefacts -> names validated, archive entries checked on unpack
- manifest -> parsed defensively; never selects a credential
- exit status -> a number, not a control instructiongo deeper
Recognise that the answer is isolation rather than input filtering, and name the obvious controls: a separate low-privilege user, a container, limits and a timeout.
Add no ambient credentials, no network by default, ephemeral instances, and the fact that job output must be treated as untrusted input.
Choose the isolation strength deliberately, justify egress control and per-job identity, and cover archive extraction and log rendering as attacks in the outbound direction.
Own the whole boundary: threat model in writing, patching and escape-response plan, per-tenant separation for catastrophic cases, the density-versus-isolation cost curve, and a policy that the orchestrator's own invocations still follow the caller-owns-structure rule.
## Accept the premise, then move the boundary Every other answer in this area says: do not let untrusted data decide structure. Here structure *is* the product — a build runner exists to run whatever the customer's configuration says. So the invariant cannot be preserved at the call site, and the design move is different in kind: you stop trying to make the code safe and instead make the *region in which it executes* untrusted by design. Everything the region can touch becomes the security control, and the security argument shifts from "the input cannot alter the structure" to "the structure can be anything, and here is precisely what it can reach." A principal-level answer states that reframing first, because it determines everything after it. ## Identity: nothing ambient The most common real-world compromise of such systems is not a sandbox escape at all — it is credential theft, because the job inherited something it should never have seen. So: - a dedicated principal per job, with no shared long-lived secret - no orchestrator tokens, deployment keys or registry credentials in the job's environment - no reachability to any local credential-issuing endpoint; treat that as a network rule, not a convention - any secret the job legitimately needs is scoped to that job, short-lived, and injected through a mediated channel rather than the ambient environment - one job's identity can never read another job's state ## Boundary: size it to the threat Isolation mechanisms are not interchangeable. A container is a set of kernel features sharing one kernel with the host, which is appropriate for *semi-trusted* workloads — your own code, your own images. For code supplied by arbitrary users, the shared kernel is a large attack surface, and the appropriate boundary is a virtual machine, a lightweight VM per job, or a user-space kernel that intercepts system calls. State the trade-off honestly: stronger boundaries cost start-up latency and density, and that is the price of the feature. Within whichever boundary you choose: drop all capabilities that are not required, apply a system-call filter, disable privilege elevation, use separate user and network namespaces, and never mount a management socket or a host path into the job — a mounted control socket is equivalent to giving the job the host. ## Filesystem, network, resources, lifetime **Filesystem**: read-only root image; a scratch area mounted without execute and without setuid honouring; no host directories; explicit quotas so a job cannot fill shared storage. **Network**: default-deny egress with a narrow allow-list of what the feature genuinely needs. This single control blocks fetching second-stage payloads, blocks exfiltration, and blocks reaching internal services that trust the network — three of the most valuable things an attacker wants. Also block the job from reaching your own control plane. **Resources**: memory and CPU limits, a process/thread cap (otherwise a fork bomb is your availability incident), an I/O limit, and a hard wall-clock timeout enforced by the orchestrator rather than by anything inside the job. **Lifetime**: one fresh, ephemeral instance per job, destroyed afterwards. Reuse across tenants is how one job poisons the next — through a shared dependency cache, a modified tool on the path, a leftover file, an environment change. If caches must be shared for performance, they are read-only to the job, or content-addressed and verified, and never writable by one tenant and readable by another. ## The direction people forget: output is untrusted input The job is adversarial, so everything crossing back out of the boundary is attacker-controlled data arriving at *your* trusted code: - **Logs** rendered into a web UI — the job controls every byte, including anything the renderer might interpret; escaping on output is the neighbouring discipline and it applies in full here. - **Artefact names and archive entries** — unpacking a job-produced archive into a directory is exactly the path-handling problem the traversal topic owns, and it is a very common real bug in runners. - **Structured results** — an exit code, a report file, a manifest the orchestrator parses. Parse defensively, bound sizes, and never let job output decide orchestrator control flow such as which credential to use next. - **Volume of output** — an unbounded log stream is a denial-of-service against your storage and your UI. ## The sandbox is now your security boundary And therefore its escape surface is your surface. That has organisational consequences you should name: you now own a kernel and hypervisor patching story with a defined urgency; you must know exactly what devices, shared namespaces or accelerators are exposed and why; you should assume escapes will eventually happen and add a second boundary for anything that would be catastrophic (per-tenant node pools, separate accounts or projects, network segmentation so an escaped job lands somewhere with nothing worth taking). Write the threat model down — what the job may do, what it may reach, what an escape would yield — because these systems accumulate convenience features that quietly erode the boundary. ## And the ordinary rules still apply inside Finally, do not lose the basics because the feature is exotic. The orchestrator's own invocation of the job is still a caller-owns-structure problem: the job specification is code-owned structure with user values in value positions, launched with an argument vector, with the end-of-options convention observed. Users supply the *program to run inside the sandbox*; they do not supply pieces of your orchestrator's command lines. Blurring that distinction is how sandboxed systems get compromised at the layer above the sandbox.
- Why is a container often the wrong boundary for genuinely user-supplied code?Because containers share the host kernel: isolation is enforced by kernel features, so the entire system-call surface is reachable by the job and a single kernel vulnerability crosses the boundary. That is an acceptable risk for your own code, which is not trying to escape. For code arbitrary users supply, the stronger option is a lightweight virtual machine per job or a user-space kernel that intercepts system calls, accepting slower start-up and lower density in exchange for a much smaller shared surface.
- What goes wrong if job instances are reused between customers for performance?One job can leave state that the next one inherits: a poisoned dependency cache, a modified tool earlier on the search path, a leftover credential or file, a changed environment. That turns a low-privilege customer into an attacker against the next customer's build, which is often more valuable than attacking you directly. If caching is essential, make caches content-addressed and verified, or read-only to the job, and never writable by one tenant and readable by another.
- Give a concrete example of the job's *output* attacking the platform.A build produces an archive of artefacts that the platform unpacks into a storage directory. If entry names are not validated, an entry can point outside the destination and overwrite platform files — the classic archive-extraction path problem, which the path-handling topic covers in detail. Log rendering is the other common one: the job controls every byte of its log, so a UI that renders it without escaping is executing attacker-chosen content in an operator's browser session.
You are not vetting what customers bring into the building; you are giving each of them an identical empty room with a locked door, no phone line, nothing of yours inside, and a cleaning crew that demolishes the room afterwards.
saying these in an interview costs you the question
- Trying to solve it with a bigger deny-list of dangerous commands.
- Mounting a container-management socket or a host directory into the job for convenience.
- Leaving cloud credentials or orchestrator tokens in the job's environment because "the sandbox will contain it."
- Allowing unrestricted outbound network, which enables both second-stage payloads and exfiltration.
- Reusing job instances across tenants for speed.
- Treating job logs, artefact names and manifests as trustworthy because they came from your own system.