skip to content

Why is there no reliable way to sandbox untrusted Python inside the same interpreter process?

level: juniorimportance: must knowfreq 40%

answer

  1. Reflection is not optional in Python
  2. Hiding names does not hide attributes
  3. The object graph is fully connected
  4. rexec and Bastion were removed
  5. The boundary is a separate process

basics

~20 s

The Python object graph is fully connected: from any value you can follow attributes back to the type system, the builtins, and every loaded module. No matter which names you hide, code can walk to them. Real isolation is an OS-level boundary — a separate process, container, or VM.

solid answer

~40 s

Because reflection in CPython is total, not optional. Any object hands you its type via `__class__`; types chain back to `object`; and `object.__subclasses__()`, any function's `__globals__`, `gc.get_objects()`, and frame objects all reach the full set of loaded modules and objects. Removing names from the execution namespace only affects name *lookup*; it leaves every attribute *edge* intact, and the graph has redundant paths, so no finite blacklist seals it. That is why 'restricted interpreter' schemes have repeatedly been broken and why the standard library has no supported sandbox. The dependable answer is architectural: run untrusted code in a separate operating-system process — ideally inside a container or VM — with dropped privileges and CPU, memory, and time limits, so a full breakout inside that interpreter still reaches nothing valuable and can be killed.

code

python · 5 lines
python
# The 'safe' namespace hides names, but a literal still reaches object:
safe_globals = {'__builtins__': {}}
code = "result = ().__class__.__bases__[0].__subclasses__()"
exec(code, safe_globals)
print(len(safe_globals['result']) > 100)   # True: reachability survived

go deeper

for a junior

State the conclusion clearly — untrusted Python cannot be safely sandboxed in the same process — and give one reason, such as reflection reaching object and the builtins. Know that the fix is a separate process.

for a middle

Back the claim with a concrete path (object.subclasses() or a function's globals) and note that rexec/Bastion were removed because in-process restriction could not be made secure.

for a senior

Explain that the escape surface is redundant, so no blacklist works, and describe the minimum safe setup: a disposable process with dropped privileges and resource limits, escalating to container or VM for arbitrary code.

for a principal

Own the policy: treat any 'restricted interpreter' proposal as non-load-bearing, mandate an OS boundary for untrusted execution, and define the resource-limit and isolation baseline the organisation runs behind.

## The question behind the question Interviewers ask this early because it separates candidates who think of Python as a set of features you can toggle from those who understand that its dynamism is structural. 'Just turn off the dangerous parts' is the wrong mental model, and this question surfaces it. ## What 'sandbox in the same process' means The tempting design is: run `exec(user_code, safe_globals)` where `safe_globals` has no `open`, no `__import__`, an emptied `__builtins__`, maybe a curated set of helpers. The hope is that by controlling the *names* the code can see, you control what it can *do*. This is in-process sandboxing, and it does not work. ## Why it does not work: reachability Controlling names controls only name *lookup*. It does nothing about *attribute access* on the objects that are unavoidably in scope. Consider what the untrusted code always has: - Any literal it writes (`()`, `''`, `0`) is a real object with a real `__class__`. - From `__class__` it reaches `__bases__`/`__mro__`, hence the builtin `object`. - `object.__subclasses__()` lists every loaded class in the process. - Any function object exposes `__globals__`, which typically contains `__builtins__` and the module's imports. - `gc.get_objects()` returns nearly every live object regardless of what was passed in. - Frame objects (`sys._getframe`) expose the caller's globals and locals. Each of these is an *edge* in the object graph, and the graph is redundant: several independent paths lead from 'a harmless literal' to 'the real `os` module'. You cannot delete these edges — CPython does not let you reliably strip attributes off builtin types, and even a perfect seal on one path leaves the others. There is no finite blacklist that closes a graph the language keeps fully connected by design. ## History backs this up Python once shipped `rexec` and `Bastion` modules intended to restrict execution. They were removed because they could not be made secure. The language documentation is explicit that there is no supported way to run untrusted code safely in-process. This is not a missing feature; it is a consequence of the object model. ## What actually works The boundary must be the operating system: - **Separate process.** Launch the untrusted code as a child process (for example via `subprocess`), so its interpreter is disjoint from yours. A breakout inside it reaches its own objects, not your application's. - **Least privilege and limits.** Drop privileges, remove network access, use a read-only or disposable filesystem, and impose CPU, memory, and wall-clock ceilings so the code cannot exhaust the host or hang forever. - **Container or VM.** For arbitrary, internet-submitted code, wrap the process in a container or a microVM for a much stronger kernel/hardware boundary. (The mechanics of containers and OS isolation belong to their own topics; the language-level point is simply that this is where the boundary lives.) ## The one-sentence version You cannot sandbox untrusted Python by hiding names, because reflection reaches everything the interpreter has loaded; put the untrusted code in its own process behind an OS boundary instead. ## Interview framing A junior is expected to state the conclusion confidently and give at least one reason (the object graph is connected; reflection reaches the builtins). A stronger answer names a concrete path such as `object.__subclasses__()` and mentions that `rexec`/`Bastion` were removed for exactly this reason, then points at the process/container/VM boundary as the fix.

  • Python used to ship rexec and Bastion — what happened to them?
    They were restricted-execution modules meant to run untrusted code in-process, and they were removed because they could not be made secure — the same reachability problems broke them. Their removal is the official acknowledgement that in-process sandboxing is not supported, which is why modern guidance is to use an operating-system boundary instead.
  • If it can't be sandboxed in-process, what is the minimum safe setup?
    A separate, short-lived operating-system process running the untrusted code with dropped privileges, no network, a disposable or read-only filesystem, and hard CPU, memory, and wall-clock limits, so a breakout inside it reaches nothing of yours and the process can be killed on a timer. Stronger isolation (container or microVM) is warranted for arbitrary, internet-submitted code.

saying these in an interview costs you the question

  • You can sandbox Python by emptying __builtins__
  • There is a supported stdlib sandbox for untrusted code
  • rexec still exists and is safe to use
  • Hiding os and open is enough to be safe
  • In-process restriction is a real security boundary

context