Beyond attribute chains, how do gc.get_objects() and frame references defeat an in-process Python sandbox?
answer
- One door is never the only door
- The collector can list the heap
- Frames climb into the caller
- gc.get_objects() and f_back
- __code__ plus types.FunctionType rebuild callables
basics
~20 sgc.get_objects() returns nearly every live object the collector tracks, so untrusted code can enumerate the whole heap and pick up references it was never handed. Separately, a frame from sys._getframe() exposes the caller's globals and locals through f_globals, f_locals and f_back. Both reach past any namespace you built.
solid answer
~50 sAttribute traversal is only one door. `gc.get_objects()` hands back a list of essentially every live Python object the cyclic collector tracks — including your host application's loaded modules, connection objects, and secrets — none of which the untrusted code was passed. `gc.get_referrers()` then walks the graph in the other direction. Independently, `sys._getframe()` yields a frame object whose `f_globals`, `f_locals`, and `f_back` let untrusted code climb the call stack into the trusted caller and read or rebind its variables. Code objects reached through a function's `__code__`, together with `types.FunctionType`, allow rebuilding callables. Every one of these is ordinary attribute and function access, so a namespace-level restriction cannot remove them. This redundancy is the point: even if you sealed attribute traversal, gc and frames provide equivalent paths. The conclusion is the same — the boundary must be the OS process, not the interpreter.
code
python · 6 linesimport gc
secret = ['api-key-nobody-handed-to-untrusted-code']
# Untrusted code never received `secret`, yet the heap scan finds it:
found = any(o is secret for o in gc.get_objects())
print(found) # Truego deeper
Know that Python offers reflection into the running program: you can list live objects and inspect the call stack. The lesson is that untrusted code sees far more than what you handed it.
Be able to name gc.get_objects() and sys._getframe with f_back/f_globals as escape paths and explain that each works independently of the attribute chain, so patching one exploit is not a fix.
Demonstrate that the escape surface is broad and redundant, that audit hooks observe rather than prevent, and that the only sound response is an OS-level boundary with resource limits — not a cleverer namespace.
Frame the tradeoff for the org: since no in-process technique is sound, standardize on a disposable-process/container/VM boundary and treat any 'restricted interpreter' proposal as defense-in-depth at best, never as the boundary.
## Why one escape technique is never enough to reason about A sandbox designer who blocks the `__subclasses__()` chain often believes the job is done. It is not, because CPython exposes several *independent* ways to reach objects the untrusted code was never handed. This leaf's second question is about those parallel doors: the garbage collector's introspection API and frame reachability. Understanding them is what separates 'I patched one exploit' from 'in-process sandboxing is unsound'. ## `gc.get_objects()`: the whole heap in one call The cyclic garbage collector must track container objects to detect reference cycles. `gc.get_objects()` returns a list of the objects it tracks — in practice, nearly every non-trivial live object in the process: modules, classes, functions, dicts, your database connections, cached credentials, the real `os` module if anything imported it. Untrusted code that can call `gc.get_objects()` does not need to *traverse* to anything; it can simply *scan* for what it wants by type or attribute: ```python import gc secret = ['api-key-nobody-passed-to-untrusted-code'] found = [o for o in gc.get_objects() if o is secret] print(bool(found)) # True ``` `gc.get_referrers(obj)` and `gc.get_referents(obj)` then let the attacker walk the object graph both directions from any object they found, reconstructing relationships. If your defense was 'I only passed the untrusted code a tiny, safe object', the heap scan makes that irrelevant. ## Frames: climbing into the trusted caller A frame object represents an execution scope. `sys._getframe()` returns the current frame; `f_back` walks up the call stack; `f_globals` and `f_locals` expose the namespaces of each frame. So untrusted code that runs *inside* a call from your trusted host can climb `f_back` until it reaches the host's frame and then read `f_globals` — the host module's globals, including whatever it imported and whatever secrets live there. Locals can be read, and in CPython globals can be rebound, so the untrusted code can even alter the trusted caller's behaviour. Frame and code-object introspection as a general subject belongs to the runtime leaf; here it matters only as an escape vector: the call stack itself is reachable data. ## Code objects and rebuilding callables Any function exposes its compiled `__code__` object. With `types.FunctionType(code, globals)` an attacker can wrap a code object in a fresh function bound to a globals mapping of their choosing — for instance, one that includes the real builtins pulled from a scanned module. This is another reason blacklisting names is futile: the raw materials to assemble a fully-powered callable are lying around as ordinary attributes. ## The unifying lesson: redundancy The reason in-process sandboxing cannot be made safe is not any single API — it is that the doors are *redundant and numerous*. Attribute traversal, `gc.get_objects()`, `gc.get_referrers()`, frame walking, and `__code__` + `types.FunctionType` all reach the same place: the full set of loaded objects and modules. Sealing one requires monkeypatching builtin behaviour that CPython does not let you remove reliably, and even a perfect seal on one leaves the others. Audit hooks (`sys.addaudithook`) can *observe* some of this, but observation is not prevention, and hooks themselves run in the same process the attacker controls. ## What holds The boundary must be the operating system. Run untrusted code in a separate, disposable process with dropped privileges, no network, a read-only or throwaway filesystem, and hard CPU/memory/wall-clock limits — or in a container or microVM for stronger isolation. Then a total breakout inside the untrusted interpreter still touches nothing of yours and can be killed on a timer. Note that per-interpreter isolation via subinterpreters (added to the stdlib in 3.14) isolates module state but shares the OS process, so it is not a security boundary either. ## Interview framing A senior candidate is expected to go past the single famous exploit and articulate that the escape surface is broad and redundant: name `gc.get_objects()` and frame reachability specifically, explain that each independently defeats a namespace restriction, and conclude with the OS boundary. Mentioning that audit hooks observe but do not contain, and that subinterpreters are not a security boundary, is the mark of production judgment.
- Do audit hooks via sys.addaudithook turn in-process execution into a sandbox?No. Audit hooks let you observe and log certain sensitive operations, which is useful for detection, but they run in the same process the attacker controls and do not cover every escape path (a heap scan or frame read raises no audit event). Observation is not containment; a determined attacker still reaches the objects. Hooks complement an OS boundary, they do not replace it.
- Are 3.14 subinterpreters a security boundary for untrusted code?No. concurrent.interpreters (PEP 734, stdlib in 3.14) gives each interpreter its own module state and, optionally, its own GIL, which helps isolate state and improve parallelism. But they share one OS process and address space, so a breakout or a C-level crash affects the whole process, and resource exhaustion is not contained. They isolate state, not trust.
saying these in an interview costs you the question
- Blocking __subclasses__ makes the sandbox safe
- Untrusted code only sees objects you pass it
- gc is a debugging tool with no security impact
- Audit hooks prevent the escape rather than observe it
- Subinterpreters isolate untrusted code securely