skip to content

Why is passing {'__builtins__': {}} as eval()'s globals not a sandbox for untrusted expressions?

level: seniorimportance: must knowfreq 55%

answer

  1. It changes lookups, not authority
  2. Every object is reflective
  3. The recipes keep getting broken
  4. Nothing there bounds time or memory
  5. Parse and inspect before deciding to run

basics

~10 s

Because it only hides convenient names. Every object still reachable from the expression carries attributes that lead back to arbitrary code, and nothing bounds CPU, memory or time. It is inconvenience, not isolation.

solid answer

~50 s

Emptying `__builtins__` removes the shortcut names an attacker reaches for first, and that is all it does. It is a **name-resolution** change, not an authority boundary: any object still visible - one you passed in the locals mapping, or one a literal produces - exposes attributes that lead back to arbitrary execution, which is why every published "restricted eval" recipe has eventually been broken. Even an airtight namespace would not bound resources: an expression with no names at all can exhaust memory or spin for hours, and `eval()` takes no timeout argument. The safe door is to stop executing the text: parse it with `ast.parse()`, reject every node type you did not deliberately allow - calls, attribute access, subscripts - and compile only what survives. If arbitrary code genuinely must run, the boundary has to be the operating system, not a dict.

code

python · 15 lines
python
import ast

def guarded_eval(src, fields):
    tree = ast.parse(src, mode="eval")
    for node in ast.walk(tree):
        if isinstance(node, (ast.Call, ast.Attribute, ast.Subscript)):
            raise ValueError(f"{type(node).__name__} is not allowed")
    code = compile(tree, "<formula>", "eval")
    return eval(code, {"__builtins__": {}}, dict(fields))

print(guarded_eval("qty * price * 1.2", {"qty": 3, "price": 10}))
try:
    guarded_eval("qty.__class__", {"qty": 3})
except ValueError as exc:
    print("rejected:", exc)

go deeper

for a junior

Take away the rule rather than the mechanics: untrusted text never goes to eval(), and an empty builtins mapping does not change that. If you see the pattern in a review, flag it and ask what the input actually is.

for a middle

Explain what the mapping argument really changes - name resolution - and why Python's reflective object model makes that an unsound boundary. Be able to sketch the ast.parse allowlist as the alternative rather than just saying "it's dangerous".

for a senior

Show production judgment: name the availability failure as well as the code-execution one, describe the allowlist plus resource-budget design, and be honest that a dict is never the boundary when arbitrary code genuinely has to run.

for a principal

Own the decision itself. Argue whether the product needs expressions at all versus a fixed set of computed fields, what the blast radius and operational budget of the chosen design are, and who reviews the allowlist when the language grammar grows.

## The setup this question always comes from A feature lands that lets people write a small formula. Take an invoice-PDF renderer where a template author can type an expression for a computed line total - `qty * unit_price * (1 + vat)`. Someone reaches for `eval()`, notices that is alarming, and writes what looks like the careful version: ```python eval(formula, {"__builtins__": {}}, fields) ``` It passes review because it *looks* locked down. It is not. ## What the empty builtins mapping actually does When `eval()` receives a globals dict without a `"__builtins__"` key, the interpreter injects a reference to the builtins module so that ordinary names work. Supplying an explicit empty mapping suppresses that injection. The effect is precisely one thing: names like the file-opening builtin, the import machinery's entry point, the object-introspection builtins and the compile/eval/exec trio are no longer resolvable *by name*. That is a name-resolution change. It removes convenience, not capability. Python's object model is uniformly reflective: every object exposes its class, and from a class the rest of the runtime is a chain of ordinary attribute accesses away. An expression that can perform attribute access on *any* live object has a path back to arbitrary callables. It does not matter that the starting object is a plain integer you passed in as a field value. This is why "restricted eval" is a recurring, and recurringly broken, genre: each published recipe blocks the previously-published escape, and a new one appears. Treat that history as the evidence - the specific escape chains are a study of their own, but the conclusion for design purposes is settled. ## The half nobody blocks: resources Suppose, arguendo, the namespace were airtight. `eval()` still has no time limit, no memory limit and no allocation ceiling; there is no timeout parameter and no release has ever added one. An expression consisting only of integer literals and operators - no names, no calls, nothing your validator would flag - can occupy a core for a very long time and drive the process into swap or into `MemoryError`. In the invoice renderer that failure does not look like an attack. It looks like an **intermittent timeout**: some PDF jobs come back, some hang until the request times out, retries make it worse because the retried job runs the same formula, and the metric that moves is render latency, not anything security-shaped. Nothing in a 27-minute test suite reproduces it, because the suite runs the formulas the team wrote, not the one a template author saved last Tuesday. Availability is the part of a restricted-eval design that fails first and gets diagnosed last. ## The safe door: parse, inspect, decide, then run `ast.parse()` is the mechanism that changes the shape of the problem. It compiles source into an abstract syntax tree **without executing anything**, which turns "run this and hope" into "look at this and decide". A validator walks the tree and rejects node types outright: ```python import ast ALLOWED = (ast.Expression, ast.BinOp, ast.UnaryOp, ast.Constant, ast.Name, ast.Load, ast.Add, ast.Sub, ast.Mult, ast.Div, ast.USub) def check(src): tree = ast.parse(src, mode="eval") for node in ast.walk(tree): if not isinstance(node, ALLOWED): raise ValueError(f"{type(node).__name__} is not allowed") return tree ``` An **allowlist** is the only defensible shape. A denylist of "dangerous" node types or substring filtering on the source text both fail for the same reason: they enumerate what you thought of. With calls, attribute access and subscripts excluded up front, the reflective paths that make the restricted-namespace approach hopeless are simply not expressible, because the *grammar* that reaches them never compiles. Then bound the rest: cap the source length and the tree depth before you walk it, cap the magnitude of numeric literals if exponentiation is allowed at all (or exclude the power operator, which is the cheapest single fix for the resource problem), and give the evaluation a wall-clock budget in a worker you can kill. ## When you genuinely need arbitrary code Sometimes the requirement really is "run the user's Python". Then the boundary must be one the operating system enforces around a separate process - not a dictionary inside your interpreter. That is a different design with its own long list of considerations, and the honest senior answer is to name it as a different design rather than to reach for a cleverer mapping. ## What to say in the interview Three beats. First: the empty builtins mapping changes name resolution and nothing else, and the object model makes name resolution an unsound boundary. Second: even a perfect namespace leaves CPU and memory unbounded, and in production that is the failure you actually meet. Third: the fix is to stop executing untrusted text - parse it, allowlist the node types, and evaluate only a subset you defined - reserving process isolation for the case where arbitrary code is a real requirement.

  • The invoice renderer starts showing intermittent render timeouts weeks after formulas shipped. How would you connect that to eval()?
    Look for a saved formula that is expensive rather than malicious. A single arithmetic expression can occupy a worker indefinitely, and because the same job is retried, one bad template poisons a queue while looking like flaky latency. Reproduce by replaying stored formulas against a budgeted worker, then add a wall-clock cap, cap literal magnitude, and exclude exponentiation.
  • If the expression may only touch the invoice's own fields, is passing them in the locals mapping enough of a restriction?
    No. Every object you place in the mapping is a live Python object with a full attribute surface, so the restriction only holds if the *grammar* cannot express attribute access, subscripting or calls. Enforce that at parse time with an allowlist, and keep what you inject down to plain numbers and strings so that even a validator gap has little to work with.
  • Why is an allowlist of AST node types preferred over stripping dangerous substrings from the source?
    Because substring filtering enumerates what you thought of, and Python has many spellings for the same thing - whitespace, string concatenation in the source, alternative encodings of a name. An allowlist inverts the burden: anything you did not explicitly permit fails, so new syntax and unfamiliar spellings are rejected by default rather than slipping through.

It is like removing the signposts from a building rather than locking the doors: a visitor who knows the layout still walks anywhere, and the fire code was never about signage.

saying these in an interview costs you the question

  • Says empty builtins makes eval() safe for user input
  • Proposes a denylist of dangerous words in the source
  • Assumes an expression without names cannot do harm
  • Forgets that eval() has no time or memory limit
  • Thinks passing only plain values closes the attribute surface
  • Treats try/except around eval() as containment

context