skip to content

Dynamic Execution Risks

Every route by which a string becomes running code, and the safe alternative to each: the compile and evaluate family, a name resolved dynamically, literal-only parsing, and why sandboxes never hold.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

16

Why is Python's eval() a code-execution risk on a string that came from a user?

level: juniorimportance: must knowfreq 65%

answer

  1. It runs the text, not just reads it
  2. One expression is still the whole language
  3. Builtins reach modules without an import statement
  4. Parsing builds a tree; evaluating runs it

basics

~20 s

eval() compiles and runs its text as a Python expression, so one expression can import a module and touch files, processes or sockets with the privileges of your process. Parse untrusted text; never evaluate it.

solid answer

~50 s

`eval(text)` compiles the text in expression mode and executes it, returning the value — it is the interpreter, not a calculator. "A single expression" still gives an attacker calls, attribute access, subscripting, comprehensions and the walrus operator, and because `__import__` is an ordinary builtin, `__import__("os")` reaches any module from inside one expression; that code runs with your process's permissions and secrets. Word blocklists lose, since an expression has unbounded spellings, and `try`/`except` catches errors rather than effects. The rule is to keep parsing separate from executing: `ast.parse()` and `compile()` build a tree or a code object without running anything, while only `eval()` or `exec()` runs it. If you need a value, accept a data format and parse it; if you must accept Python-shaped text, walk the parsed tree and refuse every node kind you did not choose to support.

code

pycon · 4 lines
pycon
>>> eval("2 + 2")
4
>>> eval("__import__('os').getpid() > 0")
True

go deeper

for a junior

Be ready to say plainly that eval() runs its text as Python and hands back the value, so user-supplied text runs as code with your program's privileges. Know that the answer to "how do I make eval() safe on this input" is to stop evaluating that input.

for a middle

Explain the mechanics: eval compiles in expression mode, executes in the caller's namespace when you pass no mappings, and reaches any module through builtins such as import. Show the parse-versus-execute boundary — ast.parse and compile run nothing; eval and exec do.

for a senior

Demonstrate production judgement. Find where evaluated text has crept in — config, templates, repr round-trips — classify it as remote code execution rather than a lint nit, and replace it with a parsed data format or a vetted tree walk, with a regression test that asserts hostile text is refused.

for a principal

Own it as policy. Evaluating external text is a design decision with a blast radius, not a review comment, so decide whether the product needs customer-supplied expressions at all; if it does, fund an owned mini-language plus process or container isolation rather than a filtered eval() someone will loosen later.

## What `eval()` actually does `eval(source)` accepts a `str`, a `bytes` object or an already-compiled code object. When it is text, CPython compiles it in expression mode and then *executes* the resulting code object, returning whatever value the expression produced. That word — **executes** — is the whole answer. `eval()` is not a parser, a calculator or a type converter; it is the same interpreter that is running the rest of your program, aimed at a string that came from somewhere else. With no mapping arguments it even runs against the caller's own namespace, so your module's names are in scope for it. ## "It is only an expression" bounds nothing People reason that `eval()` refuses statements — `eval("x = 1")` really does raise `SyntaxError` — and conclude the damage is capped at arithmetic. It is not. A Python expression may contain: - function calls, attribute access and subscripting; - conditional expressions; - comprehensions with their own loops; - and the walrus operator, which binds a name. Every builtin is in scope, and `__import__` is an ordinary builtin function: `__import__("os")` inside an expression reaches any importable module, and through it the entire standard library — reading and writing files, spawning subprocesses, opening sockets, reading environment variables. One line of user-supplied text is therefore **arbitrary code execution** running as your process, with your file permissions, your secrets in the environment and your position inside the network. ## The trust boundary, not the function, is the subject `eval()` is legitimate where the text is yours: a debugger, a REPL, a tool that generates code from source you shipped. The bug is the source of the string, and that source moves. Text you call "internal" arrives: - from config files that a deploy pipeline templates, an admin screen writes, or a mounted volume supplies; - from a "just evaluate the `repr()`" round-trip somebody added to a cache; - from a templating layer whose expressions compile to Python; - from a product feature such as an invoice-PDF renderer that lets a customer configure how a total is computed. "We control that file" is a statement about today's deployment, not an invariant of the system. ## Why the usual patches fail - **Blocklisting** words like `import` or `os` before evaluating loses immediately, because an expression has unbounded spellings — string concatenation, `chr()` arithmetic, escapes, attribute chains — so a filter is a puzzle for the attacker rather than a boundary. - Wrapping the call in `try`/`except` catches exceptions, not effects: the file was already deleted before anything was raised. - A **timeout** bounds one flavour of harm and does nothing about exfiltration. - Shrinking the mapping you hand `eval()` narrows the *convenient* names but is not a security boundary either. None of these convert code into data, which is the only change that helps. ## The safe door: keep parsing separate from executing `ast.parse(text)` builds an **abstract syntax tree** from the text and executes none of it. `compile(text, "<user>", "exec")` likewise produces a code object and runs nothing; only `eval()`, `exec()` or calling that object executes. That boundary is a real tool, not just a distinction. If you truly must accept Python-shaped text: 1. parse it, 2. walk the resulting tree, 3. reject every node kind outside a small set you deliberately chose to support, 4. and then evaluate that vetted structure with an interpreter you wrote — so the user supplies data your code interprets, never code the interpreter runs. If what you actually wanted was a value, accept a data format and use its parser; the Python interpreter was never the right reader for configuration. ## What parsing does not buy - Parsing is not a sandbox: parse a hostile tree and then `exec` it and you are exactly where you started. - Parsing is not free either — an enormous input costs memory and CPU whether or not you run it, and CPython's parser rejects input nested past its limit with a `SyntaxError`, which is still work done before the refusal. - A tree-walking evaluator you write is only as tight as its accept list, and a hand-written one that forgets to bound recursion has an availability bug. Those are resource-exhaustion and correctness problems, though — a smaller and much better-behaved class than remote code execution, and that downgrade is the whole point of the exercise. ## How to answer it in an interview 1. Lead with the mechanism: `eval()` compiles and runs the text as an expression, expressions reach builtins, so evaluating text you did not write is remote code execution. 2. Then say the fix is a **change of category** rather than a hardening step — accept data, not code. 3. Then, if pressed, show the parse-versus-execute boundary and the vetted tree walk. What loses the point is arguing about which filter would finally be long enough.

  • A product feature must let a customer supply a formula for an invoice total. How would you build that without eval()?
    Define a tiny expression language you own. Parse the submitted text with `ast.parse()` in expression mode, walk the tree, and reject any node kind outside the small set you support — numbers, the named fields you expose, arithmetic operators and comparisons — then evaluate that vetted tree with your own walker. Better still, drop free text and accept a structured formula, an operator plus operands, from the client. Either way the customer supplies data your code interprets, never code the interpreter runs.
  • Where does evaluated text creep into a codebase without anyone deciding to run user code?
    Through indirection. A cache or message format that stores a `repr()` and evaluates it on the way back; a templating layer whose expressions compile to Python; a rules or filter field in a configuration file that the deploy pipeline templates or an admin screen writes; a plugin loader that evaluates a settings string. In each case nobody wrote `eval(request.body)` — the text simply travelled far enough from its author that the trust claim stopped being true.
  • If the string only ever comes from a config file your own team writes, is eval() acceptable?
    Treat it as untrusted anyway. Config files are written by pipelines, admin tooling, mounted volumes and support staff, so "we control it" describes today's deployment rather than a property of the system, and any write path into that file becomes a code-execution path into the service. It also makes the config part of your executable surface, invisible to type checkers and linters. Prefer a data format with a real parser and compute derived values in code from named settings.

Calling eval() on user text is not offering a calculator button; it is handing the user a shell prompt that already holds your process's file access, environment secrets and network position.

saying these in an interview costs you the question

  • Says eval() only handles literals and arithmetic
  • Claims a blocklist of words like import or os makes eval() safe
  • Thinks eval() is harmless because expressions cannot contain statements
  • Believes wrapping eval() in try/except neutralises the risk
  • Assumes text from a config file or an internal caller is trusted
  • Confuses parsing or compiling the text with executing it

context

open as a page

Why is there no reliable way to sandbox untrusted Python inside the same interpreter process?

level: juniorimportance: must knowfreq 40%

basics

~20 s

The Python object graph is fully connected: from any value you can follow attributes back to the type system, the builtins, and every loaded module. No matter which names you hide, code can walk to them. Real isolation is an OS-level boundary — a separate process, container, or VM.

open as a page

Why does importlib.import_module on a user-supplied name amount to arbitrary code execution?

level: middleimportance: must knowfreq 55%

basics

~20 s

Importing a module executes that module's top-level code. If the caller names the module, the caller chooses which code runs, and search order plus every installed distribution decide what that name resolves to. Import from a fixed allowlist instead.

open as a page

Which expressions does ast.literal_eval accept, and which does it refuse?

level: middleimportance: must knowfreq 60%

basics

~10 s

ast.literal_eval evaluates only literal structures: strings, bytes, numbers, booleans, None, Ellipsis, and the tuple, list, dict and set displays built from them. Any name, attribute, operator or function call raises ValueError instead of running.

open as a page

How can Python code with no builtins reach os.system by walking __class__ and __subclasses__?

level: middleimportance: must knowfreq 45%

basics

~20 s

Every object exposes class, and from there bases reaches object, whose subclasses() lists every class the interpreter has already loaded. An attacker scans that list for a gadget class whose method or globals reaches os or subprocess, so stripping builtins never removes the path back to dangerous code.

open as a page

Why is passing {'__builtins__': {}} as eval()'s globals not a sandbox for untrusted expressions?

level: seniorimportance: must knowfreq 55%

basics

~10 s

Because it only hides convenient names. Every object still reachable from the expression carries attributes that lead back to arbitrary code, and nothing bounds CPU, memory or time. It is inconvenience, not isolation.

open as a page

Why is getattr(obj, name) unsafe when the name string comes from user input?

level: juniorimportance: should knowfreq 45%

basics

~20 s

Python's getattr does an ordinary attribute lookup, so a user-chosen string can reach any attribute the object has: private underscore names, inherited methods and dunder attributes alike. Route user input through a dict of permitted names instead.

open as a page

When should you parse a string with json.loads instead of ast.literal_eval?

level: juniorimportance: should knowfreq 46%

basics

~20 s

Whenever the string is JSON — anything produced by another system. json.loads speaks the actual format, including true, false and null, reports errors with a position, and runs roughly thirty times faster than parsing Python source into a tree.

open as a page

In Python, what do compile()'s 'eval', 'exec' and 'single' modes each produce?

level: middleimportance: should knowfreq 30%

basics

~20 s

All three return a code object without running anything. 'eval' accepts one expression whose value comes back from eval(); 'exec' accepts a module of statements; 'single' accepts one interactive statement and prints an expression's value through sys.displayhook.

open as a page

In Python, what do exec()'s globals and locals mapping arguments control, and where does a name bound by the executed code end up?

level: middleimportance: should knowfreq 45%

basics

~20 s

They are the namespaces the executed code reads and writes. Reads fall back locals to globals to builtins; new bindings land in the locals mapping. Omit both and the code uses the caller's namespaces instead.

open as a page

How do you harden a getattr-based dispatcher that genuinely has to stay dynamic?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Test the incoming name for membership in an explicit set of permitted names before the lookup, resolve against a purpose-built namespace object rather than a live service class, and verify the result is the kind of callable you expect. Never filter by rejecting underscores.

open as a page

What goes wrong when setattr writes object fields named by a client-supplied payload?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Looping setattr over a request dict lets the caller rebind any attribute the object has, including internal limits and flags, and create new ones. Nothing raises, so the effect is silent. Check each key against an allowlist of writable fields first.

open as a page

Why can ast.literal_eval exhaust memory or CPU on untrusted input?

level: seniorimportance: should knowfreq 32%

basics

~20 s

It guarantees no code runs, not that parsing is cheap. The parser builds a full syntax tree before the whitelist is consulted, so a large flat literal costs seconds of CPU and gigabytes of memory without executing anything.

open as a page

Beyond attribute chains, how do gc.get_objects() and frame references defeat an in-process Python sandbox?

level: seniorimportance: should knowfreq 33%

basics

~20 s

gc.get_objects() returns nearly every live object the collector tracks, so untrusted code can enumerate the whole heap and pick up references it was never handed. Separately, a frame from sys._getframe() exposes the caller's globals and locals through f_globals, f_locals and f_back. Both reach past any namespace you built.

open as a page

A clinical-lab result loader must run user-submitted transform scripts; how do you choose the isolation boundary and validate it against a 340-case regression pack?

level: principalimportance: should knowfreq 22%

basics

~20 s

Decide the boundary first: never in-process. Run each submitted script in a short-lived separate process inside a locked-down container (dropped privileges, no network, read-only or disposable filesystem, CPU/memory/time limits), escalating to a microVM for stronger isolation. The 340-case regression pack proves correctness, not containment — pair it with an adversarial escape suite.

open as a page

Is ast.literal_eval(repr(x)) a safe way to round-trip Python data?

level: middleimportance: nice to knowfreq 18%

basics

~20 s

No. repr is a debugging aid, not a serialization format. Only literal-shaped values survive: nan and inf come back as bare names, Decimal, datetime and frozenset repr as calls, and all of those raise on the way back in.

open as a page