skip to content

Compilation Pipeline

What sits between source text and something executable: a syntax tree you can rewrite, a code object holding constants and flags, and a cached compile result on disk.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

16

What is the difference between Python's eval() and exec()?

level: juniorimportance: must knowfreq 55%

answer

  1. One hands a value back, one does not
  2. The split follows Python's grammar
  3. Assignments and loops need the other builtin
  4. exec() returns None every single time
  5. compile() has matching 'eval' and 'exec' modes

basics

~10 s

eval() evaluates one expression and returns its value. exec() runs statements - assignments, loops, imports, definitions - and always returns None. Hand eval() a statement and it raises SyntaxError.

solid answer

~40 s

`eval()` takes a single Python **expression**, something that produces a value, and returns that value. `exec()` takes any block of **statements** and returns `None`; you observe its effect through the namespace it wrote into, not through a return value. Both accept a string, bytes, or a pre-compiled code object, and both take optional `globals` and `locals` mappings that decide which names the code can see and where new bindings land (since 3.13 those two may also be passed as keyword arguments). The split mirrors `compile()`'s mode argument: `'eval'` compiles one expression, `'exec'` compiles a sequence of statements. So `eval("cpm * 1000")` is fine, while `eval("total = 1")` raises `SyntaxError` because an assignment is a statement. If the string is really just data, `ast.literal_eval()` parses it without executing arbitrary code.

code

python · 6 lines
python
value = eval("2 + 3 * 4")
print(value)                 # 14

result = exec("total = 2 + 3")
print(result)                # None
print(total)                 # 5

go deeper

for a junior

Recall the one-line split: eval() takes a single expression and gives you its value, exec() takes statements and gives you nothing. Be ready to say which one runs x = 1 and which one runs x + 1.

for a middle

Explain the mechanics: both accept a string, bytes or a code object, both take globals and locals mappings, and compile()'s 'eval' and 'exec' modes mirror the two builtins. Show how you retrieve a value that exec() computed.

for a senior

Show judgement about when neither belongs in the code at all. Name ast.literal_eval() for data strings, note that pre-compiling avoids re-parsing on a hot path, and be clear that the namespace you pass is the only channel exec() has.

for a principal

Own the policy question: what in your codebase is allowed to execute strings at runtime, what the alternatives are (a parsed configuration format, a restricted expression grammar, a plugin entry point), and how you keep this out of code paths fed by outside input.

## The grammar line the two builtins sit on Python's grammar splits source into **expressions** and **statements**. An expression produces a value: `2 + 3`, `cpm * 1000`, a call, a comprehension, a conditional expression, even an assignment expression `(w := 5)`. A statement *does* something: `x = 5`, `for`, `if`, `import`, `def`, `class`, `return`, `with`. Every expression can stand alone as a statement (its value is then discarded), but no statement can appear where a value is required. `eval()` and `exec()` are the two builtins that run code arriving as text at runtime, and they sit on opposite sides of that line. ## eval() `eval(source, globals=None, locals=None)` parses `source` as **one expression**, evaluates it, and **returns the result**. The source may be a `str`, a `bytes` object, or a code object previously produced by `compile(..., mode='eval')`. Because the whole point is the returned value, `eval()` is the right tool when you want a number, a container, or an object back: ```python print(eval("3 * 4 + 1")) # 13 print(eval("max(bids)", {"bids": [3, 9, 4]})) # 9 ``` Give it anything the grammar calls a statement and the parse fails before evaluation ever starts: ```python eval("total = 1") # SyntaxError: invalid syntax eval("import math") # SyntaxError ``` One subtlety worth knowing: an *assignment expression* is still an expression, so `eval("(w := 5)")` succeeds and binds `w` into the globals mapping it was given. That is a real binding, not a copy. ## exec() `exec(source, globals=None, locals=None)` parses `source` as **a sequence of statements**, runs them, and **always returns `None`**. Anything you want back has to be read out of the namespace afterwards: ```python ns = {} exec("floor = 2\nbid = floor * 3", ns) print(ns["bid"]) # 6 ``` A very common beginner error is expecting `exec()` to behave like a REPL and hand back the value of the last line. It does not; the value of a bare expression statement is discarded, exactly as it is inside a normal function body. Because `exec()` parses statements, indentation matters. A string that carries leading indentation - typical when the source came out of a triple-quoted template inside an indented block - raises `IndentationError`. `textwrap.dedent()` is the usual fix. `eval()` is more forgiving only in that leading whitespace on its single line is stripped. ## What they share Both take the same two optional namespace arguments, and both accept a code object as well as text. The two builtins are best understood as *drivers* over a code object whose behaviour was fixed at compile time: `compile(src, filename, 'eval')` yields a code object that produces a value, `compile(src, filename, 'exec')` yields one that does not. The mode, not the builtin, is what decides. Consequently `eval()` will happily run an `'exec'`-mode code object - it just returns `None`, because there is no result to hand back. Both also default their namespaces to the caller's scope when you pass nothing, which is why `exec("total = 5")` typed at module level really does create a module-level name, while the same call inside a function does not create a function local. ## Choosing, and choosing neither In day-to-day code the honest answer is usually *neither*. Three questions sort it out: 1. **Is the string plain data** - a number, a list, a dict of literals? Then `ast.literal_eval()` is the right call. It parses the text and builds only literals and containers; `ast.literal_eval("len('abc')")` raises `ValueError` rather than calling anything. Note it is deliberately narrow: even `ast.literal_eval("1 + 2")` is rejected, because arithmetic is not a literal. 2. **Do you need a value back from a small formula?** That is `eval()`'s job, and pre-compiling the formula with `compile(..., 'eval')` avoids re-parsing it on every call. 3. **Do you need to define or bind things?** That is `exec()`, and you should pass an explicit dict so the results land somewhere you control. And if the string came from outside your process, the interesting question stops being *eval or exec* and becomes whether you should be executing it at all - both run arbitrary code with the full power of the interpreter.

  • If exec() always returns None, how do you get a computed value back out of the code it ran?
    Pass your own dictionary as the globals mapping and read the key out afterwards: `ns = {}; exec("bid = 6", ns); ns["bid"]`. There is no other channel - the return value is always `None`, and the value of a bare expression statement inside the string is discarded. If what you want is a single value, use `eval()` instead so the result comes back directly.
  • Can eval() run a code object that was compiled in 'exec' mode?
    Yes. The mode is baked into the code object at `compile()` time, and both builtins will execute a code object of either mode. An `'exec'`-mode object run through `eval()` executes its statements normally and returns `None`, because there is no expression result to produce. The practical consequence is that the mode argument, not the choice of builtin, decides whether you get a value.
  • When should you reach for ast.literal_eval() instead of eval()?
    Whenever the string is data rather than code - a number, a tuple, a list or dict of literals, a boolean, `None`. `ast.literal_eval()` parses the text and constructs only literals and containers, raising `ValueError` on anything else, so `ast.literal_eval("len('abc')")` fails rather than calling a function. It is deliberately narrow: even `"1 + 2"` is rejected, since arithmetic is not a literal.

saying these in an interview costs you the question

  • Says exec() returns the value of its last statement
  • Thinks eval() can run an assignment, loop or import
  • Treats exec() as just another name for eval()
  • Uses eval() to parse a config string that is plain data
  • Assumes a string passed to eval() is compiled once and cached

context

open as a page

What is Python's __pycache__ directory and when does CPython reuse a .pyc file?

level: juniorimportance: must knowfreq 55%

basics

~20 s

CPython caches each module's compiled bytecode in a pycache directory next to the source. It reuses that .pyc only when the magic number matches the interpreter and the recorded source timestamp and size still match.

open as a page

What is the difference between ast.NodeVisitor and ast.NodeTransformer in Python?

level: middleimportance: must knowfreq 50%

basics

~10 s

ast.NodeVisitor reads a tree: its visit_<NodeType> methods return whatever they like and the tree is untouched. ast.NodeTransformer rewrites it: whatever a visit method returns replaces the node, and returning None deletes it.

open as a page

Why do two functions from the same def share a code object but not defaults?

level: middleimportance: must knowfreq 40%

basics

~20 s

The code object is compiled once and is immutable, so it can be shared. Default values are ordinary expressions evaluated each time the def statement executes, so they belong to the function object created then, in defaults and kwdefaults.

open as a page

Why does a Python code-checking tool use ast.parse() instead of regular expressions over the source?

level: juniorimportance: should knowfreq 40%

basics

~20 s

ast.parse() turns source text into a tree that mirrors Python's own grammar, so a tool matches real structure - a call, an assignment, a function - instead of characters. Regexes cannot see nesting, strings or comments.

open as a page

What does a Python function's __code__ attribute give you?

level: juniorimportance: should knowfreq 28%

basics

~20 s

A function's code is its code object: the compiled bytecode for the body plus the static facts the compiler worked out — name, filename, argument count, local names, constants and flags. It is immutable, and every function built from the same def shares it.

open as a page

Why does compile() reject a tree edited with ast.NodeTransformer until you fix its locations?

level: middleimportance: should knowfreq 32%

basics

~10 s

Nodes you construct yourself have no lineno or col_offset, and compile() requires them on every node, raising TypeError: required field "lineno" missing. ast.fix_missing_locations() copies positions down from each node's parent.

open as a page

On a Python code object, how do co_varnames and co_names differ?

level: middleimportance: should knowfreq 30%

basics

~20 s

co_varnames lists the function's local variable names, parameters first, which the compiler turned into fixed numbered slots. co_names lists names that must be resolved at run time by name — globals, builtins, attributes and imported modules. Constants live in a third tuple, co_consts.

open as a page

Why does an assignment inside exec() not change a function's local variable?

level: middleimportance: should knowfreq 35%

basics

~20 s

A function's locals are compiled into fixed numbered slots, and exec() with no explicit namespace writes into a dictionary snapshot of those slots instead. The assignment lands in the throwaway mapping, so the real variable never changes.

open as a page

How do the globals and locals mappings passed to exec() decide what the code can see?

level: middleimportance: should knowfreq 30%

basics

~20 s

They are the namespace the code runs in. Pass nothing and it uses the caller's scope; pass one dict and it serves as both globals and locals; pass two and lookups fall back locals to globals to builtins, while new names bind into the locals mapping.

open as a page

What do PEP 552 hash-based .pyc files fix that timestamp .pyc files cannot?

level: middleimportance: should knowfreq 30%

basics

~20 s

A hash-based .pyc records a hash of the source bytes instead of its modification time and size, so caches survive builds that reset timestamps and stay reproducible. Checked mode re-hashes on import; unchecked trusts it.

open as a page

Why is an ast.parse() and ast.unparse() round trip a poor basis for a codemod on a real repository?

level: seniorimportance: should knowfreq 28%

basics

~10 s

The tree is an abstraction: comments, blank lines, quote style and line breaks are never in it, so ast.unparse() regenerates the whole file in its own style. Every touched file becomes a total-rewrite diff.

open as a page

Which expressions does CPython's compiler fold into co_consts before the code runs?

level: seniorimportance: should knowfreq 22%

basics

~20 s

Only expressions built entirely from literals of immutable types, and only up to a size cap: 60 * 60 * 24 becomes the constant 86400. Anything touching a name, an attribute or a call is left to run time, so hoisting those yourself is what actually pays.

open as a page

Stale .pyc bytecode caches: how would you prove they are why a redeployed catalogue importer still fails a 340-case encoding regression?

level: seniorimportance: should knowfreq 35%

basics

~20 s

First confirm which files were actually imported, then decode each cached header: magic number, stored mtime and size, compared against the source on disk. A deploy that preserves timestamps and file size makes CPython trust yesterday's bytecode.

open as a page

What are frozen modules in CPython, and why does editing their source change nothing?

level: middleimportance: nice to knowfreq 12%

basics

~10 s

A frozen module has its compiled bytecode embedded in the python executable, so importing it reads no .py and no .pyc and validates nothing. Editing the matching source changes nothing until -X frozen_modules=off.

open as a page

Why compile a string once with compile() rather than calling eval() on it repeatedly?

level: seniorimportance: nice to knowfreq 20%

basics

~20 s

eval() re-parses and re-compiles the source text on every call. compile(src, filename, 'eval') does that work once and returns a code object you evaluate per call, which for a small expression is roughly an order of magnitude cheaper.

open as a page