Why compile a string once with compile() rather than calling eval() on it repeatedly?
answer
- Running a string has two separate phases
- There is no cache for source strings
- compile() pays the front-end cost once
- The mode is baked into the code object
- filename shows up in the traceback
basics
~20 seval() re-parses and re-compiles the source text on every call. compile(src, filename, 'eval') does that work once and returns a code object you evaluate per call, which for a small expression is roughly an order of magnitude cheaper.
solid answer
~50 sWhen you hand `eval()` a string it must tokenise, parse and compile that text before executing anything, and it does so **every single call** - there is no cache. `compile(src, filename, mode)` performs the front end once and hands back a code object that `eval()` or `exec()` will run directly. On a hot path - say a bidding service replaying six hours of traffic overnight and evaluating one operator-authored pricing formula per bid - that is the difference between paying for a parse a few million times and paying for it once. The `mode` argument picks what gets compiled: `'eval'` for a single expression whose value is returned, `'exec'` for a block of statements, `'single'` for one interactive statement whose expression value is passed to `sys.displayhook` instead of being discarded, which is what the REPL uses to echo results. The `filename` argument is not cosmetic either - it names the frame in any traceback the code raises.
code
python · 10 linesimport timeit
src = "base * ctr - floor"
ns = {"base": 3, "ctr": 4, "floor": 1}
from_text = timeit.timeit(lambda: eval(src, dict(ns)), number=20000)
code = compile(src, "<pricing:camp-7>", "eval")
from_code = timeit.timeit(lambda: eval(code, dict(ns)), number=20000)
print(f"string {from_text:.3f}s code {from_code:.3f}s")go deeper
Recall that compile() turns source text into a code object, and that eval() and exec() can run that object instead of the original string. The mode argument is 'eval', 'exec' or 'single'.
Explain the two phases - parse and compile, then execute - and why repeating the first one per call is waste. Be able to say what each of the three modes produces and which one the REPL uses.
Show the operating judgement: compile at load time so syntax errors surface early and are attributed by filename, keep a fresh namespace per evaluation, and know that the code object is immutable so it explains no result drift.
Own whether runtime-compiled expressions belong in the product surface at all: who authors them, how they are validated and versioned, when a cache must be invalidated, and what a declarative alternative would cost in flexibility.
## What eval() actually does with a string Running a string is two phases: a **front end** that turns text into a code object (tokenise, parse to a syntax tree, compile to bytecode) and a **back end** that executes that code object in a namespace. `eval("base * ctr - floor", ns)` performs both, and it performs the first one from scratch on every call. Nothing memoises source strings. For a small expression the front end dominates. In a quick measurement on CPython 3.14, evaluating a three-name arithmetic expression from a string is roughly an order of magnitude slower than evaluating the same expression pre-compiled to a code object - the actual arithmetic is a handful of bytecodes, and everything else is parsing. ## Splitting the phases with compile() `compile(source, filename, mode)` gives you the code object directly: ```python formula = compile("base * ctr - floor", "<pricing:camp-7>", "eval") for bid in bids: price = eval(formula, {"base": bid.base, "ctr": bid.ctr, "floor": bid.floor}) ``` Consider a bidding service that replays six hours of the previous night's auction traffic to re-score campaigns. Each campaign carries a pricing formula authored by an operator, and each replayed bid evaluates it once. Compiling once per campaign at load time and evaluating the code object per bid removes the parse from the inner loop entirely, and it moves every syntax error in an operator's formula to load time, where it can be reported against a named campaign instead of blowing up in the middle of a long run. The `filename` argument is worth using properly. It has nothing to do with the filesystem; it is the string that appears in tracebacks and in frames, so `"<pricing:camp-7>"` tells you immediately which formula raised `ZeroDivisionError` at hour four. ## The three modes `mode` decides which grammar rule the source is parsed against, and it is baked into the resulting code object. - **`'eval'`** - the source must be a single expression. The code object produces its value, so `eval()` on it returns something useful. - **`'exec'`** - the source is a sequence of statements, like a module body. The code object returns `None`. - **`'single'`** - the source is one interactive statement. This is the mode the interactive interpreter uses, and its distinguishing behaviour is that the value of an expression statement is handed to `sys.displayhook` rather than discarded. That is what makes typing `2 + 2` at a prompt print `4`, while the identical line inside a function prints nothing. Because the mode is a property of the code object, the choice of builtin afterwards does not change it. `eval()` will run an `'exec'`-mode object and return `None`; `exec()` will run an `'eval'`-mode object and throw its value away. The mode, not the builtin, decides whether a value exists. ## What caching a code object does and does not buy you A code object is **immutable and stateless**. It holds bytecode, constants and names; it holds no variable values. That has two consequences worth stating in an interview. First, it is safe to share: one compiled formula can be evaluated concurrently from several threads, because all the mutable state lives in the namespace mapping you pass per call. Give each evaluation its own dict and there is nothing to race on. Second, caching the code object never changes results. If a replay produces different numbers than the original live run, the compiled formula is not the explanation - a formula that reads wall-clock time will see the replay's clock rather than the auction's, and that clock-skew artefact belongs to the namespace, not the compilation. The fix is to pass the original bid's timestamp in through the mapping instead of letting the expression reach for the current time. ## When not to bother If the expression is evaluated a handful of times, or if the work inside it dwarfs the parse, pre-compiling buys nothing and costs you a cache to manage - including invalidating it when an operator edits a formula. The pattern earns its keep exactly when a fixed set of strings is evaluated many times against changing data.
- What does compile()'s 'single' mode do that 'exec' mode does not?In `'single'` mode the value of an expression statement is passed to `sys.displayhook` instead of being discarded, so `compile("40 + 2", "<f>", "single")` prints `42` when executed. That is exactly how the interactive interpreter echoes results. `'exec'` mode compiles the same line to code that throws the value away, matching the behaviour of a statement inside a normal module or function body.
- A nightly replay produces different prices than the live run did for the same bids. Could caching the compiled code object explain that?No. A code object is immutable and holds no values - every input arrives through the namespace mapping you pass per evaluation. A drift like that points at the namespace instead: most often a formula that reads the current wall-clock time, so the replay sees the replay's clock rather than the auction's. Pass the original timestamp in explicitly and the expression becomes reproducible.
- Is a compiled code object safe to evaluate from several threads at once?The code object itself is, since it is immutable and carries no per-evaluation state. What is not safe is sharing one namespace mapping across concurrent evaluations - the executed code can bind names into it, so two threads would be writing the same dict. Build a fresh mapping per evaluation; it is cheap next to what you saved by not re-parsing.
saying these in an interview costs you the question
- Believes eval() caches compiled strings automatically
- Thinks compile() returns source text or a string of bytecode
- Says 'single' mode is just 'exec' mode for one line
- Expects eval() to reject an 'exec'-mode code object
- Assumes the parse cost is negligible on any hot path
- Shares one namespace dict across concurrent evaluations