skip to content

Why does a Python code-checking tool use ast.parse() instead of regular expressions over the source?

level: juniorimportance: should knowfreq 40%

answer

  1. Code as data, not as text
  2. The interpreter's own parser, stopped early
  3. Nodes are typed objects with fields
  4. ast.parse gives a Module; nothing runs
  5. Positions come from lineno and col_offset

basics

~20 s

ast.parse() turns source text into a tree that mirrors Python's own grammar, so a tool matches real structure - a call, an assignment, a function - instead of characters. Regexes cannot see nesting, strings or comments.

solid answer

~40 s

`ast.parse(source)` runs the interpreter's own parser and returns an `ast.Module` object whose nodes are typed Python objects: a call is an `ast.Call` with `func`, `args` and `keywords`; a literal is an `ast.Constant`. Nothing executes - parsing stops before bytecode. That structure is what a checker actually needs: it can ask "is this an `ast.Call` whose `func` is an `ast.Name` with `id == 'open'`, and which function is it inside?", a question text matching cannot answer because the same characters appear inside string literals and comments, and real calls wrap across lines. Each node also carries `lineno` and `col_offset`, so a finding points at a range in the file. The tree is syntactic only: it does not resolve imports or types, and it drops comments and formatting.

code

python · 5 lines
python
import ast

tree = ast.parse("total = rows + 1")
print(type(tree).__name__, type(tree.body[0]).__name__)
print(ast.dump(tree.body[0].value, indent=2))

go deeper

for a junior

Be ready to say in one sentence what ast.parse() returns and that nothing executes. Knowing that a call is an ast.Call node with a func field, and that ast.dump() shows you the shape, is enough at this level.

for a middle

Explain the traversal options - ast.walk, ast.iter_child_nodes, ast.NodeVisitor - and what each costs you in parent context. Be able to name what the tree throws away: comments, whitespace, quote style, redundant parentheses.

for a senior

Show where the syntactic tree stops: no import resolution, no types, per-file scope only. Explain how you would report a finding with lineno and col_offset, and how inline suppression markers must come from the text, not the tree.

for a principal

Own the version story. Decide which interpreter your analysis runs under, given that a tree can never contain syntax newer than the parser, and weigh a standard-library AST pass against a richer parsing layer when the tool must serve several language versions.

## What `ast.parse()` gives you `ast.parse(source)` hands a string of Python source to the interpreter's own parser and stops at the *abstract syntax tree* - the structured representation CPython builds before it compiles bytecode. You get back an `ast.Module` object. **Nothing in the source runs**; parsing is not execution. The tree is made of ordinary Python objects with typed fields. A module holds a `body` list of statements. An assignment holds `targets` and a `value`. A call node (`ast.Call`) holds `func`, `args` and `keywords`. Every literal - number, string, `True`, `None` - is an `ast.Constant` carrying its `value`. You can look at any of it with `ast.dump(node, indent=2)`, which prints the node type and its fields and is the fastest way to learn what shape a construct actually has. Because it is the interpreter's parser, the grammar is not an approximation. If `ast.parse()` accepts the text, the running interpreter would too; if it does not, you get the same `SyntaxError` an import would raise. ## Why text matching loses A regex sees characters. The tree sees constructs. Concretely, a pattern like `\bopen\(` matches: - the word inside a string literal or a comment, where no call happens at all; - a call spelled across a line break or wrapped in parentheses, which it may miss entirely; - an attribute call on some unrelated object, which it cannot tell apart from the builtin. The tree makes each of those a different question with a definite answer: is the node an `ast.Call`, is its `func` an `ast.Name` or an `ast.Attribute`, what is the `id`, which enclosing `ast.FunctionDef` contains it. Equally important, the tree is *normalised*: whitespace, redundant parentheses and quote style are gone, so `x=1` and `x = 1` parse to the same structure. A checker wants that; a formatter, as it happens, does not. ## Getting around the tree Three traversal styles cover almost everything: - `ast.walk(tree)` yields every node in the tree in no guaranteed depth-first order, with no parent context. It is the one-liner for "find all nodes of type X". - `ast.iter_child_nodes(node)` yields one level of children, for hand-written recursion. - `ast.NodeVisitor` dispatches to a `visit_<NodeType>` method per node type, which is what a real linter uses because rules become methods. ```python import ast tree = ast.parse("total = rows + 1") names = [n.id for n in ast.walk(tree) if isinstance(n, ast.Name)] ``` ## Positions, and what the tree cannot tell you Most nodes carry `lineno`, `col_offset`, `end_lineno` and `end_col_offset`, which is how a tool reports `file.py:12:5` and how an editor underlines exactly the offending expression. The limits matter as much as the power. The AST is **per file and purely syntactic**. It does not resolve imports, does not know types, and does not know whether the name `open` refers to the builtin or to something rebound three modules away - name resolution and typing are layers you build on top. It also discards comments and exact formatting, so a suppression comment such as a `# noqa`-style marker has to be recovered from the raw text or the token stream, never from the tree. Finally the grammar moves with the language, and a visitor written against an older one silently under-reports rather than failing. `ast.Match` nodes exist only from Python 3.10; PEP 695 type-parameter syntax and `ast.TypeAlias` arrived in 3.12; Python 3.14 added nodes for template strings (PEP 750). `ast.parse()` takes a `feature_version` argument to *restrict* itself to an older grammar, but it can never parse syntax newer than the interpreter running it - so a tool that must analyse code for several Python versions either runs under the newest one it supports or keeps a parser of its own. ## The practical shape of a checker Parse the file, walk or visit the tree, collect `(rule, lineno, col_offset, message)` tuples, print them. That is genuinely the whole architecture of a simple linter, and it is why `ast` is in the standard library: treating code as data is a supported, first-class thing to do in Python. ## Two details worth carrying `ast.parse()` takes a `mode` argument. The default `"exec"` parses a whole module and gives an `ast.Module`; `mode="eval"` parses a single expression and gives an `ast.Expression`, which is how you check that a configuration value is one expression rather than a block of statements. And it raises `SyntaxError` on anything the running interpreter cannot parse - including source written for a newer Python - so a tool that walks a repository must catch it per file and report the file rather than dying on the first one it does not understand.

  • Does ast.parse() execute anything in the source it is given?
    No. It parses and returns a tree, and stops well before bytecode. Even so, parsing is not free of failure modes: deeply nested or pathological source can exhaust the parser's limits, so a tool that parses arbitrary files should still handle `SyntaxError`, `RecursionError` and `MemoryError` rather than assuming a tree always comes back.
  • Why can't a tool find a suppression comment like a per-line ignore marker in the tree?
    Comments are not part of the abstract syntax tree at all - the parser discards them, along with blank lines, redundant parentheses and quote style. Tools that honour inline ignore markers read them from the source text or the token stream and correlate them with node `lineno` values.
  • What is the difference between ast.walk and ast.iter_child_nodes?
    `ast.walk(node)` yields the node and every descendant, flattened, in no guaranteed depth-first order and with no parent information. `ast.iter_child_nodes(node)` yields only the immediate children, which is what you use when you are recursing yourself and need to keep track of context such as the enclosing function.

A regex reads a sentence as a string of letters; the AST reads it as subject, verb and object. Only one of them can tell you the sentence is a question.

saying these in an interview costs you the question

  • Claiming ast.parse runs or imports the code
  • Believing the tree keeps comments and formatting
  • Thinking the AST knows types or resolves imports
  • Assuming ast.walk yields nodes in source order
  • Saying a regex is equivalent if it is careful enough

context