skip to content

Why is an ast.parse() and ast.unparse() round trip a poor basis for a codemod on a real repository?

level: seniorimportance: should knowfreq 28%

answer

  1. The A in AST is doing real work
  2. What the compiler ignores, the tree omits
  3. Regenerated source is canonical, not original
  4. A codemod's product is a reviewable diff
  5. Is the artefact a file or behaviour?

basics

~10 s

The tree is an abstraction: comments, blank lines, quote style and line breaks are never in it, so ast.unparse() regenerates the whole file in its own style. Every touched file becomes a total-rewrite diff.

solid answer

~40 s

`ast.parse()` deliberately discards everything that does not change meaning - comments, blank lines, indentation width, quote style, redundant parentheses, trailing commas, how a long expression was wrapped. `ast.unparse()` then prints the tree in its own canonical form, so a one-line change comes back as a rewritten file with all comments gone. That is unusable as a source-editing codemod: the diff is unreviewable and the loss is silent. Two viable routes exist. If you are rewriting **source**, use a concrete-syntax-tree library that keeps comments and formatting as part of the tree, or generate a text patch anchored to the `lineno` and `col_offset` values from the AST. If you are rewriting **behaviour at import time**, never touching the file, the AST is exactly right: transform, fix locations, `compile()`, execute the code object.

code

python · 9 lines
python
import ast

SOURCE = '''# keep this note
label = "ok"
data = {
    'a': 1,
}
'''
print(ast.unparse(ast.parse(SOURCE)))

go deeper

for a junior

Remember the headline: the tree has no comments and no formatting, so regenerating source from it rewrites the whole file. That alone explains why an AST round trip is not a safe way to edit code.

for a middle

List what is dropped - comments, blank lines, quote style, parentheses, line wrapping - and explain that ast.unparse() emits canonical text rather than the original. Know that a docstring survives because it is a real string node.

for a senior

Show the alternatives and pick between them: AST-located text splicing for narrow edits, a concrete-syntax-tree library when the change is structural, and tree-plus-compile when the target is behaviour rather than files. Name the collateral damage from lost suppression directives.

for a principal

Own the rollout: how a mechanical change lands across many repositories in reviewable batches, what it does to open branches and blame history, and whether a formatting-lossy pipeline is ever acceptable in your codebase.

## Abstract means lossy, on purpose The word *abstract* in abstract syntax tree is the whole answer. The tree records what the code **means** to the compiler and drops everything that does not change that meaning: - comments and docstring formatting (a docstring survives as a string node; a `#` comment does not survive at all); - blank lines, indentation width, tabs versus spaces; - quote style - `"ok"` and `'ok'` are the same `ast.Constant`; - redundant parentheses, trailing commas, and how a long call or dict was wrapped across lines. `ast.unparse(tree)`, added in Python 3.9, regenerates source *from the tree*, so it can only emit its own canonical formatting. Run a file through `ast.parse()` and back and you get a semantically equivalent file that no reviewer will recognise: ```pycon >>> print(ast.unparse(ast.parse('# note\nlabel = "ok"\ndata = {\n "a": 1,\n}\n'))) label = 'ok' data = {'a': 1} ``` The comment is gone. The quotes flipped. The multi-line dict collapsed. Nothing warned you. ## Why that disqualifies it as a source codemod A codemod's product is a **diff that humans approve**. If every touched file is rewritten end to end, the review cannot separate the intended change from the incidental churn, and blame history is destroyed for every line. Worse, the losses are semantic in practice even though they are not semantic to the compiler: type-checker directives, linter suppression markers, licence headers and the explanatory comment above the tricky branch are all comments, and all of them vanish. On a large repository this converts a mechanical change into a merge-conflict generator against every open branch. ## What to use instead, and when the AST is still right There are three honest options. **Rewrite the text, use the AST only to locate.** Parse, find the nodes you care about, read their `lineno`, `col_offset`, `end_lineno` and `end_col_offset`, and splice new text into the original source at those offsets. Everything you do not touch is byte-identical. This is the cheapest correct approach for narrow changes - renaming a call, adding an argument - and it is what many small codemods actually do. **Use a concrete syntax tree.** A CST keeps comments and formatting as nodes in the tree, so a rewrite round-trips byte-for-byte outside the edited region. That is a third-party dependency rather than a standard-library one, and it is the right call once the change is structural enough that offset splicing becomes fragile. **Do not rewrite source at all.** If the goal is changed *behaviour* rather than changed *files* - injecting instrumentation, enforcing a check, specialising code at load time - transform the tree and compile it, and leave the file alone. Take a clinical-lab result loader whose resident memory grows without bound: wrapping every loader function with allocation tracking through an `ast.NodeTransformer`, fixing locations and compiling at import time gives you the measurement without a single line of reviewable churn, and you delete the pass when the leak is found. Nothing about formatting matters, because no file is ever written. ## Judgement, not dogma The question to answer before choosing is: *is the artefact a file or a behaviour?* If a human will read the result, formatting is part of the product and the AST alone is the wrong representation. If only the interpreter will read it, the AST is the correct and standard-library-only tool. One more trap for the file case: `ast.unparse()` guarantees a tree that means the same thing, not text that looks the same, and it is not a formatter - the output will not match whatever style your repository enforces, so an unparse-based pipeline needs a formatting pass afterwards and still cannot bring the comments back. And a mechanical rewrite across thousands of files needs a rollback story anyway. Land it behind review in batches rather than as one commit that a three-week release train has to either take whole or revert whole. ## How you verify a mechanical rewrite Whichever route you take, the check is the same and it is cheap: for every file the pass touches, parse the before and after trees and compare `ast.dump()` of each, ignoring positions. That proves the rewrite changed exactly the structure you intended and nothing else, which no amount of reading the diff will prove across thousands of files. Then read the diffs anyway, in batches small enough that a reviewer can actually approve them.

  • Do docstrings survive an ast.parse and ast.unparse round trip?
    Yes, because a docstring is a string expression statement in the tree, not a comment - `ast.get_docstring()` reads it back. Its internal line breaks are preserved as part of the string value, but the quote style and any surrounding blank lines are re-emitted canonically. Ordinary `#` comments, including type-checker and linter directives, are gone entirely.
  • How would you write a narrow codemod that keeps the rest of the file byte-identical?
    Parse to locate, then edit text. Walk the tree for the nodes you want, take their `lineno`, `col_offset`, `end_lineno` and `end_col_offset`, and splice replacement text into the original source at exactly those offsets - applying edits back-to-front so earlier offsets stay valid. Untouched bytes are untouched, so the diff shows only the change.
  • When is compiling a rewritten tree the right answer rather than editing source?
    When the deliverable is behaviour, not a file: import-time instrumentation, a load-time check, or specialising code for a run. You transform the tree, call `ast.fix_missing_locations()`, `compile()` it and execute the code object. No formatting is lost because nothing is written back, and the pass is removed by deleting it rather than by reverting a repository-wide diff.

saying these in an interview costs you the question

  • Believing ast.unparse reproduces the original file
  • Assuming comments live somewhere in the tree
  • Treating an unparse round trip as a formatter
  • Ignoring that lost directives change tooling behaviour
  • Landing a repository-wide rewrite as one unreviewable commit

context