skip to content

Analyzers and AST Rewriting

Programs that read Go source as data: the analysis framework a vet pass plugs into, and go/ast with go/types for one-off rewrites and codemods.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

explore

questions

9

Why does go/parser.ParseFile need a token.FileSet, and what does it return?

level: juniorimportance: must knowfreq 40%

answer

  1. nodes store positions as plain integers
  2. many files, one numbering scheme
  3. who knows where line 12 starts?
  4. fset.Position decodes it to file:line:col

basics

~20 s

go/parser.ParseFile returns an *ast.File, the syntax tree of one source file, plus an error. Node positions are bare integers; the token.FileSet holds each file's name and line table, so those integers decode to file, line and column.

solid answer

~40 s

`parser.ParseFile(fset, filename, src, mode)` parses one file into an `*ast.File` — package clause, declarations, imports, and (with `parser.ParseComments`) comment groups. Every node exposes `Pos()` and `End()` as `token.Pos`, which is not a byte offset into that file but an index into one continuous address space owned by the `token.FileSet`. The FileSet records where each file starts and where its lines break, so `fset.Position(p)` gives you `token.Position{Filename, Offset, Line, Column}`. That is why the FileSet is a parameter rather than something the parser makes internally: you create one `token.NewFileSet()` for the whole run, parse every file into it, and keep it around for reporting positions and for printing the tree back out. A Pos is meaningless against any other FileSet.

code

go · 13 lines
go
fset := token.NewFileSet()
file, err := parser.ParseFile(fset, "handler.go", src, parser.ParseComments)
if err != nil {
	return err
}
for _, decl := range file.Decls {
	fn, ok := decl.(*ast.FuncDecl)
	if !ok {
		continue
	}
	p := fset.Position(fn.Pos()) // Filename, Offset, Line, Column
	fmt.Printf("%s:%d:%d %s\n", p.Filename, p.Line, p.Column, fn.Name.Name)
}

go deeper

for a junior

Recall the shape of the call: a FileSet you create, a filename, optional source bytes, and mode flags, returning an *ast.File and an error. Be able to say that fset.Position turns a node's position into file, line and column.

for a middle

Explain why a token.Pos is a single integer and how the FileSet's per-file base and line table decode it, and name the mode flags that change what ends up in the tree.

for a senior

Show the operational habits: one FileSet for the whole run, keeping the original bytes alongside it so Position.Offset can drive byte-level edits, and a deliberate decision about what to do with the partial tree a syntax error returns.

for a principal

Frame the choice of a syntax-only pass as a cost decision: it runs on code that does not compile and costs milliseconds per file, which is what makes a whole-repo pass viable at all, and be clear about the correctness it cannot give you.

## What ParseFile does `go/parser.ParseFile` turns the bytes of a single Go source file into a tree of `go/ast` nodes. Its signature is: ```go func ParseFile(fset *token.FileSet, filename string, src any, mode Mode) (*ast.File, error) ``` - **`fset`** — the `*token.FileSet` the new file is registered in. It must not be nil. - **`filename`** — recorded in the FileSet and used in error messages. If `src` is nil, the parser reads this path from disk; otherwise the name is purely a label. - **`src`** — nil, a `string`, a `[]byte`, or an `io.Reader` holding the source text. - **`mode`** — parser flags. `parser.ParseComments` keeps comments in the tree (without it they are thrown away). `parser.AllErrors` reports every syntax error instead of stopping after a handful. `parser.ImportsOnly` and `parser.PackageClauseOnly` stop parsing early when you only need the file's header. `parser.SkipObjectResolution` skips building the old identifier-resolution data, which is faster and is what you want when you will do your own name resolution or a type-checked pass. The result is an `*ast.File` with fields such as `Name` (the package identifier), `Decls` (top-level declarations), `Imports`, `Doc` (the package doc comment) and `Comments` (every comment group in the file). If the source could not be read at all you get a nil file; if it read but did not parse, you get a **partial tree plus an error**, and that error is a `scanner.ErrorList` you can iterate for individual positions and messages. Deciding whether to work with a partial tree or bail out is your call. One thing `ParseFile` deliberately does not do is type-check. It has no idea what any identifier refers to, what a package's exported names are, or whether the code compiles. Undefined names, wrong argument counts and type mismatches are not syntax errors and will not be reported. ## Why a FileSet exists Every AST node implements `Pos() token.Pos` and `End() token.Pos`. A `token.Pos` is a single integer, and that is a deliberate space optimisation — nodes are tiny and there are millions of them. The integer is **not** an offset within its own file. It is an index into one continuous address space that the FileSet hands out: the first file gets a base, the file occupies `size+1` values, the next file starts after it, and so on. Because the ranges do not overlap, a Pos identifies its file implicitly. The FileSet is the only thing that can decode that. It stores, per file, the name, the base, the size and the table of line-start offsets. So: ```go pos := fset.Position(node.Pos()) // pos.Filename, pos.Offset (byte index in that file), pos.Line, pos.Column ``` `fset.File(p)` gives you the `*token.File` a position belongs to, and `fset.PositionFor(p, adjusted)` lets you choose whether `//line` directives (used by code generators to point at their input) are applied. `token.NoPos` is the zero value and means "no position" — synthesised nodes you build yourself have it. The practical rules that follow: 1. **Create one FileSet per run** with `token.NewFileSet()` and parse every file into it. Do not make a fresh one per file or per print call. 2. **Never mix FileSets.** A Pos from one and a `Position` call on another yields the wrong file or nonsense; nothing panics, so the bug shows up as garbage line numbers. 3. **Keep the FileSet as long as you keep the tree.** Printing the tree back out with `go/format.Node(w, fset, file)` needs it, because the printer places comments and blank lines using the recorded positions. 4. **`Position.Offset` is your bridge to raw bytes.** If you also keep the original source, offsets let you splice text directly instead of re-printing the whole file — which is how you keep untouched lines byte-identical in a large mechanical change. ## A minimal pass ```go fset := token.NewFileSet() file, err := parser.ParseFile(fset, path, nil, parser.ParseComments) if err != nil { return err } ast.Inspect(file, func(n ast.Node) bool { if fn, ok := n.(*ast.FuncDecl); ok { fmt.Println(fset.Position(fn.Pos()), fn.Name.Name) } return true }) ``` That prints something of the form `path/to/file.go:12:1 Handler` for each function declaration — the FileSet is what makes the human-readable half of that line possible. ## Where this sits `go/parser` and `go/ast` are syntax only. `go/token` supplies positions. `go/format` prints a tree back as gofmt-formatted source. When you need to know what a name *means* — which package a selector resolves to, what type an expression has — you move up to a type-checked load, which reuses exactly the same `token.FileSet` and `*ast.File` values, so nothing you learn here is thrown away.

  • What happens if you decode a node's position using a different token.FileSet than the one it was parsed into?
    You get silent nonsense — the wrong filename and line, or a zero Position — because the integer is an index into that particular FileSet's address space. Nothing panics and nothing warns you, so mismatched FileSets show up as inexplicably wrong diagnostics. Create one FileSet per run and thread it everywhere you thread the tree.
  • Comments are missing from your parsed tree. What did you forget?
    The `parser.ParseComments` mode flag. Without it the parser drops comments entirely, since most consumers do not want them. With it, every comment group appears in `ast.File.Comments`, and the ones the parser could attach are also reachable through `Doc` and `Comment` fields on declarations, specs and struct fields.
  • ParseFile returns an error. Is the *ast.File it returns usable?
    Usually yes. If the source could not be read the file is nil, but for a syntax error the parser returns a partial tree along with a `scanner.ErrorList` holding each error's position and message. Adding `parser.AllErrors` makes it report all of them rather than stopping early. Whether a partial tree is safe to act on depends on your pass.

The FileSet is one long tape measure laid end to end across every file you parsed. A node records only the number it sits at; only the tape knows that number 40312 means line 12 of handler.go.

saying these in an interview costs you the question

  • Thinks token.Pos is a byte offset within its own file
  • Creates a fresh token.FileSet per file or per print call
  • Expects comments in the tree without parser.ParseComments
  • Believes ParseFile type-checks or resolves imports
  • Assumes a syntax error always means a nil *ast.File
open as a page

In golang.org/x/tools/go/analysis, what is an *analysis.Analyzer made of, and how does its Run function report a finding?

level: middleimportance: must knowfreq 40%

basics

~20 s

An analysis.Analyzer is a struct: Name, Doc, Requires and a Run function. Run is called once per package and calls pass.Reportf with a source position for each finding. A non-nil error from Run means the analyzer itself failed.

open as a page

In a custom go vet analyzer, what does pass.TypesInfo give you that the syntax tree alone cannot?

level: middleimportance: should knowfreq 32%

basics

~20 s

pass.TypesInfo is the package's *types.Info: it maps each identifier to the declaration it resolves to and each expression to its type. That is how a check recognises context.TODO whatever the import is aliased to, and ignores an unrelated TODO.

open as a page

How do ast.Inspect and ast.Walk differ, and what stops a traversal descending into children?

level: middleimportance: should knowfreq 32%

basics

~20 s

ast.Walk drives an ast.Visitor whose Visit returns a visitor for the children, or nil to skip them. ast.Inspect wraps that in a func(ast.Node) bool where false skips children. Both call you again with a nil node when a subtree ends.

open as a page

A custom vet analyzer every repo runs has reported nothing for a quarter. How do you prove it still fires?

level: seniorimportance: should knowfreq 22%

basics

~20 s

Treat silence as a failure until proven otherwise. Prove the matcher on an analysistest fixture with want comments, then work up the chain: analyzer linked into the shipped binary, packages actually analysed, files not hidden by build constraints, results not served from cache.

open as a page

How do you package custom analyzers into a binary that go vet -vettool can run, and what does that protocol constrain?

level: seniorimportance: should knowfreq 28%

basics

~20 s

Build a main package that calls singlechecker.Main for one analyzer or multichecker.Main for a suite; both speak the vet protocol. The go command then runs that binary once per package, giving it only that package plus its dependencies' export data.

open as a page

Your go/ast codemod rewrites 400 files, but the printed output drops some comments and attaches others to the wrong declaration. Why, and how do you fix it?

level: seniorimportance: should knowfreq 28%

basics

~20 s

Comments are not children of the nodes they document. They sit in a flat, position-sorted list on ast.File, and the printer places them by position. Editing the tree makes those positions stale. Use an ast.CommentMap, or splice the original bytes.

open as a page

Why can a go/ast-only rewrite misidentify a call, and what does a type-checked pass give you instead?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

A syntax tree records spelling, not meaning: two identical selector expressions can mean an imported package's function and a local variable's method. A type-checked load resolves each identifier to a types.Object, so a rewrite can test the real package and type.

open as a page

When should a home-grown go vet analyzer become a mandatory check for every team, and what may it forbid?

level: principalimportance: nice to knowfreq 18%

basics

~20 s

Only when the rule is decidable from one package's syntax and types, has a real defect behind it, has run non-blocking across every repository with its hits sampled, and ships with a fix and a self-serve suppression. It may forbid mistakes, never taste.

open as a page