skip to content

Why can a go/ast-only rewrite misidentify a call, and what does a type-checked pass give you instead?

level: seniorimportance: nice to knowfreq 22%

answer

  1. the tree records spelling, not meaning
  2. shadowing, aliases, dot imports
  3. identical node shape, different target
  4. an identifier resolved to the thing it names
  5. correctness bought with build time

basics

~20 s

A syntax tree records spelling, not meaning: two identical selector expressions can mean an imported package's function and a local variable's method. A type-checked load resolves each identifier to a types.Object, so a rewrite can test the real package and type.

solid answer

~40 s

`go/parser` gives you spelling only. A call written `log.Print(x)` parses to an `*ast.CallExpr` whose `Fun` is a `*ast.SelectorExpr` with `X` being the identifier `log` — and nothing in the tree says whether that `log` is the imported standard-library package, a parameter shadowing it, or an aliased import of something else. A rewrite matching on names alone edits the wrong call sites and misses the right ones. A typed pass loads packages with `golang.org/x/tools/go/packages.Load`, asking for syntax and type information, then consults `pkg.TypesInfo`: `Uses[ident]` yields the `types.Object` the identifier really refers to, so you can check `obj.Pkg().Path()` before touching anything. It costs far more — the code must compile, and checking dependencies dwarfs parsing — so use a syntax pass when a name match is provably sufficient, a typed pass when correctness depends on identity.

code

go · 13 lines
go
import "log"

type Logger struct{}

func (l *Logger) Print(v ...any) {}

func handle(log *Logger) {
	log.Print("the parameter's method")
}

func run() {
	log.Print("the imported package's function")
}

go deeper

for a junior

Recall that a parsed tree knows only the text: an identifier is a name, not a link to a declaration, and nothing in go/ast tells you which package a selector's qualifier came from.

for a middle

Name the concrete ways a name match goes wrong — import aliases, dot imports, shadowing by locals, methods promoted from embedded fields — and explain that type checking resolves each identifier to the entity it denotes.

for a senior

Make the cost call explicitly: a typed load needs the code to compile, sees one build configuration, and costs orders of magnitude more than parsing, so justify which passes need it and consider finding candidates syntactically before confirming them with types.

for a principal

Own the risk framing for a mechanical change: what a wrong edit costs, what evidence makes the diff trustworthy, and whether the team funds the slower correct tool or accepts a fast approximation with a review protocol around it.

## The limit of a syntax tree `go/parser` answers one question: what does this text say? It does not answer what any of it *means*. That distinction becomes concrete the moment a rewrite tries to find "every call to a particular function". Consider a file that both imports `log` and has a parameter named `log`. Both `log.Print("a")` call sites parse to the same shape — `*ast.CallExpr{Fun: *ast.SelectorExpr{X: *ast.Ident{Name: "log"}, Sel: *ast.Ident{Name: "Print"}}}` — and the tree carries no field that distinguishes them. The same ambiguity arises from: - **Import aliases.** `import l "log"` makes the qualifier `l`; `import "example.com/x/log"` makes it `log` while meaning something else. - **Dot imports.** `import . "strings"` puts `Contains` into file scope with no qualifier at all. - **Shadowing.** A local variable, parameter, field or method name can take over any identifier in an inner scope. - **Embedding and interfaces.** A method call may be promoted from an embedded field, or dispatched through an interface, so the receiver's spelled type is not the type that defines the method. - **Method values and function variables.** `f := pkg.Do; f()` calls the target with no selector in sight. A name-matching rewrite is therefore both unsound (edits things it should not) and incomplete (misses things it should). Whether that matters is a judgment about your repo: for a well-known qualifier that no one shadows, a syntax pass plus a review of the diff is often fine and takes seconds. For a rename that must be exact across hundreds of packages, it is not. ## What a typed pass adds Type checking resolves every identifier to a `types.Object` — the declared entity it refers to, carrying a name, a type and the package that declares it. The loading entry point is `golang.org/x/tools/go/packages`: ``` cfg := &packages.Config{Mode: packages.NeedSyntax | packages.NeedTypes | packages.NeedTypesInfo | packages.NeedName} pkgs, err := packages.Load(cfg, "./...") ``` The `Mode` bits matter more than they look: ask for too little and the fields you want are simply nil, with no error. From each `*packages.Package` you get `Fset`, `Syntax` (the `[]*ast.File`, exactly the trees `go/parser` would have produced), `Types` (the `*types.Package`) and `TypesInfo` (`*types.Info`). The maps inside `types.Info` are the payoff: - **`Uses[*ast.Ident] types.Object`** — what this reference refers to. - **`Defs[*ast.Ident] types.Object`** — what this declaring identifier declares. - **`Types[ast.Expr] types.TypeAndValue`** — the type (and constant value, if any) of an expression. - **`Selections`** — how a selector resolved, including promotion through embedded fields. So the reliable predicate replaces a string compare on the qualifier with an identity check on the object: take the selector's `Sel` identifier, look it up in `Uses`, and test `obj.Pkg()` and `obj.Name()` — and, for a method, the receiver type. That is true regardless of aliasing, shadowing or dot imports. ## What it costs - **Time and memory.** Loading with types compiles the dependency graph's type information. A whole-repo typed load is seconds to minutes and hundreds of megabytes where a parse-only pass is milliseconds per file. - **The code must build.** Type checking a package that does not compile yields errors and partial information. `pkg.Errors` must be checked — an empty `TypesInfo` on a broken package will otherwise look like "no matches found". A syntax pass, by contrast, happily runs over code mid-refactor, which is sometimes exactly why you want one. - **Build configuration is real.** Files behind build constraints, or written for another `GOOS`, are excluded from the load, so a typed pass sees one configuration at a time. Anything you must rewrite everywhere has to be run per configuration or handled syntactically. - **More machinery to get right.** You now depend on the type checker's model — assignability, instantiation of generic code, interface satisfaction — and mistakes there are subtler than a bad string match. ## Choosing, in practice A decision that holds up in review: 1. **Is the identifier realistically shadowed or aliased anywhere in this repo?** You can answer that cheaply with a search before committing to either approach. 2. **What is the blast radius of a wrong edit?** A cosmetic change that a reviewer will catch tolerates a syntax pass; a semantic change across hundreds of files does not. 3. **Does the code compile right now?** If half the repo is mid-migration, a typed load may not even run. 4. **How often will this run?** A one-shot migration can afford minutes. Something in a pre-commit path cannot. A useful middle path is to run the cheap syntax pass to find *candidates* and a typed load only over the packages that had candidates, which keeps the expensive step proportional to the number of hits rather than to the size of the repo. ## Sharing the same tree The two worlds interoperate cleanly, which is why this is a cost decision rather than an architectural one. A typed load hands back the same `*ast.File` values and the same `*token.FileSet` a plain parse would have produced, so a traversal you wrote against `ast.Inspect` keeps working unchanged — it just gains the ability to ask what each identifier means before deciding to rewrite it.

  • Which map tells you what an identifier in a loaded package actually refers to?
    `types.Info.Uses`, which maps a referencing `*ast.Ident` to the `types.Object` it denotes; `Defs` does the same for the identifier at a declaration site. From the object you get its name, its type and `Pkg()`, whose path is the reliable thing to compare against — not the qualifier as written in the source.
  • You load packages and TypesInfo comes back nil. What went wrong?
    The `packages.Config` Mode did not request it. The loader populates only the fields your Mode bits ask for, and asking for too little is silent — no error, just nil fields. Request syntax, types and type info explicitly. It is also worth checking each package's Errors, since a package that fails to type-check yields incomplete information that looks like an empty result.
  • Half the repo does not compile during a migration. Which kind of pass can still run?
    The syntax-only one. `go/parser` needs no dependencies, no build configuration and no successful type check, and it even returns a partial tree for a file with syntax errors. A typed load needs the package and its imports to check cleanly, which is exactly what a half-finished migration cannot promise.
  • How do you keep a typed pass affordable across a large repository?
    Two-stage it. Run the cheap syntax pass over everything to find candidate files or packages, then load only those packages with type information to confirm each candidate before editing. The expensive step then scales with the number of hits instead of the size of the repo, and the cheap step is trivially parallel.

saying these in an interview costs you the question

  • Matches call sites by the qualifier's spelling alone
  • Assumes an import name always equals the package name
  • Ignores shadowing, dot imports and promoted methods
  • Forgets to request type info in the loader's Mode
  • Never checks the loaded packages' errors before reporting no matches