Why does go/parser.ParseFile need a token.FileSet, and what does it return?
answer
- nodes store positions as plain integers
- many files, one numbering scheme
- who knows where line 12 starts?
- fset.Position decodes it to file:line:col
basics
~20 sgo/parser.ParseFile returns an *ast.File, the syntax tree of one source file, plus an error. Node positions are bare integers; the token.FileSet holds each file's name and line table, so those integers decode to file, line and column.
solid answer
~40 s`parser.ParseFile(fset, filename, src, mode)` parses one file into an `*ast.File` — package clause, declarations, imports, and (with `parser.ParseComments`) comment groups. Every node exposes `Pos()` and `End()` as `token.Pos`, which is not a byte offset into that file but an index into one continuous address space owned by the `token.FileSet`. The FileSet records where each file starts and where its lines break, so `fset.Position(p)` gives you `token.Position{Filename, Offset, Line, Column}`. That is why the FileSet is a parameter rather than something the parser makes internally: you create one `token.NewFileSet()` for the whole run, parse every file into it, and keep it around for reporting positions and for printing the tree back out. A Pos is meaningless against any other FileSet.
code
go · 13 linesfset := token.NewFileSet()
file, err := parser.ParseFile(fset, "handler.go", src, parser.ParseComments)
if err != nil {
return err
}
for _, decl := range file.Decls {
fn, ok := decl.(*ast.FuncDecl)
if !ok {
continue
}
p := fset.Position(fn.Pos()) // Filename, Offset, Line, Column
fmt.Printf("%s:%d:%d %s\n", p.Filename, p.Line, p.Column, fn.Name.Name)
}go deeper
Recall the shape of the call: a FileSet you create, a filename, optional source bytes, and mode flags, returning an *ast.File and an error. Be able to say that fset.Position turns a node's position into file, line and column.
Explain why a token.Pos is a single integer and how the FileSet's per-file base and line table decode it, and name the mode flags that change what ends up in the tree.
Show the operational habits: one FileSet for the whole run, keeping the original bytes alongside it so Position.Offset can drive byte-level edits, and a deliberate decision about what to do with the partial tree a syntax error returns.
Frame the choice of a syntax-only pass as a cost decision: it runs on code that does not compile and costs milliseconds per file, which is what makes a whole-repo pass viable at all, and be clear about the correctness it cannot give you.
## What ParseFile does `go/parser.ParseFile` turns the bytes of a single Go source file into a tree of `go/ast` nodes. Its signature is: ```go func ParseFile(fset *token.FileSet, filename string, src any, mode Mode) (*ast.File, error) ``` - **`fset`** — the `*token.FileSet` the new file is registered in. It must not be nil. - **`filename`** — recorded in the FileSet and used in error messages. If `src` is nil, the parser reads this path from disk; otherwise the name is purely a label. - **`src`** — nil, a `string`, a `[]byte`, or an `io.Reader` holding the source text. - **`mode`** — parser flags. `parser.ParseComments` keeps comments in the tree (without it they are thrown away). `parser.AllErrors` reports every syntax error instead of stopping after a handful. `parser.ImportsOnly` and `parser.PackageClauseOnly` stop parsing early when you only need the file's header. `parser.SkipObjectResolution` skips building the old identifier-resolution data, which is faster and is what you want when you will do your own name resolution or a type-checked pass. The result is an `*ast.File` with fields such as `Name` (the package identifier), `Decls` (top-level declarations), `Imports`, `Doc` (the package doc comment) and `Comments` (every comment group in the file). If the source could not be read at all you get a nil file; if it read but did not parse, you get a **partial tree plus an error**, and that error is a `scanner.ErrorList` you can iterate for individual positions and messages. Deciding whether to work with a partial tree or bail out is your call. One thing `ParseFile` deliberately does not do is type-check. It has no idea what any identifier refers to, what a package's exported names are, or whether the code compiles. Undefined names, wrong argument counts and type mismatches are not syntax errors and will not be reported. ## Why a FileSet exists Every AST node implements `Pos() token.Pos` and `End() token.Pos`. A `token.Pos` is a single integer, and that is a deliberate space optimisation — nodes are tiny and there are millions of them. The integer is **not** an offset within its own file. It is an index into one continuous address space that the FileSet hands out: the first file gets a base, the file occupies `size+1` values, the next file starts after it, and so on. Because the ranges do not overlap, a Pos identifies its file implicitly. The FileSet is the only thing that can decode that. It stores, per file, the name, the base, the size and the table of line-start offsets. So: ```go pos := fset.Position(node.Pos()) // pos.Filename, pos.Offset (byte index in that file), pos.Line, pos.Column ``` `fset.File(p)` gives you the `*token.File` a position belongs to, and `fset.PositionFor(p, adjusted)` lets you choose whether `//line` directives (used by code generators to point at their input) are applied. `token.NoPos` is the zero value and means "no position" — synthesised nodes you build yourself have it. The practical rules that follow: 1. **Create one FileSet per run** with `token.NewFileSet()` and parse every file into it. Do not make a fresh one per file or per print call. 2. **Never mix FileSets.** A Pos from one and a `Position` call on another yields the wrong file or nonsense; nothing panics, so the bug shows up as garbage line numbers. 3. **Keep the FileSet as long as you keep the tree.** Printing the tree back out with `go/format.Node(w, fset, file)` needs it, because the printer places comments and blank lines using the recorded positions. 4. **`Position.Offset` is your bridge to raw bytes.** If you also keep the original source, offsets let you splice text directly instead of re-printing the whole file — which is how you keep untouched lines byte-identical in a large mechanical change. ## A minimal pass ```go fset := token.NewFileSet() file, err := parser.ParseFile(fset, path, nil, parser.ParseComments) if err != nil { return err } ast.Inspect(file, func(n ast.Node) bool { if fn, ok := n.(*ast.FuncDecl); ok { fmt.Println(fset.Position(fn.Pos()), fn.Name.Name) } return true }) ``` That prints something of the form `path/to/file.go:12:1 Handler` for each function declaration — the FileSet is what makes the human-readable half of that line possible. ## Where this sits `go/parser` and `go/ast` are syntax only. `go/token` supplies positions. `go/format` prints a tree back as gofmt-formatted source. When you need to know what a name *means* — which package a selector resolves to, what type an expression has — you move up to a type-checked load, which reuses exactly the same `token.FileSet` and `*ast.File` values, so nothing you learn here is thrown away.
- What happens if you decode a node's position using a different token.FileSet than the one it was parsed into?You get silent nonsense — the wrong filename and line, or a zero Position — because the integer is an index into that particular FileSet's address space. Nothing panics and nothing warns you, so mismatched FileSets show up as inexplicably wrong diagnostics. Create one FileSet per run and thread it everywhere you thread the tree.
- Comments are missing from your parsed tree. What did you forget?The `parser.ParseComments` mode flag. Without it the parser drops comments entirely, since most consumers do not want them. With it, every comment group appears in `ast.File.Comments`, and the ones the parser could attach are also reachable through `Doc` and `Comment` fields on declarations, specs and struct fields.
- ParseFile returns an error. Is the *ast.File it returns usable?Usually yes. If the source could not be read the file is nil, but for a syntax error the parser returns a partial tree along with a `scanner.ErrorList` holding each error's position and message. Adding `parser.AllErrors` makes it report all of them rather than stopping early. Whether a partial tree is safe to act on depends on your pass.
The FileSet is one long tape measure laid end to end across every file you parsed. A node records only the number it sits at; only the tape knows that number 40312 means line 12 of handler.go.
saying these in an interview costs you the question
- Thinks token.Pos is a byte offset within its own file
- Creates a fresh token.FileSet per file or per print call
- Expects comments in the tree without parser.ParseComments
- Believes ParseFile type-checks or resolves imports
- Assumes a syntax error always means a nil *ast.File