skip to content

What does each call to xml.Decoder.Token return while streaming a large XML file in Go?

level: juniorimportance: must knowfreq 55%

answer

  1. one call, one syntactic piece
  2. the return type is an interface
  3. start, end, chardata, comment, procinst, directive
  4. indentation counts as text
  5. the loop stops on io.EOF

basics

~20 s

Token returns the next piece of the document — an xml.StartElement, xml.EndElement, xml.CharData, xml.Comment, xml.ProcInst or xml.Directive — plus an error. At the end of the input it returns io.EOF, so you loop until then.

solid answer

~40 s

`Decoder.Token` walks the document one syntactic piece at a time instead of building a tree, which is what makes it usable on a dump larger than memory. Its return type is `xml.Token`, declared as `type Token any`, so you type-switch on it: `xml.StartElement` and `xml.EndElement` for tags, `xml.CharData` for the text between them, plus `xml.Comment`, `xml.ProcInst` and `xml.Directive`. A self-closing tag still yields both a start and an end token. The loop terminates when `Token` returns `io.EOF`; any other error is a real failure, typically an `*xml.SyntaxError` on malformed or truncated input. Two things surprise people: the indentation between tags arrives as `xml.CharData` tokens, and the byte slices inside `CharData`, `Comment` and `Directive` point into the decoder's own buffer, so they are only valid until the next `Token` call.

code

go · 19 lines
go
dec := xml.NewDecoder(f) // f is an *os.File over a multi-gigabyte dump
counts := map[string]int{}

for {
	tok, err := dec.Token()
	if errors.Is(err, io.EOF) {
		break
	}
	if err != nil {
		return fmt.Errorf("xml at byte %d: %w", dec.InputOffset(), err)
	}
	switch t := tok.(type) {
	case xml.StartElement:
		counts[t.Name.Local]++
	case xml.CharData:
		// t is borrowed: valid only until the next Token call
		textBytes += int64(len(bytes.TrimSpace(t)))
	}
}

go deeper

for a junior

Be ready to name the token types and to write the loop from memory: call Token, break on io.EOF, type-switch on the result. Knowing that the return type is an interface is the part interviewers actually check.

for a middle

Explain why the loop keeps memory flat where Unmarshal does not, and account for the whitespace CharData tokens a pretty-printed file produces. Mention that self-closing tags still yield both a start and an end token.

for a senior

Show that you know the byte slices are borrowed from the decoder's buffer and that retained tokens need CopyToken. Attach Decoder.InputOffset to parse errors so a failure four gigabytes into a feed is locatable.

for a principal

Own the call between a token loop and a schema-bound decode for a whole ingestion path: streaming buys bounded memory but pushes structural validation into your own code, and that is a cost the team pays in every future change.

## Two ways into encoding/xml `xml.Unmarshal` and `Decoder.Decode` build a Go value out of a whole element: they read the element and everything under it, allocate, and hand you a struct. That is fine for a configuration file. For a multi-gigabyte export it is not — the peak footprint is the document. `Decoder.Token` is the other door. It is a *pull parser*: each call advances the decoder by one syntactic piece and returns it, and nothing accumulates. Memory stays flat no matter how big the file is, because the only thing alive at any moment is the current token plus whatever you chose to keep. ## The signature and the concrete types ```go func (d *Decoder) Token() (xml.Token, error) ``` `xml.Token` is declared as `type Token any` — an empty interface — so the value carries no useful methods and you must type-switch on it. The concrete types the decoder can hand back are: - **`xml.StartElement`** — `struct { Name xml.Name; Attr []xml.Attr }`. An opening tag, with its attributes already parsed. - **`xml.EndElement`** — `struct { Name xml.Name }`. A closing tag. - **`xml.CharData`** — a `[]byte` of character data between tags, with entity references such as `&amp;` already expanded. - **`xml.Comment`** — a `[]byte` holding the body of a `<!-- … -->`. - **`xml.ProcInst`** — `struct { Target string; Inst []byte }`, a processing instruction; the leading `<?xml version="1.0"?>` declaration arrives this way. - **`xml.Directive`** — a `[]byte` for a `<!DOCTYPE …>` and friends. A self-closing element such as `<br/>` is not a special case: the decoder synthesises a `StartElement` followed immediately by the matching `EndElement`, so a loop that tracks depth by counting starts and ends stays balanced. ## Ending the loop correctly When the input is exhausted, `Token` returns `nil, io.EOF`. That is the normal exit and must not be reported as a failure. Anything else is a genuine error — most often an `*xml.SyntaxError`, which is what you get for mismatched tags, an unexpected end of input in the middle of an element, or bytes that are not valid UTF-8. The idiomatic shape: ```go for { tok, err := dec.Token() if errors.Is(err, io.EOF) { break } if err != nil { return err } // use tok } ``` `Decoder.InputOffset()` returns the byte offset just past the most recently returned token, which is the single most useful thing to attach to an error message when the failure is 4 GB into a file. ## Whitespace is character data A pretty-printed document is mostly indentation, and every run of it comes back as an `xml.CharData` token. A loop that assumes it only sees elements will silently miscount, and code that concatenates every `CharData` it meets will collect newlines and tabs. Filter with `bytes.TrimSpace` when you only want real text. ## The borrowed-bytes rule This is the trap that turns a working prototype into a corrupt output file. The byte slices inside `CharData`, `Comment` and `Directive`, and the `Attr` slice inside `StartElement`, refer to the decoder's internal buffer. They stay valid only until the next call to `Token`. If you append tokens to a slice, stash them in a map, or hand them to another goroutine, the contents change underneath you. The fix is in the package: `xml.CopyToken(t)` returns a token that owns its memory, and each of the token types has its own `Copy` method (`StartElement.Copy`, `CharData.Copy`, and so on). Converting to a `string` also copies, because a Go string conversion of a `[]byte` allocates. ## Skipping and raw tokens When you are positioned on a `StartElement` you do not care about, `Decoder.Skip()` consumes tokens until the matching `EndElement` without materialising anything — cheaper and less error-prone than counting depth yourself. `Decoder.RawToken()` is the lower-level sibling: it returns the same token types but does not check that start and end elements match and does not translate namespace prefixes into URIs. You reach for it only when you deliberately want the document as written rather than as resolved. ## Why this matters in an interview The token loop is the answer to "how would you process an XML feed you cannot fit in RAM", and the three details that separate a confident answer from a hesitant one are: the return type is an interface, `io.EOF` is the terminator, and the bytes are borrowed.

  • Why can the bytes inside an xml.CharData token you saved earlier turn into something else later in the loop?
    Because they are not yours. `CharData`, `Comment`, `Directive` and the `Attr` slice inside a `StartElement` point into the decoder's internal buffer, which is reused on the next `Token` call. Anything you retain past that call must be copied first — `xml.CopyToken`, the token's own `Copy` method, or a conversion to `string`.
  • You are sitting on a StartElement whose subtree you do not want. What is the cheapest way to get past it?
    Call `Decoder.Skip()`. It consumes tokens until the `EndElement` matching the start element you just read, handling nesting for you, and it allocates nothing for the discarded content. Doing it by hand means counting starts and ends yourself, which is where off-by-one bugs on self-closing tags creep in.
  • Why do you see xml.CharData tokens that contain only whitespace?
    Because in XML the whitespace between tags is character data, and the decoder reports it faithfully. A pretty-printed document produces one `CharData` token per run of indentation. Trim with `bytes.TrimSpace` and drop the empties if you only want real text; do not assume the parser has stripped them.

saying these in an interview costs you the question

  • Claims Token reads the whole document into memory first
  • Treats io.EOF as a failure, or loops forever without checking it
  • Assumes only elements come back and ignores CharData
  • Keeps a CharData slice past the next Token call without copying
  • Thinks Token returns a struct, so no type switch is needed
  • Believes a self-closing tag yields only a start token