How do you use xml.Decoder.DecodeElement to pull one repeated subtree out of a huge XML dump?
answer
- stream outside, unmarshal inside
- the tag you already read is an argument
- one record live at a time
- a pointer target, and the address of the start element
- the decoder resumes after the closing tag
basics
~20 sWalk the document with Decoder.Token; when a start element matches the one you want, call Decoder.DecodeElement(&v, &start) with that same start element. It reads the subtree through its matching end tag into v, so only one record is live at a time.
solid answer
~50 sThe hybrid loop is the standard answer for a file bigger than memory. You call `Decoder.Token` until you get an `xml.StartElement` whose name is the record element, then hand that start element back to the decoder: `dec.DecodeElement(&rec, &se)`. `DecodeElement` behaves like `Unmarshal` for the subtree rooted at `se` — it consumes everything through the matching `EndElement` and fills the Go value — and then the outer loop resumes at the next sibling. Peak memory is one record plus the decoder's buffer, whatever the file's size. Two details matter: you must pass a **pointer** to the target, and you must pass `&se` — the start element you already consumed. Calling `Decoder.Decode` instead would go looking for the *next* start element, which is now the record's first child, so you would silently decode the wrong thing. For records you do not want, `Decoder.Skip()` discards the subtree without allocating it.
code
go · 24 linestype Record struct {
ID string // matches a child element named id, case-insensitively
Name string
}
dec := xml.NewDecoder(r)
for {
tok, err := dec.Token()
if errors.Is(err, io.EOF) {
break
}
if err != nil {
return err
}
se, ok := tok.(xml.StartElement)
if !ok || se.Name.Local != "record" {
continue
}
var rec Record
if err := dec.DecodeElement(&rec, &se); err != nil {
return fmt.Errorf("record at byte %d: %w", dec.InputOffset(), err)
}
emit(rec) // rec becomes garbage on the next iteration
}go deeper
Know that a big XML file does not have to be unmarshalled in one go, and be able to say that Go can decode one element at a time into a struct while streaming the rest. The exact call is fair to look up.
Write the loop on a whiteboard: Token, type-assert the start element, match the name, DecodeElement with a pointer and the start element's address. Explain why the start element has to be handed back rather than re-read.
Argue the memory profile out loud — live set is one record regardless of file size — and have a per-record error policy ready that distinguishes a bad field from a corrupt stream position.
Decide whether an ingestion path should stream at all: bounded memory buys you predictable capacity, but moves structural validation into your loop and makes partial-failure semantics something your team has to define and support.
## The problem this solves A token loop keeps memory flat but gives you a firehose of syntactic pieces; assembling a record out of start, text and end tokens by hand is tedious and easy to get wrong once elements nest. Unmarshalling gives you a clean Go value but wants to read a whole element — and if that element is the root of a ten-gigabyte export, the whole file. `Decoder.DecodeElement` is the seam between the two. You stream at the top level and unmarshal at the record level. ```go func (d *Decoder) DecodeElement(v any, start *xml.StartElement) error ``` ## The loop 1. Create the decoder over the reader: `dec := xml.NewDecoder(r)`. 2. Call `dec.Token()` until it returns an `xml.StartElement`. 3. Test its `Name` — usually `Name.Local`, and `Name.Space` too if the document is namespaced. 4. On a match, declare a fresh target value and call `dec.DecodeElement(&rec, &se)`. 5. On no match, either continue the loop (to descend into it) or call `dec.Skip()` (to discard it). 6. Break on `io.EOF`. When `DecodeElement` returns, the decoder is positioned immediately after the record's closing tag, so the loop naturally lands on the next sibling. The record you just filled is an ordinary Go value that you write out, aggregate, or hand to a channel — and then it becomes garbage. That is the whole memory argument: live set is one record, not one document. ## Why the start element must be passed back The decoder has already consumed the opening tag; it cannot un-read it. `DecodeElement` therefore takes that tag as an argument so it knows the element's name and attributes, and knows which `EndElement` terminates the subtree. This is exactly why `Decode` is the wrong call here. `Decode` "works like Unmarshal, except it reads the decoder stream to find the start element" — it goes hunting for the *next* opening tag. Having already consumed `<record>`, the next opening tag is `<record>`'s first child, so `Decode` would decode that child and leave the loop misaligned. The failure is quiet: no error, just wrong data, which is the worst kind. ## The target value `v` must be a non-nil pointer; passing a struct by value gets you an error rather than a filled value. Allocate a fresh target per iteration, or zero a reused one — `DecodeElement` fills fields that appear in the document but does not clear fields that do not, so a reused struct can carry a previous record's values into a document that omitted them. Field matching follows the same rules `xml.Unmarshal` uses for any element, so a struct whose exported field names correspond to the child element names works with no further ceremony. ## Attributes on the record element Because you hand the whole `StartElement` in, its attributes are available to `DecodeElement` and land in the target the same way they would under `Unmarshal`. You can also read them yourself from `se.Attr` before decoding — useful when one attribute decides *whether* you want the record at all, letting you `Skip()` the ones you do not want without building them. If you keep a `StartElement` around beyond the next `Token` call rather than passing it straight to `DecodeElement`, copy it first with `se.Copy()`: the `Attr` slice points into the decoder's buffer. ## Errors An error from `DecodeElement` concerns one record. On a dirty multi-gigabyte feed, the operationally useful choice is usually to log it with `dec.InputOffset()`, count it, and carry on, rather than abandoning the run at record 4,000,000. But note that a *syntax* error leaves the decoder's position untrustworthy, whereas a type-conversion error inside a well-formed element does not — the recovery policy differs, and saying so is what separates a middle answer from a senior one. ## Shape of the result A pipeline built this way is a plain `for` loop over a reader with a fixed live set: one record, the decoder's buffer, and whatever your sink holds. It composes with a compressed or network reader unchanged, because the decoder only ever asks for the next bytes.
- What goes wrong if you call Decoder.Decode instead of DecodeElement after consuming the start element?`Decode` searches the stream for the next start element, and since you already consumed `<record>`, that is `<record>`'s first child. You decode the child into your record type and the loop is left misaligned in the subtree. There is usually no error at all — just wrong data, which makes it a nasty bug to find.
- How do you cheaply discard the records you are not interested in?Test what you can from the `StartElement` itself — its `Name` and its `Attr` values are already parsed — and call `Decoder.Skip()` when it fails the test. Skip consumes through the matching end tag without building any Go value, so a filtered pass over a huge feed costs almost nothing per rejected record.
- Is it safe to reuse one struct variable across iterations instead of declaring a fresh one?Only if you zero it. `DecodeElement` sets the fields present in the document and leaves the others untouched, so a record that omits an optional element will silently inherit the previous record's value. Declaring the variable inside the loop is the cheap, obviously correct option; the escape analysis usually makes it free anyway.
- A single record fails to decode halfway through a ten-gigabyte feed. Do you abort?It depends on the error. A type error inside a well-formed element affects one record, so logging it with `Decoder.InputOffset()`, counting it and continuing is usually right. A syntax error means the decoder's position in the document is no longer trustworthy, and continuing produces garbage — that one should stop the run.
saying these in an interview costs you the question
- Reaches for Unmarshal over the whole file and hopes memory holds
- Calls Decode after already consuming the start element
- Passes the target by value instead of a pointer
- Forgets to pass the start element's address at all
- Reuses one struct across records without zeroing it
- Expects the loop to see the record's EndElement afterwards