skip to content

Why does re-emitting decoded tokens through xml.Encoder rewrite a document's namespace prefixes?

level: seniorimportance: should knowfreq 26%

answer

  1. the round trip loses one thing only
  2. identity survives, spelling does not
  3. the encoder keeps no prefix map
  4. elements get a default declaration each
  5. attributes get an invented prefix
  6. compare token streams, not bytes

basics

~20 s

Decoder.Token resolves each prefix to its namespace URI and discards the prefix, and xml.Encoder has no prefix table to restore it. It writes an element's namespace as a default xmlns declaration on that element and invents a prefix for namespaced attributes.

solid answer

~50 s

The prefix does not survive the round trip because nothing carries it. `Token` resolves `a:item` against the `xmlns` declarations in scope and gives you `Name{Space: "urn:example:v1", Local: "item"}` — correct XML semantics, since a prefix is only local shorthand. On the way out, `xml.Encoder` writes `Name.Local` as the tag and, when `Name.Space` is non-empty, emits it as a **default** `xmlns="…"` declaration on that element; for a namespaced *attribute* it generates a prefix of its own and declares it inline. So `<a:item xmlns:a="urn:example:v1">` comes back as `<item xmlns="urn:example:v1">`. The document is namespace-equivalent and the diff against the source is enormous. A consumer matching on the literal string `a:item` breaks — that consumer is wrong by the XML spec, but you still have to ship. The practical options are: pass bytes through untouched for the parts you are not editing, or emit tokens with the prefix baked into `Local` and the `xmlns` attribute written by hand.

code

text · 5 lines
text
source:      <a:item xmlns:a="urn:example:v1">x</a:item>

token:       StartElement{Name: {Space: "urn:example:v1", Local: "item"}}

re-emitted:  <item xmlns="urn:example:v1">x</item>

go deeper

for a junior

Take away one fact: decoding and re-encoding XML in Go gives you an equivalent document, not the same bytes. The namespace is preserved; the prefix you saw in the file is not.

for a middle

Explain both halves of the mechanism — Token resolves the prefix to a URI and drops it, and the encoder re-spells it as a default xmlns declaration per element, with a generated prefix for namespaced attributes.

for a senior

Diagnose the broken consumer without hand-waving: name the encoder as the source of the rewrite, propose passing untouched bytes through rather than round-tripping, and replace the byte diff with a token-stream comparison as the real check.

for a principal

Decide what fidelity your integration promises. Byte-preserving pass-through, semantic equivalence, or canonical output are three different contracts with different costs, and choosing one for the platform is what stops each team inventing its own.

## The scenario You are reconciling two systems. A multi-gigabyte namespaced export comes in, a Go pipeline streams it with `Decoder.Token`, edits a handful of elements, and writes the rest back out with `Encoder.EncodeToken`. The output is diffed against the source as a sanity check, and the diff is the whole file. Downstream, a consumer that matches on `a:item` reports zero records. Nothing is corrupt. Two design decisions, one on each side of the round trip, combined. ## Decoding throws the prefix away — on purpose In XML, a prefix is shorthand scoped to where it is declared; the namespace **URI** is the identity. `<a:item xmlns:a="urn:example:v1">` and `<x:item xmlns:x="urn:example:v1">` are the same document. `Decoder.Token` models this correctly: it tracks the declarations in scope, resolves the name, and reports `Name{Space: "urn:example:v1", Local: "item"}`. The character `a` appears nowhere in the token — except incidentally, in the `xmlns:a` attribute that is still present in `StartElement.Attr`. (`Decoder.RawToken` is the exception: it performs no prefix translation, so it reports the document as written. It also stops verifying that start and end elements match, which is the price.) ## Encoding invents its own spelling `xml.Encoder` has no notion of a prefix map you can populate. When it writes a start element it emits `Name.Local` as the tag name, and if `Name.Space` is non-empty it appends a **default** declaration `xmlns="<Space>"` to that element. It does this on every element that carries a `Space`, so a deeply nested document ends up repeating the declaration rather than binding it once at the root. Namespaced attributes cannot use a default declaration — in XML, a default `xmlns` never applies to unprefixed attribute names — so the encoder generates a prefix for the URI and declares it inline the first time it needs one. That generated prefix is derived from the URI and bears no relationship to whatever the source used. The result is namespace-equivalent to the input and textually unrecognisable. ## Why the diff is the wrong instrument A byte diff of the re-emitted document against the source will light up on every element. It is measuring the wrong thing. The check that actually answers "did I change the meaning" is a comparison of the two **token streams**: decode both files, and compare the sequence of resolved `(Space, Local)` pairs, attribute sets and trimmed character data. If those match, the documents say the same thing, whatever the diff says. Build that comparison once and it becomes the regression test for the whole pipeline. ## The downstream consumer A consumer matching on `a:item` is relying on spelling, not meaning; by the namespaces specification it is broken and would break on any conforming rewriter, in any language. Say that plainly — and then fix the problem, because "they are wrong" does not restore their feed. Three ways out, in rough order of preference: **1. Do not round-trip what you are not changing.** If the job edits a few elements, the safest pipeline copies source bytes through for everything else, so the untouched regions are byte-identical by construction. Decode-and-re-encode is a rewrite of the whole document even when you meant to change one field. **2. Control the spelling by hand.** The encoder writes `Name.Local` verbatim. So if you construct tokens with `Local: "a:item"` and `Space: ""`, and add the `xmlns:a` binding yourself as an ordinary attribute, you get exactly the output you want. It is manual and easy to get wrong across nesting, but it is deterministic, and it is the standard workaround for the encoder having no prefix control. **3. Fix the consumer to match on the URI.** Correct, cheapest in the long run, and often not available on the timeline you have — which is the real content of the decision. ## The case where you must not re-emit at all If the document carries a signature or a digest computed over its serialised form, any re-serialisation invalidates it, regardless of prefixes. Whitespace, attribute order and declaration placement all move. A pipeline that has to preserve such a document treats it as opaque bytes and never runs it through a decoder-and-encoder pair. ## Loose ends worth knowing - `Encoder.Close` (added in Go 1.20) flushes buffered output and reports any element left unclosed; `Encoder.Flush` only flushes. Ending a hand-built token stream without one of them can truncate the output. - The original prefixes are still visible to you: the `xmlns` declarations arrive as ordinary entries in `StartElement.Attr`, so a pipeline that must reproduce them can read them from there.

  • The downstream team matches on the literal string a:item. Whose bug is it, and what do you actually do?
    Theirs, by the namespaces specification: a prefix is local shorthand and any conforming rewriter may change it. But that argument does not restore their feed. In practice you either stop round-tripping the parts you are not editing so the bytes pass through unchanged, or you construct the output tokens with the prefix baked into `Name.Local`, while helping them move to URI matching.
  • How do you prove the re-emitted document means the same as the source when the byte diff is the whole file?
    Compare token streams instead of text. Decode both documents and walk them in parallel, comparing the sequence of resolved `(Space, Local)` pairs, attribute sets and trimmed character data. That is the comparison that reflects XML's own equality rules, and it is cheap enough to keep as the pipeline's regression test.
  • Is there a document you should refuse to pass through a decode-and-re-encode pipeline at all?
    One whose serialised bytes are load-bearing — anything carrying a signature or digest computed over the document text. Re-serialisation moves whitespace, attribute order and declaration placement even when the meaning is untouched, so the verification fails. Treat such documents as opaque bytes and edit them, if at all, without a full round trip.
  • Where can you still find the prefixes the source document used?
    In `StartElement.Attr`. The `xmlns` declarations are not swallowed by the decoder; they arrive as ordinary attributes, a prefixed one carrying the prefix in `Name.Local`. A pipeline that has to reproduce the original spelling reads its prefix map from there and applies it when constructing output tokens.

saying these in an interview costs you the question

  • Claims the round trip loses the namespace itself, not just the prefix
  • Thinks setting Name.Space chooses the output prefix
  • Assumes decode-then-encode is byte-preserving
  • Blames the decoder when the encoder chooses the spelling
  • Concludes from a byte diff that the document changed meaning
  • Re-emits a signed document and wonders why verification fails