skip to content

What does the Space field of an xml.Name hold for a token returned by xml.Decoder.Token?

level: middleimportance: should knowfreq 38%

answer

  1. two fields, one of them surprises people
  2. the prefix is only shorthand
  3. identity lives in the declaration's value
  4. a default declaration means no prefix appears
  5. RawToken is the unresolved view

basics

~20 s

Space holds the resolved namespace URI, not the prefix written in the document. Token expands each prefix through the xmlns declarations in scope and discards it, leaving Local as the bare element name. Elements in no namespace have an empty Space.

solid answer

~50 s

`xml.Name` is `struct { Space, Local string }`, and `Decoder.Token` fills it *after* namespace resolution. For `<a:item xmlns:a="urn:example:v1">` you get `Space == "urn:example:v1"` and `Local == "item"` — the literal prefix `a` is gone. A default declaration works the same way: under `xmlns="urn:example:v1"`, a plain `<item>` also gives that URI. This is the correct XML model: the prefix is a local shorthand, the URI is the identity, and two documents that spell the same namespace with different prefixes are the same document. So match on `Space` and `Local` together and never on the prefix. Elements outside any namespace get an empty `Space`, and `Decoder.DefaultSpace` lets you supply a URI for unadorned tags in a document that omits the declaration. `Decoder.RawToken` is the escape hatch: it does not translate prefixes, so it reports the document as written.

code

go · 12 lines
go
const atom = "http://www.w3.org/2005/Atom"

switch tok := tok.(type) {
case xml.StartElement:
	if tok.Name.Space == atom && tok.Name.Local == "entry" {
		// matches <entry> under a default xmlns and <a:entry> alike
		var e Entry
		if err := dec.DecodeElement(&e, &tok); err != nil {
			return err
		}
	}
}

go deeper

for a junior

Remember that xml.Name has two fields and that the one called Space carries the namespace URI, not the short prefix you see in the file. Matching uses both fields together.

for a middle

Explain how Token resolves prefixes against the xmlns declarations in scope, and why a default declaration means an element can be namespaced with no prefix on it at all. Be able to write the two-field comparison.

for a senior

Show why a Local-only match is a data-correctness risk in a feed that mixes vocabularies, and know the levers — Decoder.DefaultSpace for an exporter that omits declarations, RawToken when the literal spelling is what you are auditing.

for a principal

Decide what your ingestion contract is: pinning to namespace URIs makes you robust to a supplier's cosmetic changes but couples you to their versioning scheme, since a new URI is how XML vocabularies signal a breaking change.

## The type ```go type Name struct { Space, Local string } ``` Every `xml.StartElement`, `xml.EndElement` and `xml.Attr` carries one. `Local` is the bare name. `Space` is the part people get backwards. ## Space is the URI, not the prefix XML namespaces work in two layers. In the document you write a short **prefix** and bind it with an `xmlns:` attribute: ```xml <a:item xmlns:a="urn:example:v1">…</a:item> ``` The prefix `a` is arbitrary and scoped to where it is declared. The **URI** `urn:example:v1` is the actual identity of the vocabulary. Rewriting every `a:` to `x:` and changing the declaration to match produces a different byte sequence and *the same document*. `Decoder.Token` implements that model. It tracks the `xmlns` declarations in scope, resolves each name it returns, and hands you the URI in `Space`. The prefix is not preserved anywhere in the token. So after `Token`, both of these give you `Space == "urn:example:v1"`, `Local == "item"`: ```xml <a:item xmlns:a="urn:example:v1"/> <item xmlns="urn:example:v1"/> ``` That second form — a **default** namespace declaration, with no prefix — is why matching on the presence of a colon is hopeless: an element can be in a namespace without any prefix appearing on it at all, and the declaration may sit several levels up the tree. ## How to match Compare both fields against constants: ```go const atom = "http://www.w3.org/2005/Atom" if se.Name.Space == atom && se.Name.Local == "entry" { … } ``` Matching on `Local` alone is the common shortcut, and it is fine for a single-vocabulary document you control. It stops being fine the moment a feed mixes vocabularies — two schemas can both define `<title>` or `<id>`, and a `Local`-only match will happily pull the wrong one. On a reconciliation job between two systems that is precisely the class of bug that surfaces as "a few thousand rows have the wrong value" rather than as a crash. ## Where the declarations themselves go The `xmlns` attributes are not silently swallowed: they arrive in `StartElement.Attr` alongside ordinary attributes. A prefixed declaration `xmlns:a="…"` comes through with `Local` equal to the prefix `a` and `Space` equal to `xmlns`; a default declaration `xmlns="…"` comes through with `Local` equal to `xmlns`. You rarely need them, but if you are re-emitting a document you will care, because they are the only place the original prefixes still exist. ## Unprefixed attributes are not in the default namespace A subtlety worth knowing: in XML, a default `xmlns` declaration applies to elements but **not** to unprefixed attributes. So under `<item xmlns="urn:example:v1" id="7">`, the element's `Space` is the URI while the `id` attribute's `Space` is empty. Code that expects attributes to inherit the element's namespace will fail to match them. ## When there is no namespace A document with no declarations at all yields an empty `Space` on everything, which is a perfectly valid state, not an error. If you are consuming a feed that *should* be namespaced but is not — an old exporter, a hand-written file — `Decoder.DefaultSpace` lets you set the URI to assume for unadorned tags, so downstream matching code can be written once against the namespaced form. ## RawToken `Decoder.RawToken` returns the same token types but performs no namespace resolution and no start/end matching check. With it, `<a:item>` reports the prefix, not the URI. That is what you want when the document's literal spelling is the thing you are inspecting; it is the wrong default for anything that has to interpret the content.

  • Under a default declaration such as <item xmlns="urn:example:v1" id="7">, what is the id attribute's Space?
    Empty. A default `xmlns` declaration applies to element names but not to unprefixed attribute names — that is an XML rule, not a Go quirk. The element's `Space` is `urn:example:v1` while the attribute's is `""`. Code that assumes attributes inherit the element's namespace quietly stops matching them.
  • The feed you consume should be namespaced but the exporter emits bare tags. What can you do at the decoder?
    Set `Decoder.DefaultSpace` to the URI those unadorned tags ought to have. The decoder then reports them as if the document were wrapped in an element carrying that default declaration, so one matching path handles both the broken exporter and the correct one without a second code branch.
  • When would you use Decoder.RawToken rather than Token here?
    When the literal spelling is the subject — auditing which prefixes a supplier actually uses, or reproducing a document as written. `RawToken` does no prefix-to-URI translation and does not verify that start and end elements match, so it reports the document rather than interpreting it. For anything that consumes content, `Token` is the right call.

saying these in an interview costs you the question

  • Says Space holds the prefix, such as a or soap
  • Matches elements on Local only in a mixed-vocabulary feed
  • Believes changing a prefix changes the document's meaning
  • Tests for a colon in the name to detect a namespace
  • Assumes unprefixed attributes inherit the default namespace
  • Thinks an empty Space means the document is malformed