In an encoding/xml struct tag, what do the ,attr and ,chardata options and an a>b name mean?
answer
- name before the comma, options after
- one option picks a different XML place
- text is not a named child
- the arrow walks down a level
- one option hands back bytes untouched
basics
~20 sIn an encoding/xml tag, ,attr binds the field to an XML attribute rather than a child element, ,chardata binds it to the element's own text, and a name such as metadata>title descends through nested elements to the innermost one.
solid answer
~40 sThe tag is a name plus comma-separated options. `xml:"id,attr"` matches the attribute `id` on the current element instead of a child element. `xml:",chardata"` — a name-less tag with the option only — collects the character data directly inside the element, which is how you capture the text of `<price currency="EUR">9.99</price>` alongside its attribute. A tag whose name contains `>`, such as `xml:"metadata>title"`, tells the decoder to descend through the intermediate elements and bind the innermost one, so you can flatten a nested document without declaring a struct per level; on the encoding side, sibling fields sharing a path prefix are written inside one parent element. Two related options are `,innerxml`, which hands you an element's raw unparsed XML as a `string` or `[]byte`, and `,any`, which catches sub-elements no other field claimed.
code
go · 11 linestype Price struct {
Currency string `xml:"currency,attr"` // <price currency="EUR">
Amount string `xml:",chardata"` // 9.99
}
type Item struct {
XMLName xml.Name `xml:"item"`
ID string `xml:"id,attr"`
Title string `xml:"metadata>title"` // <metadata><title>...
Price Price `xml:"price"`
}go deeper
Learn the three forms by sight: a tag ending in ,attr reads an attribute, a tag that is only ,chardata reads the element's text, and a name with an arrow reaches a value nested one or more levels down.
Be able to write the tags for a document you are shown, and explain why XML needs options that a flat key/value format does not — attributes, text alongside children, and nesting.
Show judgment about how much of a vendor document to model: which subtrees get real tags, which are captured raw with ,innerxml and passed through, and how the choice affects what breaks when the vendor changes the schema.
Weigh a hand-maintained tag surface against generating types from a schema, and decide how much of a third-party document your codebase should encode as Go types that reviewers must then keep correct.
## The shape of the tag An `encoding/xml` field tag is a string of the form `xml:"name,option,option"`. The part before the first comma is the XML name to match; everything after it is options. Either half may be empty: `xml:"id,attr"` gives both, `xml:"title"` gives a name only, and `xml:",chardata"` gives an option only, leaving the name unused. ## ,attr — the value lives on the element, not under it XML can carry a value in two places, and they are not interchangeable: ```xml <item id="7"><title>Bolt</title></item> ``` `id` is an **attribute** of `<item>`; `title` is a **child element**. A field is matched against attributes only when its tag says `,attr`. This is the difference that most often produces an empty field for someone coming from a format with no attributes: `ID string \u0060xml:"id"\u0060` looks right and matches nothing. On output, `xml.Marshal` writes an `,attr` field into the start tag rather than as a nested element, so the same struct round-trips. ## ,chardata — the element's own text An element that has attributes *and* text needs somewhere to put the text, because the text is not a named child: ```xml <price currency="EUR">9.99</price> ``` Map it with a small struct: one field tagged `currency,attr` and one tagged `,chardata`. The `,chardata` field accumulates the character data inside the element. It works on a `string` or `[]byte` field. A caution worth knowing: if a struct has both a `,chardata` field and child-element fields, the character data field also collects the whitespace *between* those child elements, because that whitespace is character data too. Use `,chardata` on leaf-ish elements — ones whose content is text — and not on a container. ## a>b — descending without a struct per level Documents produced by enterprise tooling nest deeply, and declaring a Go struct for every wrapper level is tedious. A tag name containing `>` walks the path for you: ```go type Item struct { Title string `xml:"metadata>title"` Author string `xml:"metadata>author"` } ``` That binds `<item><metadata><title>…</title><author>…</author></metadata></item>` to two flat fields. A tag beginning with `>` is shorthand for the field's own name followed by `>`. On the encoding side the rule is mirrored: fields that name the same parent path and appear next to each other are enclosed in one shared parent element, so marshalling the struct above emits a single `<metadata>` containing both children. The path form binds the innermost name; it does not let you skip an arbitrary number of levels or use wildcards. Every step must be spelled out. ## ,innerxml — keep the bytes ```go type Entry struct { Summary string `xml:"summary"` Raw []byte `xml:",innerxml"` } ``` A `,innerxml` field receives the raw, unprocessed XML nested inside the element, verbatim. The other fields are still populated normally — `,innerxml` is additive, not exclusive. It is the escape hatch for a subtree whose shape you do not want to model: pass it through untouched, store it, or decode it later with a second struct. It is also the fastest way to see what a vendor document actually contains at a given level. ## ,any and the rest `,any` marks a field as the catch-all for sub-elements that matched no other field, which is useful for a heterogeneous list. `,comment` collects XML comments. `,cdata` (on output) writes the field's text inside a CDATA section. `,omitempty` affects marshalling only: it suppresses an empty field from the output, and it applies to attributes as well as elements. ## Why the grammar looks like this Every one of these options exists because XML's data model is richer than a struct's: an element has a name, ordered children, unordered attributes, text, comments and namespaces, while a Go struct has only named fields. The tag is the small language that projects one onto the other, and each option names one XML construct that would otherwise have no Go home. Learning the five that matter — `,attr`, `,chardata`, the `a>b` path, `,innerxml`, `,any` — covers nearly every real document.
- When would you reach for ,innerxml instead of modelling the subtree?When you must pass a fragment through unchanged — storing a vendor payload you do not own, or forwarding it — and when you are exploring an undocumented document and want to see the bytes at a level before writing tags for it. The field takes a `string` or `[]byte` and the other fields still decode normally.
- What happens if a struct has both a ,chardata field and child-element fields?The `,chardata` field collects all character data directly inside the element, including the newlines and indentation between the child elements, so you typically get whitespace rather than a useful value. Put `,chardata` on a small struct that models a text-carrying element, not on a container.
- Does the a>b path form work for encoding as well as decoding?Yes, and symmetrically: `xml.Marshal` writes the field nested inside the named parent elements, and adjacent fields that share a parent path are wrapped in a single shared parent rather than one each. That keeps a flattened struct round-tripping back to the document shape it came from.
saying these in an interview costs you the question
- Uses a plain element tag for a value that is an attribute
- Thinks a>b can skip an unknown number of levels or wildcard
- Puts ,chardata on a container and is surprised by whitespace
- Believes ,innerxml replaces the other fields instead of adding to them
- Assumes ,omitempty affects decoding rather than only output