xml.Unmarshal returns no error on a vendor's sample feed, yet one struct field is always empty. How do you find the cause?
answer
- a nil error only means well-formed
- replay the vendor's own sample in a test
- exported, spelled right, attribute or element
- one toolchain command checks the tag syntax
- capture the raw bytes one level up
basics
~20 sencoding/xml never errors on input it cannot place, so an empty field means nothing matched it. Check that the field is exported, that the tag name equals the element name, that attributes carry ,attr, and that a nested value needs a path tag.
solid answer
~50 sStart from the fact that a nil error only means the bytes were well-formed XML — `encoding/xml` silently drops anything that matches no field. Pin the vendor's sample document into a test so the failure is reproducible, then walk the matching rules in order: is the field exported; does the tag name match the element's name exactly; is the value actually an attribute, which needs `,attr`; is it nested deeper, needing an `a>b` path; is the element repeated into a non-slice field so only the last one survives; did the tag get written with `"` instead of backticks or a space after `xml:`, which makes it not a tag at all. Run `go vet`, whose struct-tag check catches that last class. If the rules all look right, add a `,innerxml` field at the level above and print it — that shows exactly what reached that point.
code
go · 10 linestype probe struct {
XMLName xml.Name
Raw []byte `xml:",innerxml"`
}
var p probe
if err := xml.Unmarshal(sample, &p); err != nil {
t.Fatal(err)
}
t.Logf("element=%v raw=%s", p.XMLName, p.Raw)go deeper
Know the first two things to check when a decoded field is empty: the field must be exported, and the tag name must match the element's name. An empty field is normal behaviour, not a crash.
Be able to run the whole matching checklist out loud — exported, tag syntax, name, attribute versus element, nesting depth, repetition — and explain why none of it is caught at compile time.
Show a method, not a guess: replay the vendor's sample in a test, use go vet for tag syntax, capture the raw bytes with an innerxml probe, then leave behind per-field assertions so the next schema change fails loudly.
Set the standard for depending on a document you do not control: what evidence a converter must carry before it ships, and whether the team invests in schema-generated types rather than hand-maintained tags a reviewer must verify by eye.
## Why there is no error to read `encoding/xml` reports malformed XML, a Go type it cannot fill, and an `XMLName` tag that does not match the element it was pointed at. It does **not** report input it could not place, and it does not report fields nothing filled. So on a converter replacing a hand-rolled parser, the classic symptom is exactly this: `xml.Unmarshal` returns `nil`, most of the struct is populated, and one field is stubbornly zero. The bug is always in the matching rules, and the job is to walk them in order. ## Make it reproducible first Before theorising, put the vendor's own sample document into a test — a table of `[]byte` fixtures decoded into the struct with assertions on the fields you care about. Two reasons. First, you will change the tags several times and need a fast loop. Second, the assertion you are about to write is the one that was missing: *this field is non-empty after decoding*. A converter that only asserts `err == nil` will regress the same way again the next time the vendor edits the schema. ## The checklist, cheapest first **1. Is the field exported?** An unexported field is invisible to reflection and stays at its zero value with no complaint. This is the number-one cause and costs nothing to check. **2. Is the tag actually a tag?** A struct tag has to be a back-quoted string with no space after the key's colon. Written with double quotes, or as `xml: "id,attr"`, it is just a string the reflect package cannot parse into pairs — and the field silently falls back to matching on its Go name. `go vet` has a struct-tag check that reports exactly this, so run it before staring at the document. **3. Does the tag name match the element name?** The comparison is against the element's own name; a typo, a different spelling or a plural you assumed is enough to miss. **4. Is the value an attribute?** `<item id="7">` needs `xml:"id,attr"`. A plain `xml:"id"` matches a child element named `id` that does not exist. Legacy vendor documents are attribute-heavy, so this one recurs. **5. Is it nested deeper than you think?** Wrapper elements are easy to miss when reading a large sample. Either declare a struct for the wrapper or use a path tag, `xml:"metadata>title"`. **6. Is the element repeated?** Three matching elements into a plain `string` leave you with the last one — which looks like a decode bug when the value you were checking came first. Repetition belongs in a slice field. **7. Did you pin a namespace?** A tag written as `xml:"http://example.com/ns title"` constrains the match to that namespace, so a document using a different one will not match. A tag with no namespace matches whatever the document uses, which is why the plain form is usually the right default when you are decoding. ## When the checklist runs out: look at the bytes If the tags look right, stop guessing about the document and make it show you. Add a field tagged `,innerxml` to the struct one level *above* the missing field and print it after decoding: ```go type probe struct { XMLName xml.Name Raw []byte `xml:",innerxml"` } ``` The raw bytes answer every remaining question at once: whether the element is there at all, what it is really called, whether the value is an attribute, and how deep it sits. An untagged `XMLName` field additionally records the name of the element you actually landed on, which catches the case where the whole struct is being fed a different element from the one you assumed. A second, blunter trick: put a tagged `XMLName` on the wire struct. That converts one silent class of mistake — pointing the decoder at the wrong element — into a returned error, permanently. ## Close the loop When you find it, fix the tag *and* keep the assertion. A converter reading a document you do not control should have, for each field that matters, a test over the vendor's sample proving the field is populated. That is the only durable defence against a codec whose contract is to say nothing when it does not understand you.
- How do you stop this class of bug from reaching production again?Check the vendor's sample document into the repository and assert, per field that matters, that it is populated after decoding — not merely that the error is nil. Add a tagged `XMLName` field so pointing the decoder at the wrong element becomes a returned error, and keep `go vet` in the build so a malformed struct tag fails there.
- Which of these mistakes does the compiler catch?None of them. A struct tag is an ordinary string literal, so a typo, a wrong option or the wrong quoting compiles cleanly; the mismatch only shows at run time as an empty field. `go vet`'s struct-tag check catches malformed syntax and duplicate names, which is the closest thing to compile-time help you get.
- The vendor sends the same element three times but your field is a string — what do you actually see?The text of the last occurrence, because each match overwrites the previous one and nothing reports the collision. If your test happened to check the first occurrence's value it reads as a wrong-value bug rather than a shape bug. The fix is a slice field, which keeps all three in document order.
saying these in an interview costs you the question
- Concludes the document is malformed because the field is empty
- Trusts a nil error from xml.Unmarshal as proof the data landed
- Never checks whether the value is an attribute rather than an element
- Assumes the compiler validates struct tag syntax
- Fixes the tag but leaves the test asserting only err == nil