How do you make xml.Decoder read a document whose declaration says encoding="ISO-8859-1"?
answer
- the standard library parses one encoding
- there is a hook, not a table
- you hand back a reader
- unset means a refusal, not a guess
- no declaration means the hook never fires
basics
~10 sSet Decoder.CharsetReader to a function that wraps the input reader and returns UTF-8. encoding/xml parses only UTF-8, so when a declaration names another encoding and CharsetReader is unset the decoder fails instead of guessing.
solid answer
~50 s`encoding/xml` decodes UTF-8 and nothing else. When the XML declaration names a different encoding, the decoder calls `Decoder.CharsetReader`, whose signature is `func(charset string, input io.Reader) (io.Reader, error)`. You return a reader that transcodes the bytes to UTF-8; the decoder then parses that. If the field is left unset, parsing stops with an error saying the encoding was declared but `Decoder.CharsetReader` is nil — deliberate, because silently misreading a Latin-1 document as UTF-8 corrupts every accented character. Return an error from your function for labels you do not support, and it surfaces from `Token`. The `charset` package under `golang.org/x/net/html` and the encoding packages under `golang.org/x/text` provide ready converters. One catch: `CharsetReader` is only consulted for an encoding **declared** in the document. A file with no declaration whose bytes are actually Windows-1252 will instead fail with a syntax error about invalid UTF-8, and you have to wrap the reader yourself before `xml.NewDecoder`.
code
go · 10 linesdec := xml.NewDecoder(f)
dec.CharsetReader = func(label string, in io.Reader) (io.Reader, error) {
switch {
case strings.EqualFold(label, "ISO-8859-1"), strings.EqualFold(label, "latin1"):
return newLatin1Reader(in), nil // each byte becomes the rune of the same value
case strings.EqualFold(label, "UTF-8"):
return in, nil
}
return nil, fmt.Errorf("unsupported xml charset %q", label)
}go deeper
Know that Go's XML decoder handles UTF-8 only and will refuse rather than guess when a document declares something else. The field to set is Decoder.CharsetReader.
State the hook's signature and what your function has to return: a reader yielding UTF-8. Explain why the library ships no charset tables and where the converters actually live.
Distinguish the declared case from the undeclared one, since only the first reaches the hook, and make the source encoding an explicit input to a reconciliation job rather than trusting a supplier's declaration.
Own the policy for a fleet of feeds: whether you normalise everything to UTF-8 at the boundary, who is accountable when a supplier's declaration lies, and what a run does when it meets a charset nobody configured.
## Why the decoder refuses Go source is UTF-8, Go strings are UTF-8 by convention, and `encoding/xml` parses UTF-8. XML in the wild is not so tidy: exports from older systems routinely arrive as ISO-8859-1 or Windows-1252, announced in the declaration: ```xml <?xml version="1.0" encoding="ISO-8859-1"?> ``` The standard library does not carry the world's character-set tables — those live in the `golang.org/x/text` repositories — so it cannot transcode by itself. What it does instead is give you a hook and, if you have not filled it in, fail loudly. The error reads along the lines of `xml: encoding "ISO-8859-1" declared but Decoder.CharsetReader is nil`. Failing is the right behaviour. Latin-1 and UTF-8 agree on the ASCII range, so a naive read of a Latin-1 document *mostly* works — until a name with an accent or a price with a currency symbol goes through, and then you get replacement characters or a parse error somewhere unrelated. Corruption that only affects a fraction of rows is far more expensive to discover than a refusal at byte zero. ## The hook ```go type Decoder struct { CharsetReader func(charset string, input io.Reader) (io.Reader, error) // … other fields } ``` The decoder calls it with the charset label exactly as the document spelled it — so compare case-insensitively, and be ready for the aliases (`latin1`, `ISO-8859-1`, `windows-1252`, `cp1252`). You return an `io.Reader` that yields the same content as UTF-8; the decoder replaces its own reader with yours and carries on. Returning an error rejects the document, and that error surfaces from the next `Token` call. For the actual conversion, the quasi-standard repositories cover it: the `charset` package under `golang.org/x/net/html` maps a label straight to a decoding reader, and the encoding packages under `golang.org/x/text` expose the individual code pages. For the single-byte Latin-1 case you can also write it by hand in a few lines, since every byte maps to the rune of the same value. ## The case CharsetReader does not cover `CharsetReader` is consulted only when the document **declares** a non-UTF-8 encoding. Two gaps follow: 1. **No declaration at all.** A bare `<records>` root over Windows-1252 bytes never triggers the hook. The decoder assumes UTF-8, hits a byte sequence that is not valid UTF-8, and returns an `*xml.SyntaxError` complaining about invalid UTF-8 — at whatever offset the first accented character happens to sit, which on a large export can be minutes into the run. The fix is outside the decoder: wrap the reader in a transcoder before you call `xml.NewDecoder`. 2. **A lying declaration.** If the file says UTF-8 and is not, you get the same invalid-UTF-8 syntax error. Only out-of-band knowledge, or sniffing, resolves it. When you control neither the exporter nor the metadata, the durable answer on a reconciliation pipeline is to make the source encoding an explicit input to the job — a flag or a per-source configuration value — rather than trusting the declaration and discovering the truth in production. ## The other direction There is no symmetric hook on the encoder. `xml.Encoder` always writes UTF-8, and the `xml.Header` constant it is conventional to write first declares exactly that. So a decode-and-re-emit pass over a Latin-1 document produces a UTF-8 document. That is almost always what you want, but if a downstream consumer insists on the original code page you must transcode the encoder's output yourself and rewrite the declaration to match. ## Practical checklist - Compare the label case-insensitively; do not switch on an exact string. - Return a real error for unsupported labels instead of returning the raw reader unchanged — returning the input untouched is the sneaky failure, because it looks like it worked. - Remember the hook fires per document, not per element, and only after the declaration has been read. - If the source has no declaration, transcode before the decoder, not inside it.
- The file has no XML declaration at all but the bytes are Windows-1252. Does CharsetReader help?No. The hook is only consulted for an encoding the document declares, so it is never called. The decoder assumes UTF-8, meets an invalid byte sequence at the first accented character and returns a syntax error about invalid UTF-8. You have to wrap the reader in a transcoder before `xml.NewDecoder` sees it.
- What should CharsetReader return for a label you do not support?An error. It propagates out of the next `Token` call and stops the parse, which is the honest outcome. The tempting mistake is returning the input reader unchanged: the parse then appears to succeed and quietly mangles every non-ASCII character, which is far harder to notice than a failed run.
- Your pipeline reads a Latin-1 document and re-emits it with xml.Encoder. What encoding comes out?UTF-8. The encoder has no charset hook and always writes UTF-8, and the `xml.Header` constant declares that. Usually that is the improvement you want, but if a downstream consumer requires the original code page you must transcode the encoder's output yourself and emit a matching declaration.
saying these in an interview costs you the question
- Thinks encoding/xml sniffs or auto-detects the charset
- Sets CharsetReader but returns the input reader unchanged
- Expects the hook to fire when no encoding is declared
- Assumes the encoder honours the source document's encoding
- Compares the charset label with an exact case-sensitive match
- Reads Latin-1 as UTF-8 and calls the mangled output a Go bug