skip to content

How do you make xml.Decoder read a document whose declaration says encoding="ISO-8859-1"?

level: middleimportance: nice to knowfreq 30%

answer

  1. the standard library parses one encoding
  2. there is a hook, not a table
  3. you hand back a reader
  4. unset means a refusal, not a guess
  5. no declaration means the hook never fires

basics

~10 s

Set Decoder.CharsetReader to a function that wraps the input reader and returns UTF-8. encoding/xml parses only UTF-8, so when a declaration names another encoding and CharsetReader is unset the decoder fails instead of guessing.

solid answer

~50 s

`encoding/xml` decodes UTF-8 and nothing else. When the XML declaration names a different encoding, the decoder calls `Decoder.CharsetReader`, whose signature is `func(charset string, input io.Reader) (io.Reader, error)`. You return a reader that transcodes the bytes to UTF-8; the decoder then parses that. If the field is left unset, parsing stops with an error saying the encoding was declared but `Decoder.CharsetReader` is nil — deliberate, because silently misreading a Latin-1 document as UTF-8 corrupts every accented character. Return an error from your function for labels you do not support, and it surfaces from `Token`. The `charset` package under `golang.org/x/net/html` and the encoding packages under `golang.org/x/text` provide ready converters. One catch: `CharsetReader` is only consulted for an encoding **declared** in the document. A file with no declaration whose bytes are actually Windows-1252 will instead fail with a syntax error about invalid UTF-8, and you have to wrap the reader yourself before `xml.NewDecoder`.

code

go · 10 lines
go
dec := xml.NewDecoder(f)
dec.CharsetReader = func(label string, in io.Reader) (io.Reader, error) {
	switch {
	case strings.EqualFold(label, "ISO-8859-1"), strings.EqualFold(label, "latin1"):
		return newLatin1Reader(in), nil // each byte becomes the rune of the same value
	case strings.EqualFold(label, "UTF-8"):
		return in, nil
	}
	return nil, fmt.Errorf("unsupported xml charset %q", label)
}

go deeper

for a junior

Know that Go's XML decoder handles UTF-8 only and will refuse rather than guess when a document declares something else. The field to set is Decoder.CharsetReader.

for a middle

State the hook's signature and what your function has to return: a reader yielding UTF-8. Explain why the library ships no charset tables and where the converters actually live.

for a senior

Distinguish the declared case from the undeclared one, since only the first reaches the hook, and make the source encoding an explicit input to a reconciliation job rather than trusting a supplier's declaration.

for a principal

Own the policy for a fleet of feeds: whether you normalise everything to UTF-8 at the boundary, who is accountable when a supplier's declaration lies, and what a run does when it meets a charset nobody configured.

## Why the decoder refuses Go source is UTF-8, Go strings are UTF-8 by convention, and `encoding/xml` parses UTF-8. XML in the wild is not so tidy: exports from older systems routinely arrive as ISO-8859-1 or Windows-1252, announced in the declaration: ```xml <?xml version="1.0" encoding="ISO-8859-1"?> ``` The standard library does not carry the world's character-set tables — those live in the `golang.org/x/text` repositories — so it cannot transcode by itself. What it does instead is give you a hook and, if you have not filled it in, fail loudly. The error reads along the lines of `xml: encoding "ISO-8859-1" declared but Decoder.CharsetReader is nil`. Failing is the right behaviour. Latin-1 and UTF-8 agree on the ASCII range, so a naive read of a Latin-1 document *mostly* works — until a name with an accent or a price with a currency symbol goes through, and then you get replacement characters or a parse error somewhere unrelated. Corruption that only affects a fraction of rows is far more expensive to discover than a refusal at byte zero. ## The hook ```go type Decoder struct { CharsetReader func(charset string, input io.Reader) (io.Reader, error) // … other fields } ``` The decoder calls it with the charset label exactly as the document spelled it — so compare case-insensitively, and be ready for the aliases (`latin1`, `ISO-8859-1`, `windows-1252`, `cp1252`). You return an `io.Reader` that yields the same content as UTF-8; the decoder replaces its own reader with yours and carries on. Returning an error rejects the document, and that error surfaces from the next `Token` call. For the actual conversion, the quasi-standard repositories cover it: the `charset` package under `golang.org/x/net/html` maps a label straight to a decoding reader, and the encoding packages under `golang.org/x/text` expose the individual code pages. For the single-byte Latin-1 case you can also write it by hand in a few lines, since every byte maps to the rune of the same value. ## The case CharsetReader does not cover `CharsetReader` is consulted only when the document **declares** a non-UTF-8 encoding. Two gaps follow: 1. **No declaration at all.** A bare `<records>` root over Windows-1252 bytes never triggers the hook. The decoder assumes UTF-8, hits a byte sequence that is not valid UTF-8, and returns an `*xml.SyntaxError` complaining about invalid UTF-8 — at whatever offset the first accented character happens to sit, which on a large export can be minutes into the run. The fix is outside the decoder: wrap the reader in a transcoder before you call `xml.NewDecoder`. 2. **A lying declaration.** If the file says UTF-8 and is not, you get the same invalid-UTF-8 syntax error. Only out-of-band knowledge, or sniffing, resolves it. When you control neither the exporter nor the metadata, the durable answer on a reconciliation pipeline is to make the source encoding an explicit input to the job — a flag or a per-source configuration value — rather than trusting the declaration and discovering the truth in production. ## The other direction There is no symmetric hook on the encoder. `xml.Encoder` always writes UTF-8, and the `xml.Header` constant it is conventional to write first declares exactly that. So a decode-and-re-emit pass over a Latin-1 document produces a UTF-8 document. That is almost always what you want, but if a downstream consumer insists on the original code page you must transcode the encoder's output yourself and rewrite the declaration to match. ## Practical checklist - Compare the label case-insensitively; do not switch on an exact string. - Return a real error for unsupported labels instead of returning the raw reader unchanged — returning the input untouched is the sneaky failure, because it looks like it worked. - Remember the hook fires per document, not per element, and only after the declaration has been read. - If the source has no declaration, transcode before the decoder, not inside it.

  • The file has no XML declaration at all but the bytes are Windows-1252. Does CharsetReader help?
    No. The hook is only consulted for an encoding the document declares, so it is never called. The decoder assumes UTF-8, meets an invalid byte sequence at the first accented character and returns a syntax error about invalid UTF-8. You have to wrap the reader in a transcoder before `xml.NewDecoder` sees it.
  • What should CharsetReader return for a label you do not support?
    An error. It propagates out of the next `Token` call and stops the parse, which is the honest outcome. The tempting mistake is returning the input reader unchanged: the parse then appears to succeed and quietly mangles every non-ASCII character, which is far harder to notice than a failed run.
  • Your pipeline reads a Latin-1 document and re-emits it with xml.Encoder. What encoding comes out?
    UTF-8. The encoder has no charset hook and always writes UTF-8, and the `xml.Header` constant declares that. Usually that is the improvement you want, but if a downstream consumer requires the original code page you must transcode the encoder's output yourself and emit a matching declaration.

saying these in an interview costs you the question

  • Thinks encoding/xml sniffs or auto-detects the charset
  • Sets CharsetReader but returns the input reader unchanged
  • Expects the hook to fire when no encoding is declared
  • Assumes the encoder honours the source document's encoding
  • Compares the charset label with an exact case-sensitive match
  • Reads Latin-1 as UTF-8 and calls the mangled output a Go bug