What does implementing encoding.TextMarshaler and TextUnmarshaler change about a type's encoding/json output?
answer
- one method in, one method out
- the encoder supplies the quotes
- the parsing side must be able to write back
- MarshalText returns us, not "us"
basics
~20 sencoding/json writes the type as one JSON string built from the bytes MarshalText returns, instead of encoding its fields or its underlying number. Decoding hands the unquoted bytes back to UnmarshalText. json adds the quotes and escaping itself.
solid answer
~50 sThe two interfaces are one method each: `MarshalText() ([]byte, error)` and `UnmarshalText(text []byte) error`. When a type implements them, `encoding/json` stops encoding its structure and writes the returned bytes as a single JSON string, quoting and escaping them for you — so `MarshalText` must return `us`, never `"us"`. On the way back, json strips the quotes, unescapes, and calls `UnmarshalText` with the raw bytes. `UnmarshalText` has to be declared on a pointer receiver, because it must store the parsed value in the caller's variable; `MarshalText` is normally on a value receiver so it works on values and pointers alike. The same pair is what makes a custom type usable as a JSON object key, and `encoding/xml` and `flag.TextVar` consult it too. There is a parallel pair, `encoding.BinaryMarshaler` and `BinaryUnmarshaler`, for non-human-readable forms; `encoding/gob` uses that one.
code
go · 26 lines// Region is a value type a shared package exports.
type Region uint8
const (
US Region = iota
EU
)
var regionNames = [...]string{US: "us", EU: "eu"}
func (r Region) MarshalText() ([]byte, error) {
if int(r) >= len(regionNames) {
return nil, fmt.Errorf("unknown region %d", uint8(r))
}
return []byte(regionNames[r]), nil // no quotes: json adds them
}
func (r *Region) UnmarshalText(text []byte) error {
for i, name := range regionNames {
if name == string(text) {
*r = Region(i)
return nil
}
}
return fmt.Errorf("unknown region %q", text)
}go deeper
Be ready to write both method signatures from memory and to say that the encoder, not your method, supplies the surrounding quotes. Know that the unmarshal side needs a pointer receiver.
Explain the round-trip contract, why the text slice must be copied if retained, and that an error from either method fails the entire Marshal or Unmarshal call rather than just one field.
Show that you know the blast radius: implementing the interface silently changes JSON values, JSON map keys, XML output and flag parsing at once, and any error path in MarshalText makes every enclosing type's encoding fallible.
Own the decision of whether a type should have a canonical text form at all, and treat the exact bytes as a published contract other teams will store, key on and grep for.
## The two interfaces The `encoding` package in the standard library is tiny: it declares four interfaces and nothing else. Two of them are the text pair. ```go type TextMarshaler interface { MarshalText() (text []byte, err error) } type TextUnmarshaler interface { UnmarshalText(text []byte) error } ``` They express a single idea: *this type has one canonical, human-readable text form, and that form is enough to reconstruct the value.* A type that satisfies both is saying its text form is lossless. ## What encoding/json does with them Without them, `encoding/json` encodes a type structurally — a struct becomes a JSON object of its exported fields, a named integer becomes a JSON number, a named string becomes a JSON string. With `MarshalText`, json stops looking at the structure. It calls the method, takes the `[]byte` you return, and writes it as **one JSON string**. Crucially, json supplies the surrounding double quotes and does the JSON escaping (of quotes, backslashes, control characters, and by default of `<`, `>` and `&`). Your job is to return the bare text. Returning bytes that already contain quotes gives you a doubly quoted string such as `"\"us\""`. Decoding mirrors this. json requires the incoming value to be a JSON string, unquotes and unescapes it, and passes the resulting bytes to `UnmarshalText`. If your method returns an error, the whole `json.Unmarshal` call fails with it. Likewise, an error out of `MarshalText` aborts the whole `json.Marshal`, not just that one field — so a marshaler that can fail makes every enclosing value's encoding fallible. ## Receivers The conventional pairing is a **value receiver** on `MarshalText` and a **pointer receiver** on `UnmarshalText`. `UnmarshalText` must be on a pointer because its entire purpose is to write the parsed value into the caller's variable; a value receiver gets a copy and the assignment is thrown away. `MarshalText` only reads, so a value receiver is right, and it also means both `T` and `*T` satisfy `TextMarshaler` — the method set of `*T` contains methods declared on `T`, but not the other way round. Declaring `MarshalText` on a pointer receiver is a common self-inflicted bug, because json will silently skip it whenever it is holding a value it cannot take the address of. ## The contract on the bytes Two rules travel with these interfaces: 1. `UnmarshalText` must be able to decode whatever `MarshalText` produced. That is the round-trip promise, and it is the thing worth testing with a table-driven test over your interesting values, including the zero value and the boundary values. 2. The `text` slice passed to `UnmarshalText` is **not yours to keep**. The caller may reuse or overwrite that memory after the call returns. If you want to retain the bytes, copy them — `string(text)` copies, `append([]byte(nil), text...)` copies, keeping the slice itself does not. `MarshalText` should likewise return a slice the caller can hold; returning a slice into a shared buffer you will overwrite is a data-corruption bug waiting for a busy encoder. ## Who else calls them This is the part that surprises people: implementing the interface changes the behaviour of packages you never mentioned. - `encoding/json` uses it for values **and** for map keys — a map whose key type implements `TextMarshaler` can be marshaled at all, and the produced text becomes the JSON object key. - `encoding/xml` uses it for element character data and attribute values. - `flag.TextVar` binds a command-line flag straight to any `encoding.TextUnmarshaler`, so your type becomes a flag type for free. - Plenty of third-party config and database layers probe for the same interfaces. Many standard-library types already implement them, which is why they behave the way they do: `time.Time` marshals as an RFC 3339 timestamp, `net.IP` and `netip.Addr` marshal as their familiar dotted or colon-separated text, `big.Int` as digits. ## The binary pair `encoding.BinaryMarshaler` (`MarshalBinary() ([]byte, error)`) and `encoding.BinaryUnmarshaler` (`UnmarshalBinary(data []byte) error`) are the same shape with the opposite goal: a compact, machine-oriented byte form that is not meant to be read by a human and does not have to be valid UTF-8 or valid JSON string content. `encoding/gob` uses this pair, and `time.Time` implements it with its own versioned binary layout that carries the monotonic-clock-stripped wall time and zone offset. Do not reach for the binary pair to feed JSON: JSON has no byte-string type, so a binary blob still has to be armoured before it can travel. Recent Go also added append-style companions, `encoding.TextAppender` and `encoding.BinaryAppender`, whose `AppendText`/`AppendBinary` methods write into a caller-supplied buffer so hot encoding paths can avoid an allocation per value.
- Why must UnmarshalText be declared on a pointer receiver?Because it has to store the parsed value in the caller's variable. A value receiver gets a copy, so the assignment vanishes when the method returns and the field is left at its zero value with no error reported. Only `*T`'s method set can mutate, and `*T` already includes the methods declared on `T`, so a pointer receiver costs nothing.
- When would you implement MarshalBinary instead of MarshalText?When the form is for machines, not people: a compact fixed-width encoding for `encoding/gob`, a cache entry, or an on-disk record. Binary output has no UTF-8 or printability requirement, so it can be smaller and cheaper. It is the wrong choice for JSON, config files and log lines, where a human has to read the value and JSON has no byte-string type.
- May UnmarshalText keep the slice it was handed?No. The caller may reuse or overwrite that memory after the call returns, so retaining the slice can silently corrupt your value later. Copy it — `string(text)` and `append([]byte(nil), text...)` both copy. This is the same rule that applies to every `[]byte` a decoder lends you, and it is easy to miss because a test that decodes one document never reuses the buffer.
MarshalText is the label you write on a box; json is the shipping clerk who puts the label in a plastic sleeve. Write the words, not the sleeve.
saying these in an interview costs you the question
- Returns bytes that already include the JSON quotes
- Declares UnmarshalText on a value receiver
- Thinks the interfaces only affect encoding/json
- Retains the text slice instead of copying it
- Expects MarshalText output to be encoded as a JSON object