skip to content

Non-JSON Wire Formats

Encoding that is not JSON: encoding/xml attributes and character data, encoding/csv quoting, encoding/binary byte order, gob, base64 and hex. Each has a trap that quietly corrupts data.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

Why does a []byte struct field marshal to a base64 string in encoding/json?

level: juniorimportance: must knowfreq 55%

answer

  1. JSON has no byte type
  2. bytes must survive as text
  3. four characters per three bytes
  4. standard alphabet, padded
  5. a [32]byte is not a slice

basics

~20 s

JSON has no type for raw bytes, so encoding/json marshals a []byte field as a base64 string using the standard padded alphabet. A nil []byte becomes null and an empty one becomes an empty string.

solid answer

~40 s

JSON's value types are object, array, string, number, boolean and null — none of them carry raw bytes. So `encoding/json` special-cases a slice of `byte`: it emits a JSON string holding the base64 of those bytes, using `base64.StdEncoding` (the RFC 4648 standard alphabet with `=` padding), and reverses that on unmarshal. Two details bite people. A `nil` []byte marshals to `null` while an empty non-nil one marshals to `""`. And a fixed-size array is not a slice to the encoder: a `[32]byte` digest marshals as a JSON array of 32 numbers, not a base64 string. The alphabet is not configurable; if you want hex or the URL-safe alphabet on the wire, use a `string` field or give the type its own `MarshalJSON`.

code

go · 8 lines
go
type Blob struct {
	Digest []byte
	Tag    [4]byte
}

b, _ := json.Marshal(Blob{Digest: []byte("hi"), Tag: [4]byte{1, 2, 3, 4}})
fmt.Println(string(b))
// {"Digest":"aGk=","Tag":[1,2,3,4]}

go deeper

for a junior

Be ready to state the rule out loud: a []byte field turns into a base64 string because JSON cannot carry raw bytes. Knowing that nil becomes null is the usual second beat.

for a middle

Explain the mechanics: which alphabet and padding are used, why an array of bytes behaves differently from a slice, and how a type's own MarshalJSON overrides the default.

for a senior

Show the operational angle — that clients cannot infer the encoding from the string, so the representation belongs in the API contract, and that nil versus empty changes what a consumer's decoder sees.

for a principal

Frame it as a published contract decision: once a field is base64 on the wire, the alphabet and the null-versus-empty behaviour are things partners depend on, and changing them later is a breaking change.

## The problem JSON leaves you with JSON defines exactly six value kinds: object, array, string, number, boolean and null. There is no byte-string type. Anything binary — a hash, a digest, an encrypted token, a small image — has to be *armoured* into characters before it can travel inside a JSON document. `encoding/json` makes that decision for you. When it meets a value whose type is a slice of `byte` (that is, `[]byte`, or any named type whose underlying type is `[]uint8`), it writes a JSON **string** containing the **base64** encoding of the bytes, and it uses `base64.StdEncoding`: the RFC 4648 standard alphabet `A–Z a–z 0–9 + /`, padded with `=` so the length is a multiple of four. ```go type Blob struct { Digest []byte } // json.Marshal(Blob{Digest: []byte("hi")}) -> {"Digest":"aGk="} ``` Unmarshaling is the mirror image: a JSON string being decoded into a `[]byte` field is base64-decoded, and a string that is not valid base64 produces an error rather than a garbage slice. ## Why base64 and not hex Base64 costs four output characters for every three input bytes — about 33% inflation — where hex costs two characters per byte, or 100%. Every character in the standard base64 alphabet is also a character JSON can put in a string without escaping. For a format whose job is to move opaque payloads compactly, base64 is the obvious default, and it is the same choice most JSON APIs make by hand. ## The edge cases worth memorising - **nil versus empty.** `var b []byte` is nil and marshals to `null`. `b := []byte{}` is non-nil and marshals to `""`. Round-tripping does not preserve the distinction in the direction you might hope: decoding `null` into a `[]byte` field leaves it nil, and decoding `""` gives you an empty non-nil slice. - **Arrays are not slices.** `[4]byte` and `[32]byte` are arrays. The encoder's byte-slice rule does not apply to them, so they marshal as a JSON array of numbers: `[1,2,3,4]`. This surprises people who store a `sha256.Sum256` result — which is a `[32]byte` — directly in a wire struct. Slice it (`sum[:]`) if you want the base64 string. - **Named types still count.** `type Token []byte` marshals as base64 too, because the encoder looks at the underlying type. - **`json.RawMessage` is the exception.** It is defined as a `[]byte`, but it declares `MarshalJSON`, so the encoder uses that instead of the byte-slice rule and splices the bytes in as literal JSON. - **Map values too.** A `map[string][]byte` gets base64 strings for its values by the same rule. ## Changing the representation The alphabet is fixed; there is no struct-tag option for URL-safe or unpadded base64, and none for hex. Three practical ways out: 1. Declare the field as a `string` and encode it yourself — `base64.RawURLEncoding.EncodeToString(b)` or `hex.EncodeToString(b)` — at the boundary where you build the wire struct. 2. Give the field a named type with its own `MarshalJSON`/`UnmarshalJSON`, or a text marshaler, so every use of the type agrees on the representation. 3. Keep the wire type separate from the domain type, so the domain keeps `[]byte` and only the wire struct carries the encoded string. Whichever you pick, write it down for consumers: a partner integrating against your API cannot tell from a 22-character string whether it is standard base64, URL-safe base64 or hex, and picking the wrong decoder is one of the most common integration bug reports on binary fields. ## What an interviewer is listening for That you know the rule (`[]byte` → base64 string), that you know *why* (JSON has no byte type), and that you do not confuse it with a byte array. The follow-up is usually the nil/empty distinction, because it decides whether a client sees `null` or `""` and therefore whether their own decoder blows up.

  • Does a [16]byte array field marshal the same way as a []byte field?
    No. The base64 rule applies to slices of byte only. An array marshals as a JSON array of numbers, so a `[16]byte` becomes `[12,34,...]` with sixteen entries. If you want the base64 string, slice it first with `arr[:]` when you build the wire struct.
  • Which base64 alphabet does encoding/json use, and can you switch it to the URL-safe one?
    It uses `base64.StdEncoding` — the standard alphabet with `+`, `/` and `=` padding. There is no tag or option to change it. To put a URL-safe value on the wire, declare a `string` field and call `base64.RawURLEncoding.EncodeToString` yourself, or give the type its own `MarshalJSON`.
  • Why doesn't json.RawMessage come out as base64, even though it is a []byte?
    Because `json.RawMessage` declares a `MarshalJSON` method that returns its bytes unchanged. A marshaler method takes precedence over the encoder's default rules, so the raw bytes are spliced into the output as literal JSON instead of being armoured into a string.
  • What happens when a client sends a string that is not valid base64 into a []byte field?
    `json.Unmarshal` returns an error rather than filling the field with partial data — the base64 decode failure is reported like any other type error for that field. Everything decoded before the failure may already be set, so treat the whole result as unusable on error.

JSON is a text-only envelope. Raw bytes have to be turned into letters before they can be posted in it — the same reason a binary attachment is encoded to text before it travels in an email body.

saying these in an interview costs you the question

  • Says JSON has a native byte-array type
  • Expects a []byte to appear as a list of numbers
  • Thinks nil and empty byte slices both encode as null
  • Assumes the URL-safe alphabet is used
  • Passes a [32]byte digest and expects a base64 string
open as a page

What do binary.BigEndian.Uint32 and binary.LittleEndian.Uint32 return for the bytes 00 00 01 00?

level: juniorimportance: must knowfreq 45%

basics

~10 s

binary.BigEndian.Uint32 reads the most significant byte first and returns 256. binary.LittleEndian.Uint32 reads the least significant byte first and returns 65536. Same four bytes, opposite orders, different numbers.

open as a page

Why must a gzip.Writer be closed before the file is valid, and what breaks if its Close error is dropped?

level: juniorimportance: must knowfreq 48%

basics

~20 s

A gzip.Writer buffers compressed data and writes the gzip trailer, a CRC32 plus the uncompressed length, only on Close. Skip Close and the file is truncated; ignore Close's error and a failed final write vanishes silently.

open as a page

In Go's encoding/csv, when would you use Reader.Read in a loop instead of Reader.ReadAll?

level: juniorimportance: must knowfreq 55%

basics

~20 s

ReadAll builds the whole file as one [][]string in memory and returns nil records if any row fails to parse. Read hands back one record per call and ends with io.EOF, so large or partly malformed files stay workable.

open as a page

What is Go's encoding/gob package, and how do you encode and decode a value with it?

level: juniorimportance: must knowfreq 45%

basics

~20 s

encoding/gob is Go's own binary serialization format. Wrap a writer with gob.NewEncoder and call Encode to send a value; wrap a reader with gob.NewDecoder and call Decode with a pointer to receive it. Only exported fields travel.

open as a page

How does Go's encoding/xml decide which struct field an XML element or attribute fills?

level: juniorimportance: must knowfreq 55%

basics

~20 s

encoding/xml looks only at exported fields. An element fills the field whose xml tag name, or field name if untagged, matches it; an attribute needs a tag ending in ,attr. Anything unmatched is dropped silently.

open as a page

What does each call to xml.Decoder.Token return while streaming a large XML file in Go?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Token returns the next piece of the document — an xml.StartElement, xml.EndElement, xml.CharData, xml.Comment, xml.ProcInst or xml.Directive — plus an error. At the end of the input it returns io.EOF, so you loop until then.

open as a page

How do you use xml.Decoder.DecodeElement to pull one repeated subtree out of a huge XML dump?

level: middleimportance: must knowfreq 48%

basics

~20 s

Walk the document with Decoder.Token; when a start element matches the one you want, call Decoder.DecodeElement(&v, &start) with that same start element. It reads the subtree through its matching end tag into v, so only one record is live at a time.

open as a page

When would you pick base64.RawURLEncoding over base64.StdEncoding in Go?

level: middleimportance: should knowfreq 48%

basics

~20 s

Pick base64.RawURLEncoding when the value goes into a URL path, query or filename: it swaps '+' and '/' for '-' and '_' and drops the '=' padding, so nothing needs percent-escaping. Use base64.StdEncoding for ordinary opaque payloads.

open as a page

Why does binary.Write refuse a struct containing a string or an int field?

level: middleimportance: should knowfreq 55%

basics

~20 s

binary.Write only encodes fixed-size data, whose byte width is known from the type alone. A string has no fixed width and int has an implementation-defined one, so Write returns an error and binary.Size reports -1 for such a value.

open as a page

How does gzip.Writer's Flush differ from its Close, and when is calling Flush worth the cost?

level: middleimportance: should knowfreq 34%

basics

~20 s

Flush pushes everything compressed so far to the underlying writer and ends the deflate block, so a reader can decompress it immediately; no trailer is written and the stream continues. Close finishes the member. Each flush costs ratio.

open as a page

In Go's encoding/csv, what does Reader.FieldsPerRecord do at 0, at a positive value, and at -1?

level: middleimportance: should knowfreq 42%

basics

~10 s

At 0 the reader takes the field count from the first record and requires every later record to match. A positive value demands exactly that many fields. A negative value switches the check off.

open as a page

Why must a Go csv.Writer be flushed with Flush and then checked with Error?

level: middleimportance: should knowfreq 48%

basics

~20 s

csv.Writer buffers records, so Write can return nil while the destination is already failing, and unflushed rows never reach the file at all. Flush pushes the buffer out but returns nothing, so Error is what reports the failure.

open as a page

What does a gob Encoder transmit the first time it encodes a given struct type?

level: middleimportance: should knowfreq 38%

basics

~20 s

A description of the type: its name, and each exported field's name and type, tagged with a numeric type id. Later values of that type on the same stream reference the id instead of repeating the description.

open as a page

What does the Space field of an xml.Name hold for a token returned by xml.Decoder.Token?

level: middleimportance: should knowfreq 38%

basics

~20 s

Space holds the resolved namespace URI, not the prefix written in the document. Token expands each prefix through the xmlns declarations in scope and discards it, leaving Local as the bare element name. Elements in no namespace have an empty Space.

open as a page

In an encoding/xml struct tag, what do the ,attr and ,chardata options and an a>b name mean?

level: middleimportance: should knowfreq 45%

basics

~20 s

In an encoding/xml tag, ,attr binds the field to an XML attribute rather than a child element, ,chardata binds it to the element's own text, and a name such as metadata>title descends through nested elements to the innermost one.

open as a page

A golden-file diff shows base64.NewEncoder output truncated at the tail. Why?

level: seniorimportance: should knowfreq 38%

basics

~10 s

The encoder was never closed. base64.NewEncoder returns an io.WriteCloser that buffers up to two leftover source bytes; Close is what emits that final group and its padding, so without it the tail never arrives.

open as a page

A telemetry decoder using binary.Read returns plausible but wrong numbers. How do you tell a byte-order bug from a struct-layout bug?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Hexdump the raw frame and compare it field by field with the specification. Byte-swapped values across every field point to the wrong binary.ByteOrder. A correct first field followed by garbage points to a width or padding mismatch in the Go struct.

open as a page

In Go, how do you cap how much data a gzip.Reader can produce from an untrusted .gz file?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Wrap the gzip.Reader, not the compressed input, in io.LimitReader with your budget plus one byte, then reject any stream that reaches that extra byte. io.LimitReader signals a plain io.EOF at its limit, so the caller must check the count.

open as a page

A partner CSV fails in Go with csv.ErrBareQuote. What does setting Reader.LazyQuotes change, and what does it cost?

level: seniorimportance: should knowfreq 36%

basics

~20 s

LazyQuotes makes csv.Reader accept a quote inside an unquoted field and an undoubled quote inside a quoted field instead of failing. The cost is the signal: a malformed file now parses into fields that may be split wrongly, silently.

open as a page

Why must gob.Register be called for a struct field whose type is an interface?

level: seniorimportance: should knowfreq 32%

basics

~20 s

The wire has to name the concrete type behind the interface, and the decoder has to turn that name back into a Go type it can allocate. gob.Register records the name-to-type mapping, and both programs must call it.

open as a page

xml.Unmarshal returns no error on a vendor's sample feed, yet one struct field is always empty. How do you find the cause?

level: seniorimportance: should knowfreq 40%

basics

~20 s

encoding/xml never errors on input it cannot place, so an empty field means nothing matched it. Check that the field is exported, that the tag name equals the element name, that attributes carry ,attr, and that a nested value needs a path tag.

open as a page

Why does re-emitting decoded tokens through xml.Encoder rewrite a document's namespace prefixes?

level: seniorimportance: should knowfreq 26%

basics

~20 s

Decoder.Token resolves each prefix to its namespace URI and discards the prefix, and xml.Encoder has no prefix table to restore it. It writes an element's namespace as a default xmlns declaration on that element and invents a prefix for namespaced attributes.

open as a page

Your Go log-shipping agent gzips rotated files before upload. How do you choose the compression level, and who can overrule you?

level: principalimportance: should knowfreq 30%

basics

~20 s

Choose the level from a benchmark over real log lines that records bytes out and CPU per gigabyte, not from a default. The CPU lands on the service owner's node and the saving on someone else's bill.

open as a page

For a [32]byte digest, how does hex.EncodeToString differ from fmt.Sprintf with %x?

level: middleimportance: nice to knowfreq 30%

basics

~10 s

Both produce the same 64 lowercase hex characters. encoding/hex takes a slice, so an array needs sum[:], and it ships a decoder with named errors; fmt reflects over any value and has no inverse.

open as a page

What do the two return values of binary.Uvarint mean when the byte count is zero or negative?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

binary.Uvarint returns the decoded uint64 and the number of bytes consumed. A positive count means success. Zero means the buffer held an incomplete varint and more bytes are needed. A negative count means the value overflowed 64 bits.

open as a page

In encoding/gob, why does decoding into a reused struct variable leave stale field values?

level: middleimportance: nice to knowfreq 22%

basics

~20 s

gob omits struct fields holding their type's zero value, and the decoder writes only the fields the stream contains. Fields absent from a record are left untouched, so a reused destination keeps the previous record's values.

open as a page

Why can encoding/xml not fill a map field, and what struct shape do you use instead?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

encoding/xml has no map support: xml.Marshal reports an unsupported type for a map, and Unmarshal has no rule for filling one. Decode repeated key/value elements into a slice of small structs, then build the map yourself.

open as a page

How do you make xml.Decoder read a document whose declaration says encoding="ISO-8859-1"?

level: middleimportance: nice to knowfreq 30%

basics

~10 s

Set Decoder.CharsetReader to a function that wraps the input reader and returns UTF-8. encoding/xml parses only UTF-8, so when a declaration names another encoding and CharsetReader is unset the decoder fails instead of guessing.

open as a page

What does gzip.Reader.Multistream(false) change when reading a file of concatenated gzip members?

level: seniorimportance: nice to knowfreq 22%

basics

~10 s

By default Go's gzip.Reader reads concatenated gzip members as one continuous stream. Multistream(false) makes Read stop with io.EOF at the end of each member, so you can handle members individually and advance with Reset.

open as a page

showing 1–30 of 32