What does a gob Encoder transmit the first time it encodes a given struct type?
answer
- the stream explains itself
- sent once, then referenced
- type ids, not repeated field names
- the second value of a type is smaller
- one encoder, one stream, read from the front
basics
~20 sA description of the type: its name, and each exported field's name and type, tagged with a numeric type id. Later values of that type on the same stream reference the id instead of repeating the description.
solid answer
~50 sThe first value of a type carries a *type descriptor* ahead of it: the type's name and, for a struct, every exported field's name and type, bound to a numeric type id. Every later value of that same type on that same encoder is sent as data referencing the id, so the description is paid for once per type per stream. That is what "self-describing" means here — the decoder reconstructs the layout from the bytes, with no schema file shipped alongside. It also makes a gob stream *stateful*: it must be read from the front by one `gob.Decoder`, you cannot seek into the middle of a gob file, and you cannot concatenate the output of two encoders, because the second one re-defines the same type ids and the decoder rejects the duplicate. Note the decoder still needs a compiled Go type to decode *into*; the descriptors tell it how to match, not what to invent.
code
go · 11 linesvar buf bytes.Buffer
enc := gob.NewEncoder(&buf)
_ = enc.Encode(Entry{Key: "a"})
first := buf.Len()
_ = enc.Encode(Entry{Key: "b"})
second := buf.Len() - first
// second is much smaller than first: the type description for Entry
// was written once, ahead of the first value, and referenced by id after that.go deeper
Know the headline: a gob stream carries its own type descriptions, so there is no schema file to generate and ship alongside the data.
Explain the once-per-type-per-stream rule and the numeric type ids that follow it, and why that makes a gob stream stateful rather than a bag of independent records.
Show the operational consequence: a gob file must be read from its start by one decoder, so appending with a fresh encoder each time, or splicing two files together, produces bytes no decoder will accept.
Weigh what you are committing to when you pick this as a storage format: a stream that must be read whole from the front rules out random access and partial recovery, which is a real constraint on a shared on-disk cache.
## The prologue before the data When a `gob.Encoder` is asked to encode a value of a type it has not sent before on this stream, it first writes a description of that type. For a struct, that description is the type's name plus, for each exported field, the field's name and the id of its type. The description itself is bound to a numeric **type id**; the basic types (int, string, bool, float64, byte slices and so on) have predefined ids that both sides know without being told, and user-defined types get ids allocated as they are first encountered. After that, values of that type are sent as a type id followed by the field data. The second value of a type is therefore noticeably smaller on the wire than the first, and the hundredth costs nothing extra in description. Because a struct's fields may themselves be user-defined types, sending one type can pull in a small tree of descriptions — the compound type and everything it names, each defined once. ## What "self-describing" buys, and what it does not The payoff is that there is no separate schema artefact. Nothing has to be generated, versioned, shipped alongside the data, or kept in sync between two repositories: a gob file explains its own layout, and the decoder learns the sender's field names and types by reading them. What it does *not* buy is the ability to materialise types out of nothing. The decoder still needs a real Go type to decode into. The descriptors are used to *match* what arrived against what the receiving program declares — field name against field name — and to decide what to convert, what to ignore and what to leave alone. A stream of unknown shape cannot be turned into a usable Go value just because it describes itself. ## The stream is stateful, and that has consequences Type ids are stream-local state built up as the encoder goes. Three practical consequences fall out of that, and they are the reason this question gets asked at all: **You must read from the beginning.** A `gob.Decoder` accumulates the type table as it reads. Seek into the middle of a gob file and you land on data referencing ids you never saw defined; there is no way to resynchronise. So gob is a poor fit for a format you want random access into, and a reasonable one for a file you always read whole. **One encoder per stream.** Every new `gob.Encoder` starts its own type table and allocates user type ids from the same base. If a process appends to a cache file by opening it in append mode and constructing a fresh encoder for each record, the resulting file contains the same type id defined more than once. A single decoder reading it fails with a duplicate-type error as soon as it reaches the second definition. The fix is to keep one encoder alive for the life of the writes, or to frame each record independently — a length prefix, one decoder per record — and accept the repeated descriptions. **Concatenation is not composition.** For the same reason, two independently produced gob files cannot simply be `cat`-ed together and read as one stream, even though each is individually valid. ## Matching is by field name What the descriptor makes possible is name-based matching rather than position-based. The receiving struct does not have to be laid out in the same order, and does not have to be the same *named* type; what matters is that the field names line up. Fields present on the wire but absent from the destination are discarded, and fields present in the destination but absent from the stream are left alone. (Type names do become load-bearing in one place — a value sent through an interface-typed field, where the registered name of the concrete type is what the decoder looks up.) ## Reading it in practice A quick way to see the mechanism is to encode the same struct twice into a `bytes.Buffer` and compare how much the buffer grew each time. The first `Encode` pays for the description; the second pays only for the data. That single experiment explains most of gob's behaviour: why the format is compact for repeated records, why it is unsuited to random access, and why a stream cannot be spliced.
- A service appends to a cache file by opening it in append mode and creating a new gob.Encoder for each record. What breaks when one decoder reads it back?Each encoder starts a fresh type table and allocates user type ids from the same base, so the file ends up defining the same id more than once. A single `gob.Decoder` reads the first stream fine and then fails with a duplicate-type error. Either keep one encoder alive for the whole file, or frame each record yourself and give each one its own decoder.
- Must the receiving Go struct have the same type name as the encoded one?No. For plain struct values gob matches on field names, so the destination type may be named differently and declare its fields in another order; extra fields on either side are ignored or left alone. Type names do matter for a value sent through an interface-typed field, where the concrete type's registered name is what the decoder resolves.
- Can you seek to a record in the middle of a gob file?Not with gob alone. The decoder builds its type table as it reads, so data in the middle references ids it has not seen defined. If you need random access, put the framing outside gob — an index of offsets plus a self-contained stream per record — or use a format designed for it.
It is a shipment where the first crate contains the assembly diagram and every later crate just says "see diagram 65". Lose the front of the shipment and nothing after it can be assembled.
saying these in an interview costs you the question
- Thinks the type description is repeated with every value
- Says both sides must ship a generated schema file
- Assumes you can seek into the middle of a gob file
- Concatenates two encoders' output into one file
- Believes the decoder can build a type it does not have