skip to content

A golden-file diff shows base64.NewEncoder output truncated at the tail. Why?

level: seniorimportance: should knowfreq 38%

answer

  1. base64 consumes three bytes at a time
  2. one or two bytes can be left over
  3. who emits the '=' characters
  4. it is a WriteCloser for a reason
  5. a deferred Close hides the error

basics

~10 s

The encoder was never closed. base64.NewEncoder returns an io.WriteCloser that buffers up to two leftover source bytes; Close is what emits that final group and its padding, so without it the tail never arrives.

solid answer

~50 s

Base64 works in quanta of three source bytes to four output characters, so `base64.NewEncoder` holds back one or two bytes whenever the amount written so far is not a multiple of three. That residue is flushed only by `Close`, which also writes the `=` padding for a padded encoding. Code that writes and returns without calling `Close` produces output that is correct up to the last complete group and then simply stops — which is exactly what a golden-file diff shows as a truncated tail. Two things make it hard to spot: payloads whose length happens to be a multiple of three encode perfectly, so a fixture corpus can hide the bug for a long time; and `Close` returns an error that a bare `defer enc.Close()` throws away. Note that closing the encoder flushes it but does not close the underlying writer — you still close the file yourself, after the encoder.

code

go · 9 lines
go
func writeArmoured(w io.Writer, digest []byte) error {
	enc := base64.NewEncoder(base64.StdEncoding, w)
	if _, err := enc.Write(digest); err != nil {
		return err
	}
	// Close flushes the trailing 1-2 bytes and the padding.
	// It does not close w.
	return enc.Close()
}

go deeper

for a junior

Remember that base64.NewEncoder gives you an io.WriteCloser and that Close is not optional — it is what writes the last few characters.

for a middle

Explain the three-byte-to-four-character quantum, why one or two bytes must be buffered, and what Close does that Write cannot.

for a senior

Show the diagnosis: a valid-but-short tail, a golden-file or expected-length check that pins it, and the realisation that fixtures sized in multiples of three hid the bug.

for a principal

Argue for the design that removes the whole class: keep the encoder's lifetime inside one function that owns its Close and returns its error, and use EncodeToString when the payload does not need streaming.

## Why a base64 stream has residue at all Base64 maps every three input bytes onto four output characters, because three bytes are 24 bits and four base64 characters are 4 × 6 bits. A streaming encoder therefore cannot emit anything for a partial group: given one or two bytes it does not yet know whether the third byte is coming or whether the stream is over, and the answer changes the output (a final group of one byte encodes to two characters plus `==`, two bytes to three characters plus `=`). So `base64.NewEncoder(enc *base64.Encoding, w io.Writer) io.WriteCloser` buffers that residue internally. Everything up to the last complete three-byte group is written through to `w` as you go; between zero and two source bytes are held back at all times. `Close` is the signal that no more input is coming. It encodes the residue as a final short quantum, appends the padding characters if the encoding has padding, writes them to `w`, and returns any error from that write. ## The failure it produces ```go func writeArmoured(w io.Writer, digest []byte) error { enc := base64.NewEncoder(base64.StdEncoding, w) _, err := enc.Write(digest) return err // BUG: no Close, so up to four characters never appear } ``` The output is not corrupt in the middle and it is not empty. It is a valid prefix that stops early — the single most confusing shape a serialisation bug can take, because everything you look at first (the alphabet, the content, the first line of the file) is right. A golden-file diff is the diagnostic that pins it quickly: encode a fixed corpus, compare byte-for-byte against a stored expected output, and the diff points straight at a tail that is short by two to four characters. Comparing lengths alone is enough: `base64.StdEncoding.EncodedLen(len(src))` tells you exactly how many characters the file should hold. ## Why it survives testing When `len(src) % 3 == 0` there is no residue and the missing `Close` costs nothing. A test corpus of three-byte, six-byte and nine-byte fixtures will pass forever. The bug then shows up in production, on a payload of an awkward length, as a bug report from a partner whose decoder rejects the identifier — and their error message says "illegal base64 data" or "unexpected end of input", which points at the *decoder* and sends triage in the wrong direction. Reproducing with an input whose length is not a multiple of three is the fastest way to turn their report into your bug. ## Getting the shutdown order right Three rules cover almost every real case. 1. **Close the encoder, and close it before the thing underneath.** Deferred calls run last-in-first-out, so `defer f.Close()` followed by `defer enc.Close()` does the right thing: the encoder flushes into a file that is still open. Written the other way round, the flush lands on a closed file. 2. **Do not swallow the error.** `defer enc.Close()` discards the return value, and that return value is where a failed final write is reported. In code that must not lose data, call `Close` explicitly on the success path and check it, keeping a deferred close only as the cleanup path. 3. **Flush any buffering layer after the encoder.** If a `bufio.Writer` sits between the encoder and the file, the order is: close the encoder, then `Flush` the `bufio.Writer`, then close the file. Flushing first simply flushes bytes the encoder has not produced yet. A closely related trap: `Close` on the encoder does **not** close `w`. The encoder does not own the writer it was handed. If your handler writes armoured bytes into an HTTP response writer, closing the encoder is still mandatory to get the tail out, and there is nothing to close underneath. ## Making the class of bug impossible For small payloads, skip the streaming encoder entirely: `base64.StdEncoding.EncodeToString(b)` has no flush semantics to forget, and neither does `Encode` into a caller-sized buffer. Reach for `NewEncoder` when the input is genuinely a stream you do not want to hold in memory, and when you do, keep the encoder's lifetime inside one small function that owns its `Close` and returns its error. ## What an interviewer is listening for That you name the quantum (three bytes to four characters), that you know `Close` is a flush and not a close-of-the-underlying-writer, that you know the error must be checked, and that you can explain why the bug hides in tests. Candidates who reach straight for "the encoding is wrong" or "the file was truncated on disk" have not internalised that a `WriteCloser` with buffering has a mandatory finish step.

  • Does closing the encoder from base64.NewEncoder also close the underlying io.Writer?
    No. It flushes the buffered residue and writes the padding, then returns. The encoder does not own the writer it was handed, so a file or a response body still has to be closed by whoever opened it — after the encoder, not before.
  • How much output can go missing when Close is skipped?
    At most one four-character group. Up to two source bytes are held back, and with a padded encoding they would have produced three characters plus one `=`, or two characters plus `==`. With a Raw encoding it is two or three characters. Small enough to look like anything, which is why it hides.
  • Why did the truncation only reproduce for some payloads?
    Inputs whose length is a multiple of three leave no residue, so the missing Close costs nothing and the output is byte-identical to the golden file. Only lengths that are not multiples of three expose it — which is what to add to the fixture corpus first.
  • A bufio.Writer sits between the encoder and the file. What is the correct shutdown order?
    Close the base64 encoder first so its residue reaches the bufio.Writer, then Flush the bufio.Writer so those bytes reach the file, then close the file. Flushing before closing the encoder pushes out only what the encoder had already emitted.

saying these in an interview costs you the question

  • Assumes every Write emits all its bytes
  • Closes the file but never the encoder
  • Thinks closing the encoder closes the underlying writer
  • Flushes the bufio.Writer before closing the encoder
  • Ignores the error returned by Close
  • Blames the decoder because its error message arrived first