skip to content

How do you stop refactors from silently rewording a Go CLI's user-facing error text?

level: seniorimportance: should knowfreq 34%

answer

  1. messages are output, and output gets tested
  2. the wording should show up in review
  3. record the rendered line, diff it
  4. an update flag or nobody maintains it
  5. do not pin the half you did not write

basics

~10 s

Treat rendered error text as program output and pin it. A golden-file test compares err.Error() against a recorded string, with an -update flag to re-record, so any reworded message arrives as a reviewable diff.

solid answer

~50 s

Error prose is the part of a CLI most users actually read, and it is the part no compiler checks — a refactor that moves a wrap one function up rewrites what a user sees, invisibly. The fix is a golden-file test per failure path: construct the failure, take `err.Error()`, compare it against a file under `testdata`, and add a `-update` flag that rewrites those files. Now the wording lives in the repository and a reviewer sees the exact before-and-after line. Two disciplines make it work rather than annoy. First, keep the leading clauses stable, because that is the fragment a triager pastes into a search box; churn there is the expensive kind. Second, be careful what you pin: the tail often comes from the operating system or a dependency and varies by platform and version, so assert your own prefix and normalise or exclude the rest instead of freezing text you do not own.

code

go · 18 lines
go
var update = flag.Bool("update", false, "rewrite golden files")

func TestSyncErrorText(t *testing.T) {
	got := syncRemote("origin").Error()
	golden := filepath.Join("testdata", "sync_error.txt")
	if *update {
		if err := os.WriteFile(golden, []byte(got), 0o644); err != nil {
			t.Fatal(err)
		}
	}
	want, err := os.ReadFile(golden)
	if err != nil {
		t.Fatal(err)
	}
	if got != string(want) {
		t.Errorf("Error() = %q, want %q", got, want)
	}
}

go deeper

for a junior

Know that error text is program output and can be tested like any other output: build the failure, compare the rendered string, keep the expected text in the repository.

for a middle

Explain the mechanics of a golden-file test and the role of an update flag, and describe how you would produce a deterministic message for comparison.

for a senior

Show the judgement: which parts of the chain you own and may pin, which come from the platform and must be normalised, which failure paths are worth recording, and why the leading clause is the expensive thing to reword.

for a principal

Own error prose as a user-visible surface with a house style and a review gate, and be ready to argue how much test maintenance that surface justifies against the support load it removes.

## The problem Error messages are simultaneously the most-read output of a command-line tool and the least-tested. Nothing in the type system describes them. A perfectly reasonable refactor — extracting a helper, hoisting a wrap one level, renaming an operation in a prefix — changes the exact line a user sees, and no test fails. The wording drifts release by release, and the first person to notice is someone comparing two support tickets. ## Pin the rendered output The standard remedy in Go is a golden-file test. For each failure path worth caring about, drive the code to that failure, render the error, and compare against a recorded file: ```go var update = flag.Bool("update", false, "rewrite golden files") ``` The `-update` flag is what makes this pleasant rather than a chore: when a change to the wording is intended, you re-run with `-update`, and the new text appears in the diff. The reviewer then sees exactly what the user will now read, which is the whole point. Without the flag, engineers hand-edit golden files or, worse, delete the test. The test itself is ordinary: build the error, call `Error()`, read the file, compare, and report a mismatch with `%q` on both sides so an invisible difference — a stray space, a trailing newline that crept back in — is visible in the failure output. ## What is safe to pin, and what is not This is the judgement that separates a useful golden test from a flaky one. The rendered chain is a concatenation of clauses from several owners: - **Your own prefixes** — `sync origin`, `fetch refs`, `dial 10.0.0.5:9418`. You own these completely. Pin them. - **Text from another package or the operating system** — the tail, such as a syscall's message. You do not own it. It differs between platforms, and it can change when the runtime or a dependency is upgraded. Pinning it produces a test that fails on someone else's machine for reasons unrelated to the change. - **Volatile values** — timestamps, ports chosen at random, temporary directory paths, durations. These must be normalised before comparison or kept out of the message's stable part. So the practical shape is: assert the full text for failures you construct entirely yourself; for failures that end in foreign text, assert your prefix and normalise the tail, or inject a fixed sentinel cause in the test so the whole line becomes deterministic. ## Why the front of the message deserves special care The person triaging a bug report pastes a fragment into a search box. What they select is the beginning of the line — the distinctive operation words — because the tail is often generic (`connection refused` matches everything). That makes the leading clauses the de facto identifier of a failure mode across tickets, chat threads, and your own issue tracker. Rewording them is cheap for you and expensive for everyone who has an older ticket open. A golden test makes that cost visible at review time so it is a decision rather than an accident. ## Supporting practices - **A written house style.** Lowercase, verb-plus-object per clause, no trailing punctuation or newline, quoted interpolated values. A golden test locks in what the wording *is*; the style rule keeps new messages consistent with it. - **Cover the paths users hit, not all of them.** Golden files for the dozen failures that generate support load; ordinary assertions elsewhere. A golden file per error in the codebase is maintenance nobody will keep up. - **Keep the files readable.** One file per scenario, plain text, checked in. Reviewers should be able to read the file and hear the message. - **Include the empty and hostile cases.** The message produced when a user passes an empty name or one containing odd bytes is exactly the one likeliest to render wrong, and it is cheap to record. ## The trap to avoid A golden test is a check on *your own* output, run inside your own repository, where you may change the recorded text whenever you decide to. Do not let it grow into an assumption that the text is a stable contract for other code to consume — the test exists so that a human reviews wording changes, not so that anything can rely on the wording. Keeping the golden files scoped to your own package's prose, and normalising anything you did not write, keeps that line clear. ## In short Messages that users read are output. Record them, diff them, and let review decide when they change — while pinning only the parts you actually own.

  • Your golden file contains an operating-system error string at the end of the line. Why is that brittle?
    Because you do not own that text. Syscall wording differs between platforms and can change with a runtime or dependency upgrade, so the test fails for reasons unrelated to the change under review. Assert the prefix you produced, normalise the tail, or inject a fixed sentinel cause so the whole line is deterministic.
  • How do you keep a message useful to someone triaging a pasted bug report?
    Keep the leading clauses stable and distinctive, since that is the fragment people search on, and avoid putting volatile values such as timestamps or random ports at the front. Uniform verb-plus-object prefixes mean one search leads to one line of source rather than a hundred generic matches.
  • Why add a -update flag rather than editing golden files by hand?
    Because hand-editing is tedious enough that people disable the test instead. With a flag, re-recording is one command and the intended new wording lands in the diff, which is where you want a human to look at it.
  • Which failure paths deserve a golden file?
    The ones users actually hit and write tickets about — usually a dozen or so — plus the awkward inputs such as empty or oddly encoded values, where rendering is most likely to go wrong. Recording every error in the codebase creates maintenance nobody sustains.

saying these in an interview costs you the question

  • Treats error prose as untestable and relies on review alone
  • Pins an operating-system message and calls the test flaky when it breaks
  • Leaves timestamps or random ports in the compared text
  • Records golden files by hand so the test gets disabled
  • Rewords the leading clause of a common failure without noticing the cost
  • Writes a golden file for every error path in the codebase