skip to content

Your CLI tests err == sql.ErrNoRows and now reports "lookup failed" for ids that do not exist. How do you diagnose and fix it?

level: seniorimportance: should knowfreq 54%

answer

  1. right message, wrong branch
  2. ask both questions at the failing line
  3. the sentinel is still down there somewhere
  4. no compiler or vet check for this
  5. the fake returned a pristine value

basics

~20 s

A layer between the query and the CLI now returns a new error carrying the original as its cause, so the outermost value is no longer sql.ErrNoRows and == silently stops matching. Fix it with errors.Is.

solid answer

~50 s

The symptom says a normal outcome is being reported as a hard failure, so I start at the branch that classifies it. `err == sql.ErrNoRows` only matches when the error reaches the CLI untouched; if any layer in between returns a new error carrying the original as its cause, the comparison is false and the empty-result path is skipped. I confirm it by printing the chain — walking `errors.Unwrap` or just checking whether `errors.Is(err, sql.ErrNoRows)` is true where `==` is false — which tells me immediately that the sentinel is present but no longer outermost. The fix is `errors.Is`. Then I write the test that would have caught it: wrap `sql.ErrNoRows` one layer deep and assert the classification still holds, because nothing in the compiler or `go vet` flags the `==` form and a test calling the repository directly passes either way.

code

go · 6 lines
go
func TestMissingRowIsReportedAsEmpty(t *testing.T) {
	err := fmt.Errorf("find user 42: %w", sql.ErrNoRows)
	if !errors.Is(err, sql.ErrNoRows) {
		t.Fatal("missing-row condition is no longer detectable")
	}
}

go deeper

for a junior

Recognise the signature: the printed message still names the real condition but the program took the failure branch. That means the classification test, not the query, is what went wrong.

for a middle

Explain mechanically why the comparison stopped matching once a layer returned a new error holding the original as its cause, and show the chain walk that proves the original is still there.

for a senior

Demonstrate the whole loop: diagnose from the symptom, fix with the matching helper, sweep the codebase for the same form, and add a test whose fixture is shaped the way production actually produces the error rather than a pristine sentinel.

for a principal

Own the systemic answer — a convention that classification always uses the matching helper, fakes return realistically shaped errors, and layers preserve causes — because the compiler will never catch this class and review alone has already failed once.

## The shape of the bug A small CLI looks up one record by id and prints it. The repository runs a single-row query and returns whatever comes back; the command layer decides how to report it: ```go rec, err := repo.Find(ctx, id) if err == sql.ErrNoRows { fmt.Println("no record with that id") return nil } if err != nil { return fmt.Errorf("lookup failed: %w", err) } ``` This worked. Then someone improved the repository so its errors say which query failed, returning a new error that carries the original as its cause. Nothing in the CLI changed, nothing failed to compile, no test went red — and now a missing id exits non-zero with `lookup failed: find user 42: sql: no rows in result set`. `sql.ErrNoRows` is a sentinel: one package-level value returned by `(*sql.Row).Scan` when a single-row query selected nothing. `==` asks whether the value in your hand *is* that value. Once a layer returns a new outer error, it is not — even though the sentinel is still reachable underneath. ## Diagnosing it The symptom points straight at classification, because the *message* is right and the *branch* is wrong: the text still ends in `sql: no rows in result set`, so the information survived; only the decision went astray. That combination — correct message, wrong branch — is the signature of a broken sentinel comparison, and it is worth recognising on sight. To confirm, ask both questions at the failing branch: ```go fmt.Println(err == sql.ErrNoRows) // false fmt.Println(errors.Is(err, sql.ErrNoRows)) // true ``` When those disagree, the diagnosis is complete: the sentinel is in the chain but not on top. If you want to see the shape, walk it with `errors.Unwrap` in a loop, printing `%T` and the message at each step, until it returns nil. `%+v` on the error will not show you the structure — only the concatenated message — which is why the loop is worth writing once. A second possibility is worth ruling out: the sentinel might not be in the chain at all, because a layer rebuilt the error from its text and dropped the cause. Then `errors.Is` is also false, and the fix belongs in that layer, not in the CLI. ## Fixing it ```go rec, err := repo.Find(ctx, id) if errors.Is(err, sql.ErrNoRows) { fmt.Println("no record with that id") return nil } ``` Two further judgments belong with the fix. **Sweep, do not spot-fix.** The same `==` form is almost certainly elsewhere in the codebase, and every instance is a latent version of this bug. Grep for comparisons against sentinel values and convert them all in one change, including the ones that currently work. **Consider whether the CLI should be matching a database sentinel at all.** `sql.ErrNoRows` leaking into the command layer means the repository's abstraction is thin: every caller now has to import `database/sql` to classify a normal outcome. A common improvement is for the repository to translate it into its own condition at its boundary. That is a design call about the package surface, but the diagnosis and the immediate fix stand on their own. ## The guard Why did the existing tests pass? Because a repository test calls the repository and gets its error directly; a command test with a fake repository returns whatever the fake was told to return, which was the bare sentinel. In both, nothing wrapped, so `==` and `errors.Is` agree. The test that catches this is the one that reproduces the *assembled* shape — the sentinel with at least one layer of context on it: ```go func TestMissingRowIsReportedAsEmpty(t *testing.T) { err := fmt.Errorf("find user 42: %w", sql.ErrNoRows) if !errors.Is(err, sql.ErrNoRows) { t.Fatal("missing-row condition is no longer detectable") } } ``` Better still, apply that to the real classification helper rather than to `errors.Is` itself, so the test exercises your code: give the fake repository an error that is wrapped one layer deep and assert the CLI exits zero with the empty-result message. The general rule this leaves behind: **fakes and fixtures should return errors shaped the way production produces them.** A fake that hands back a pristine sentinel is testing a call chain that does not exist. ## Why nothing caught it earlier - The compiler is happy: both operands are of interface type `error`. - `go vet` has no check for it — this is not a type error, it is a semantic one. - The race detector, profiles and traces are all irrelevant; nothing is slow or concurrent, the program simply decided wrongly. - Code review misses it because `err == sql.ErrNoRows` reads like the obvious thing, and it *was* correct when it was written. That leaves review discipline plus a test as the only defences, which is exactly why the convention is to write `errors.Is` unconditionally rather than deciding case by case whether a chain exists today.

  • Why did the existing unit tests stay green through this regression?
    They asserted on an error taken straight from the layer that produced it, or from a fake told to return the bare sentinel. Nothing wrapped it, so == and errors.Is agreed. The bug only exists in the assembled chain, so the test has to reproduce that chain — a sentinel with at least one layer of context on it.
  • If errors.Is(err, sql.ErrNoRows) is also false, what has gone wrong?
    Some layer rebuilt the error instead of keeping the original as its cause — typically by formatting the message into a fresh error. The chain is cut there, so nothing at the top can recover the condition. Fix the layer that dropped the cause; do not compensate by matching on message text.
  • Should the command layer be matching sql.ErrNoRows directly at all?
    It works, but it means every caller imports database/sql to recognise a normal outcome, and the storage choice has leaked upward. Translating the condition into the repository's own vocabulary at its boundary keeps the classification available without the coupling — at the cost of one more thing the package documents and must keep returning.
  • How do you find the other instances of this bug in a large codebase?
    Search for equality comparisons against error values — anything of the form err == and a name starting with Err, plus the common standard-library sentinels. Convert every one, including those that currently behave correctly, because whether a chain exists today is an accident of the call graph rather than a property of that line.

saying these in an interview costs you the question

  • Concluding the database or the driver is broken
  • Adding strings.Contains on the error message as the fix
  • Removing the added context so the == comparison works again
  • Fixing only the one reported line and not sweeping the codebase
  • Expecting go vet or the compiler to have caught it
  • Claiming the sentinel was lost when it was merely not outermost