skip to content

A CLI hashes each file with sha256.New and io.Copy into a content-addressed store, but some digests are wrong. How do you find the bug?

level: seniorimportance: should knowfreq 33%

answer

  1. the hash is not the suspect
  2. ask what bytes it was actually fed
  3. one call in that line cannot fail
  4. an ignored error still yields a digest
  5. compare against a value you did not compute

basics

~20 s

Wrong digests mean the hash was fed the wrong bytes, not that SHA-256 misbehaved. hash.Hash.Write never fails, so the only error io.Copy returns is the read error - drop it and you digest a prefix. Golden-vector tests pin it down.

solid answer

~50 s

Start from the fact that a hash is a total function: the same bytes always give the same digest, so a wrong digest is an **input-plumbing** bug. First reproduce it with a golden vector - hash a fixed fixture file whose SHA-256 you obtained from an independent implementation and assert the hex - because a test that only checks your code against itself passes happily on a truncated feed. Then audit every place bytes reach the hash: the dropped `io.Copy(h, f)` error, a reused `hash.Hash` that is never `Reset`, directory entries or symlinks walked into the hasher as empty input, and files being rewritten while the walk is in progress. The reason this survives review is that nothing errors: `hash.Hash.Write` is documented never to fail, and `Sum` always hands back a well-formed 32-byte value that encodes to plausible hex.

code

go · 10 lines
go
func digest(path string) ([]byte, error) {
	f, err := os.Open(path)
	if err != nil {
		return nil, err
	}
	defer f.Close()
	h := sha256.New()
	io.Copy(h, f) // error dropped: a short read is invisible
	return h.Sum(nil), nil
}

go deeper

for a junior

Take one habit away: never write io.Copy(h, f) without handling the returned error. A dropped read error still leaves you a perfectly valid-looking digest over half a file.

for a middle

Explain why the failure is silent - hash.Hash.Write is documented never to fail and Sum always returns Size() bytes - so nothing in the hashing path can signal that the input was short.

for a senior

Drive the diagnosis: reproduce against a fixture tree, assert digests produced by an independent implementation, and audit every entry point where bytes reach the hash before you go anywhere near the crypto.

for a principal

Decide what the store guarantees when a digest can be wrong. Whether the index is rebuilt, verified on read, or trusted forever is a durability call, and it sets how much verification the tool is allowed to pay for.

## The digest is never the suspect SHA-256 is deterministic and has no failure mode. If your indexer produces a digest that does not match the one a reference tool produces for the same file, then your program hashed **different bytes**. The investigation is entirely about the input path, and none of it is cryptography. ### Why the failure is silent Three properties of Go's hashing API combine to hide it: 1. `hash.Hash` embeds `io.Writer`, and its `Write` is documented **never to return an error**. So `io.Copy(h, f)` can only fail on the read side - and code that "knows the hash cannot fail" often drops the return values entirely. 2. `Sum(nil)` always returns `Size()` bytes. A digest over four kilobytes of a four-gigabyte file is exactly as well-formed as the correct one. 3. Hex encoding launders it further. Every wrong digest still looks like a digest in a log line, a filename, or a database column. So nothing panics, nothing logs, and the store fills with confidently wrong names. ### The usual causes, in the order worth checking **The dropped copy error.** `io.Copy(h, f)` returns `(int64, error)`. A mid-stream read failure - a network filesystem hiccup, a file removed under you, an unreadable block - stops the copy early with a non-nil error and leaves the hash holding the prefix. ```go h := sha256.New() io.Copy(h, f) // both return values discarded sum := h.Sum(nil) ``` **A hash reused without Reset.** `Sum` deliberately does not touch the hash state, so a single hasher shared across a walk accumulates every file it has ever seen. The first file's digest is right, which makes the bug look data-dependent. **Entries that are not files.** A directory walk that opens and hashes every entry gets zero bytes from a directory or from a symlink resolved to nothing, and the digest of the empty input is a real, valid SHA-256 value. Seeing the same digest repeated across unrelated entries is the tell. **The file changed under you.** A tree being written while it is walked yields a digest for a state that never fully existed. This one is not a code bug and needs a different answer - snapshot, lock, or re-verify. ### The diagnostic that actually settles it A **golden-vector test**: a fixed input with a digest you did not compute with this code. ```go func TestDigestGolden(t *testing.T) { dir := t.TempDir() path := filepath.Join(dir, "fixture.bin") if err := os.WriteFile(path, []byte("hello\n"), 0o600); err != nil { t.Fatal(err) } got, err := digest(path) if err != nil { t.Fatal(err) } if hex.EncodeToString(got) != wantHexFromReferenceTool { t.Fatalf("digest mismatch: %x", got) } } ``` The crucial word is *independent*. The two tests teams reach for first prove nothing here: - Hashing the same input twice and comparing - a truncated feed is perfectly deterministic and both runs agree. - Checking `len(digest) == 32` - always true, including for the empty input. A table of vectors covering an empty file, a one-byte file, a file larger than the read buffer, and a file larger than any in-memory limit catches the whole class. Add a fixture tree that includes a subdirectory and a symlink and the walk bugs fall out too. ### Making the failure loud instead Two cheap invariants, both per file: - **Return the copy error.** Always. `n, err := io.Copy(h, f)`; wrap it with the path using `fmt.Errorf("hash %s: %w", path, err)` so the operator knows which entry died. - **Check the byte count.** Compare `n` against `fi.Size()` from `f.Stat()`. It costs a syscall and turns a silent prefix digest into an explicit "read 4096 of 12345678 bytes" error. A mismatch that is not an error is still worth surfacing, because it means the file changed mid-read. A newcomer to the codebase should not have to know any of this. Hide it once behind a `digest(io.Reader) ([]byte, error)` or `digestFile(path string) ([]byte, error)` helper that handles the error, resets or constructs its own hash, and is covered by the golden vectors - then the walk code cannot get it wrong. ### How to talk about it Say the diagnosis in one line - "a hash cannot be wrong, so the input was wrong" - then name the ignored error, then reach for the vector test. That ordering is what distinguishes someone who has debugged this from someone reciting API facts.

  • The copy error is handled properly and one directory still produces wrong digests. What do you check next?
    The walk itself. Directory entries and dangling symlinks opened and hashed as empty input, unreadable files skipped inconsistently, one `hash.Hash` reused across entries without `Reset`, and files being written while the tree is walked. Reproduce it against a fixed fixture tree under `t.TempDir()` containing a subdirectory, a symlink and an empty file.
  • Why can't a test that hashes data and re-hashes it catch this?
    Self-consistency only proves your code agrees with itself. A truncated feed is entirely deterministic, so both runs produce the same wrong digest and the test passes. You need a vector whose expected digest came from an independent implementation, so the assertion is anchored outside the code under test.
  • What two cheap checks would you add so this fails loudly next time?
    Return `io.Copy`'s error wrapped with the file path, and compare the bytes actually copied against `f.Stat()`'s reported size. Together they cost one syscall per file and convert a silent prefix digest into an explicit error naming the entry and the byte counts, which is what an operator needs at 3am.
  • Where would you put the fix so a newcomer cannot reintroduce it?
    Behind one helper - `digest(io.Reader) ([]byte, error)` or `digestFile(path string)` - that owns the hash construction, the error handling and the byte-count check, covered by the golden vectors. The walk code then never touches a `hash.Hash` directly, so there is no place left to drop an error.

saying these in an interview costs you the question

  • Suspects a bug in crypto/sha256 rather than the input path
  • Checks the error from hash.Hash.Write instead of io.Copy's
  • Tests only that two runs produce the same digest
  • Treats a well-formed 32-byte digest as evidence the input was complete
  • Reuses one hash.Hash across files without calling Reset
  • Hashes directory entries and symlinks as if they were regular files