skip to content

An overnight `go test -fuzz` job crashed a frame parser, but nobody can reproduce it now. What went wrong?

level: seniorimportance: should knowfreq 30%

answer

  1. the artefact was in a directory that no longer exists
  2. two things died with the container
  3. a log line is not a corpus entry
  4. make the file leave the machine
  5. persist the cache, commit the entry

basics

~20 s

The failing input was written into the job's working tree under testdata/fuzz, and the coverage-guided corpus into that machine's build cache. A throwaway container destroys both, so only a log line survives and the crash is effectively gone.

solid answer

~50 s

Go writes a discovered crasher to `testdata/fuzz/<FuzzName>/<hash>` in the package's source directory, and keeps the coverage-guided corpus that led it there under the build cache directory `go env GOCACHE` reports. In an ephemeral CI container both die with the workspace, leaving a log line that cannot be replayed. The fix is to treat the corpus entry as the job's output: on failure, print the file and upload `testdata/fuzz/` as a build artefact, or push it on a branch so a human can commit it beside the fix. Persist the build cache between scheduled fuzz runs so the search resumes instead of starting cold every night, and never run `go clean -fuzzcache` on that machine. Finally, make sure the target is deterministic - a crasher that reproduces one run in ten is not a corpus problem, it is a second bug.

code

text · 5 lines
text
--- FAIL: FuzzParseFrame (0.02s)
    Failing input written to testdata/fuzz/FuzzParseFrame/1f0a9c2e...

$ cat proxy/testdata/fuzz/FuzzParseFrame/1f0a9c2e...   # last-resort backup in the log
$ tar czf crashers.tgz proxy/testdata/fuzz             # upload as a build artefact

go deeper

for a junior

The takeaway to hold on to is that the failing input is a file in the working directory, so if that directory is thrown away, the crash is gone with it.

for a middle

Be able to name both artefacts and where each lives - the entry in the package's testdata/fuzz and the coverage corpus in the build cache - and say why neither survives a container.

for a senior

This is the level being tested: diagnose it as a lost artefact rather than flakiness, and lay out the pipeline changes - export the entry, persist the cache, enforce a deterministic target - that make the next crash reproducible.

for a principal

Own the framing that a fuzzing job's deliverable is the corpus entry, and that paying for discovery without a path from crash to committed test is spending capacity for nothing.

### Where the evidence actually lived Two artefacts exist after a fuzzing run finds a failure in a binary frame parser, and neither of them is durable by default. The first is the **failing input**, minimized and written to `testdata/fuzz/FuzzParseFrame/<hash>` inside the package directory of whatever working copy the job checked out. In CI that working copy is inside a container that is deleted when the job ends. The second is the **generated corpus** - the inputs that expanded coverage during the search - held under the build cache directory reported by `go env GOCACHE`. In a fresh container that is a fresh empty cache, populated during the run and discarded with it. So the log said "failing input written to testdata/fuzz/FuzzParseFrame/1f0a9c2e...", and that path no longer exists anywhere. That is the whole diagnosis. It is not flakiness, not a compiler bug, not a mysterious heisenbug: the artefact was written to a disk that was thrown away. ### Why re-running is not a fix The instinct is to just fuzz again for another eight hours. That often fails, and understanding why is the senior part of this answer. Fuzzing is coverage-guided, and the guidance is exactly the cached corpus that was also discarded. A cold start has to rediscover the whole path from scratch: a valid header, a plausible length prefix, then the specific malformed combination that reaches the broken branch. Mutation is randomized, so there is no guarantee it lands in the same place, or in any acceptable time. You are re-rolling the dice, not re-running a test. ### Making the next crash survivable Four changes, roughly in order of value: 1. **Export the corpus entry.** On failure, the job should upload the package's `testdata/fuzz/` directory as a build artefact, and additionally print the file's contents into the log - it is short text, so a `cat` of it is a serviceable last-resort backup. Some teams go further and have the job open a branch or a pull request containing the new entry, which makes the artefact impossible to overlook. 2. **Persist the build cache across scheduled runs.** Restore and save the fuzzing machine's cache between runs so the search is cumulative rather than starting cold nightly. On a dedicated fuzzing host, simply do not wipe it, and do not put `go clean -fuzzcache` in any routine cleanup. 3. **Fix determinism in the target.** Replay only works if the target is a pure function of its arguments. State kept between calls, dependence on map iteration order, wall-clock or timing sensitivity, or shared globals all mean the saved entry may not fail on the next machine - and that undermines the entire corpus mechanism. 4. **Pin the toolchain for the job.** If the crash depends on behaviour that changed between Go releases, a reproduction attempt on a different toolchain version is a different experiment. Recording the version used, and pinning it for the fuzzing job, removes one variable from an already awkward investigation. ### What the recovery looks like once it is set up The next morning you download the artefact, drop the entry into `proxy/testdata/fuzz/FuzzParseFrame/` on a branch, and run: `go test -run='FuzzParseFrame/1f0a9c2e' ./proxy` No `-fuzz`, no randomness, no time budget - a deterministic red test in under a second. You fix the parser, watch that same command turn green, and land the entry together with the fix so the case runs on every build forever after. ### The reframing worth saying out loud A fuzzing job's product is not its exit status; it is the corpus entry. A pipeline that reports a fuzz failure but does not carry the input out of the container has not actually found a bug - it has found and then thrown away a bug. Treating the entry as the deliverable, with the same care you would give a build artefact, is what turns an expensive nightly job into something that improves the codebase.

  • The nightly job now archives testdata/fuzz. What do you do with the artefact the next morning?
    Download it, drop the entry into `proxy/testdata/fuzz/FuzzParseFrame/` on a branch, and confirm a plain `go test ./proxy` fails on it with no `-fuzz` flag - that gives a deterministic red test in seconds. Then fix the parser and land the entry with the fix in the same change, so the input becomes a permanent regression case.
  • Why does re-running the fuzzer for another eight hours often fail to find the same crash?
    Because fuzzing is coverage-guided and the guidance lives in the cached corpus that was discarded. A cold start must rediscover the whole path - a valid header, a plausible length prefix, then the specific malformed combination - and mutation is randomized, so there is no guarantee it reaches the same place at all. Persisting the cache is what makes progress cumulative.
  • Is persisting the Go build cache between CI runs risky?
    For a fuzzing job it is the point. The build cache is content-addressed, so stale entries are simply unused rather than wrong. The real concerns are size, which needs periodic trimming, and sharing: keep the fuzzing cache on the fuzzing machine or on a key dedicated to that job rather than mixing it into every ordinary test runner.
  • The saved entry reproduces only sometimes on a developer's laptop. Where do you look?
    At the target, not the corpus. Look for state kept between invocations, package-level variables, dependence on map iteration order, timing or wall-clock use, or anything touching the filesystem or network. A fuzz target has to be a pure function of its arguments; until it is, no saved input is a trustworthy regression test.

saying these in an interview costs you the question

  • Calls it flaky before asking where the input was written
  • Assumes the crasher was safely in the repository
  • Expects another fuzzing run to rediscover the input
  • Treats the failure log as sufficient reproduction
  • Wipes the fuzzing machine's cache as routine cleanup