How does the seed corpus under `testdata/fuzz` differ from the fuzzing corpus Go keeps in the build cache?
answer
- two corpora, only one is yours
- one is committed, one is machine-local
- the cached one matters only while fuzzing
- go clean -fuzzcache touches only the discovered set
- testdata entries run on every plain go test
basics
~20 stestdata/fuzz is checked-in source: it travels with the repository and runs on every plain go test. The generated corpus lives in the build cache directory, is machine-local, is only used while fuzzing, and go clean -fuzzcache deletes it.
solid answer
~40 sGo fuzzing keeps two corpora. The **seed corpus** is what you wrote down: the values seeded in the target plus every file under `testdata/fuzz/<FuzzName>/`. It is ordinary source - committed, reviewed, shared with everyone who clones the repo, and executed as subtests by a plain `go test` with no `-fuzz` flag. The **generated corpus** is what the engine discovered: inputs that expanded code coverage during a `-fuzz` run, cached under the directory `go env GOCACHE` reports. It is per-machine, never committed, invisible to a normal test run, and exists only to make further fuzzing on that machine resume where it left off. `go clean -fuzzcache` removes it, which costs you fuzzing progress but never a test. The practical consequence: anything you want teammates or CI to keep has to be moved into `testdata/fuzz`.
code
text · 8 lines$ go env GOCACHE
/home/dev/.cache/go-build
$ ls proxy/testdata/fuzz/FuzzParseFrame
1f0a9c2e... 9b3d77aa...
$ go clean -fuzzcache # drops discovered inputs only
$ go test ./proxy # still replays both testdata entriesgo deeper
Remember there are two corpora: the one in your repository under testdata/fuzz, and one the toolchain caches on your own machine while fuzzing. Only the first is shared.
Be able to name where each lives, which one a plain go test executes, and what go clean -fuzzcache removes and leaves alone.
Show the operational consequence: the cached corpus is a machine's memory of its search, so it should be persisted where fuzzing runs continuously and never relied on for reproduction.
Frame it as a storage policy question - what belongs in the repository and is paid for on every build, versus what stays a restorable local artefact owned by the fuzzing infrastructure.
### Two corpora, two lifetimes Go's fuzzing has a seed corpus and a generated corpus, and almost every confusing thing about corpus management comes from mixing them up. **The seed corpus is yours.** It is made of two parts: the values you seed programmatically inside the fuzz target, and every file in `testdata/fuzz/<FuzzName>/` next to the test. Both are source. They live in the repository, go through review, and are identical on every machine that clones it. When `go test -fuzz=FuzzParseFrame` starts, the seed corpus is what it runs first and what it mutates from. When a plain `go test ./proxy` runs - no fuzzing flag at all - the seed corpus is *still* executed: the target runs once per entry, deterministically, as a subtest named after the file. **The generated corpus belongs to the machine.** While fuzzing, the engine keeps any input that reached new code coverage, because such inputs are useful stepping stones toward deeper states of the parser. These are written under the fuzz area of the Go build cache - the directory `go env GOCACHE` prints. They are not part of your module, are not seen by `git`, and are never executed by a normal `go test`. Their only job is to let the next `-fuzz` run on that same machine resume a search instead of starting from scratch. ### Why the split exists Coverage-guided fuzzing is cumulative. Getting a frame parser past a valid magic number, past a plausible length prefix, and into the branch that mishandles a truncated body can take a long search. The generated corpus is the memory of that search. But it is also large, noisy and uninteresting to humans: thousands of near-identical byte strings that mean nothing to a reviewer. Committing that would make every clone and every checkout heavier while adding no clarity. The seed corpus is the opposite: small, curated, human-meaningful. Each entry is either a case you chose or a crash the engine found and a human decided to keep. ### `go clean -fuzzcache` `go clean -fuzzcache` removes the cached inputs used for fuzzing. It is the right tool when the cache has grown too large, when you want to measure how a fuzz run performs from a cold start, or when the corpus is full of inputs for a target whose shape has completely changed. What it does **not** touch is anything under `testdata/fuzz`. Those are source files; `go clean` has no business deleting them. So the honest cost of `-fuzzcache` is fuzzing progress on that machine, not test coverage: your regression cases all still run, the next fuzz run just has to rediscover its way back in. This is also the trap the flag creates. If the only copy of an interesting input was in the cache, cleaning it - or simply moving to a different machine, or a fresh CI container - loses it permanently. Anything that matters must be promoted into `testdata/fuzz`. ### Practical rules that fall out of this - **Commit crashers, not the cache.** A failing input is written into `testdata/fuzz` for you; committing it is the deliberate act that makes it survive. - **Do not commit the generated corpus wholesale.** If you want a shared head start, hand-pick a handful of entries and promote them, or archive the cache directory as a build artefact and restore it on the machine that fuzzes. - **Persist the cache where fuzzing runs continuously.** A dedicated fuzzing box or a scheduled job benefits enormously from keeping its cache between runs; wiping it every night means every night starts cold. - **Expect nothing from the cache in CI.** A pull-request test job never reads it. Determinism in CI comes from the committed seed corpus, which is exactly why the seed corpus is the thing you invest in. ### Checking where things are `go env GOCACHE` tells you the build cache root that holds the generated corpus. `ls proxy/testdata/fuzz/FuzzParseFrame` shows the committed entries. Those two commands settle most arguments about "where did my corpus go" in a few seconds.
- What exactly is lost when someone runs `go clean -fuzzcache`?Only the coverage-expanding inputs the engine discovered on that machine. Nothing under `testdata/fuzz` is touched, so no test disappears and no committed regression case is lost. The next `-fuzz` run simply starts cold and has to work its way back into the interesting parts of the parser. On a long-lived fuzzing machine that can mean hours of lost search; on a laptop it is usually irrelevant.
- Should the generated corpus be committed to the repository?Normally no. It can be thousands of files, it is noise in review, and it grows without bound. If you want a shared starting point, promote a small hand-picked set into `testdata/fuzz`, or archive the cache directory as a build artefact and restore it on the fuzzing machine. Committing raw cache output makes every clone and every test run heavier for everyone else.
- Does CI need the generated corpus to reproduce a known failure?No. A saved entry in `testdata/fuzz` reproduces the failure deterministically under a plain `go test`, in milliseconds, with no fuzzing at all. The generated corpus only helps the engine find *new* failures. That is precisely why the fix for an unreproducible crash is to persist the testdata entry rather than the cache.
saying these in an interview costs you the question
- Thinks testdata/fuzz and the cached corpus are the same store
- Believes go clean -fuzzcache deletes committed corpus files
- Assumes the discovered corpus is shared across machines
- Expects a plain go test to run cache-only inputs
- Proposes committing the whole generated corpus