When a Go fuzz run finds a crasher, how do you decide whether it gets committed under `testdata/fuzz`?
answer
- a corpus entry is forever, not free
- who pays for it, on every build
- a public repo changes the timing
- one entry per root cause, not per hash
- discovery budget is a separate call
basics
~20 sDefault yes, but with conditions. Every committed entry runs on every go test forever, and in a public repository it publishes a working trigger. So commit minimized entries, one per distinct root cause, and sequence publication after the fix ships.
solid answer
~50 sDefault yes, with conditions, because a `testdata/fuzz` entry is not just a file - it is a permanent cost and, in a public repository, a disclosure. Every entry is a seed that runs on every plain `go test` for every developer and every build from now on, so I want it minimized, one entry per distinct root cause rather than one per hash, and the total seed corpus held to a wall-clock budget somebody notices when it breaks. On timing: if the crash is a denial of service in a parser other people deploy, the input lands after the fixed release is out, not alongside the report. Volume that does not clear that bar stays in an archived corpus the fuzzing job restores, never in the repository. And the discovery budget is a separate decision - bounded scheduled runs on a dedicated machine, never nondeterministic fuzzing gating a pull request.
go deeper
The part to internalise is that a committed corpus entry is a real test that runs on every build, so adding one is a change with consequences, not just saving a file.
Be ready to explain both costs concretely: the per-build time each entry adds, and the fact that the entry is a readable, working input that reproduces the bug.
Show you can operate a policy - minimized entries, one per root cause, a measured wall-clock budget for the seed corpus, and discovery kept off the pull-request path.
This is your question: state the default, the conditions, who owns the disclosure timing and who owns the pipeline spend, and what evidence would make you change the policy.
### What committing an entry actually commits you to A file under `testdata/fuzz/FuzzParseFrame/` is a seed corpus entry, and seed entries run during **every ordinary `go test`**, not only during fuzzing. So "just commit it" is a decision to spend a small amount of everyone's time, on every build, indefinitely - and, because nobody ever dares delete a regression case, effectively permanently. One entry is free. Four hundred entries on a parser that takes a millisecond each is a visible tax on the inner loop, and it arrived without anyone deciding. There is a second, sharper cost when the repository is public. A minimized crasher for a binary frame parser is a working trigger for the bug, expressed in the most convenient possible form. If that parser sits in a proxy other people deploy, publishing the entry before a fixed release is out hands an input to anyone watching the repository, aimed at users who cannot yet upgrade. So the decision has two axes - recurring cost and disclosure timing - and both belong to someone. ### The policy I would argue for - **Commit, by default.** A crash that is not captured as a test comes back. The engineer fixing the bug adds the entry in the same change as the fix, so the test is red before the fix and green after. - **Commit the minimized entry, not the raw one.** Small entries are cheap to run, reviewable, and less useful as ready-made attack material. - **One entry per distinct root cause, not one per hash.** Fuzzing will happily produce a dozen inputs that all reach the same broken branch. Keep the smallest, archive the rest. - **Budget the seed corpus in wall-clock, not in file count.** Measure how long the seed entries take in a plain `go test`, set a ceiling, and make exceeding it a review conversation rather than something that silently accretes. - **Sequence publication in a public repository.** Fix, release, then publish the input - or publish it in the same change that ships the fixed release. For an internal repository this constraint mostly disappears, and the policy can be "always immediately". - **Everything else lives outside the repo.** An archived corpus that the fuzzing job restores gives you the volume without charging every clone for it. ### The discovery budget is a different decision How much wall clock the organisation spends looking for new inputs is not the same question, and it should not be answered by the same reflex. Discovery is nondeterministic and unbounded by nature, so it belongs on a schedule and on a dedicated machine, not on the pull-request path. A fuzzing run that gates a merge introduces failures that appear and disappear for reasons unrelated to the change under review, which is the fastest way to teach a team to ignore red builds. What runs per pull request is the committed seed corpus, under a plain `go test` - deterministic, fast and meaningful. The scheduled job then needs a persisted build cache so its search is cumulative, and a path that carries any new crasher out of the container as an artefact, otherwise the spend produces nothing durable. ### Who owns it, and who can overrule The repository maintainer owns the default policy: they carry the review load and the test-suite health. Two people can legitimately overrule them. A **release or security owner** can overrule on disclosure timing, because they own what a published input does to unpatched users; the entry waits, or lands with the release. A **platform owner** can overrule on spend, because a nightly fuzz job and a bloated seed corpus consume shared capacity; the answer there is a smaller discovery budget or a pruned corpus, not abandoning the practice. What should not happen is either overrule arriving as a one-off objection on a single pull request. If the cost is real, it changes the written policy - the corpus budget, the schedule, the disclosure rule - so the next crasher does not restart the argument. ### How you know the policy is working Three signals. The seed corpus runs inside its wall-clock budget. Every crash a scheduled run found in the last quarter is either a committed entry or a closed decision not to keep it. And no engineer has quietly stopped running fuzzing because the corpus made their test loop slow.
- Someone argues every crasher should be committed immediately, no exceptions. What is the strongest version of that argument?That any policy with a human judgment step leaks: entries get forgotten, the bug reappears, and the team relearns it the expensive way. Immediate commits are simple, auditable, and make the regression test exist before anyone is tempted to skip it. I would accept it outright for a private repository or an internal library; the exception I would keep is a parser other people deploy, where the input is a usable trigger before the fix ships.
- How do you keep the seed corpus from growing until the test suite is slow?Treat it as a budget rather than a folder. Measure how long the seed entries add to a plain `go test`, set a ceiling, and when a new entry would push past it somebody prunes instead of merging. Prune by root cause - keep the smallest entry covering a branch, archive the rest outside the repository where the fuzzing job can still restore them.
- Who should be able to overrule the maintainer on this?Whoever owns the cost being spent. A release or security owner can overrule on disclosure timing, because they own what a published input does to unpatched users. A platform owner can overrule on pipeline spend when a scheduled job or a bloated corpus eats shared capacity. The maintainer owns the default; an overrule should change the written policy, not just one file's fate.
- Should the nightly fuzzing job be moved onto every pull request instead?No. Discovery is nondeterministic and unbounded, so gating merges on it produces red builds unrelated to the change under review and trains people to ignore them. Pull requests run the committed seed corpus under a plain `go test` - fast and deterministic. Discovery stays on a schedule with a persisted cache and a path that exports any new crasher.
saying these in an interview costs you the question
- Treats committing a crasher as free and automatic
- Ignores that a public repository publishes a working trigger
- Lets the seed corpus grow with one entry per hash
- Puts nondeterministic fuzzing on the pull-request path
- Frames it as personal preference rather than owned policy
- Never revisits or prunes entries once merged