What makes a good seed corpus for a fuzzing campaign, and what does corpus minimisation preserve?
answer
- Where the search starts from
- Real traffic beats invented examples
- One per feature, not a thousand copies
- Distil the set against reached coverage
- Two meanings: shrink the set, shrink an input
basics
~20 sA good seed corpus is real, small, fast, scrubbed of sensitive data, and spans one input per distinct feature rather than many near-duplicates. Corpus minimisation keeps the smallest subset that preserves the coverage the whole set reached.
solid answer
~50 sSeeds are the starting points a campaign explores from, so they should be captured from real traffic rather than invented, cover each optional block, encoding and rare record type once instead of a thousand near-identical samples, be as small and fast as possible because execution cost scales with input size, and be scrubbed of credentials and personal data before they are committed or shared. Past failing inputs go back in permanently as both regression cases and known-interesting starting points. **Corpus minimisation**, or distillation, then selects the smallest subset that preserves the coverage the full set achieved — worth doing because the fuzzer cycles through the corpus, so a bloated set means every genuinely interesting entry is mutated far less often. Keep it distinct from **test-case reduction**, which shrinks one failing input while preserving the same failure.
go deeper
Be ready to say what a seed corpus is and why the campaign starts from real example inputs rather than empty bytes. Knowing that seeds should be small, valid and varied is enough at this level.
Explain the mechanics and the vocabulary trap: minimisation of the corpus preserves coverage, reduction of a single input preserves the failure. Expect to justify why one seed per feature beats a thousand near-duplicates.
Show that you have operated a corpus over time — persisting it between runs, re-distilling periodically rather than constantly, re-seeding after a format change, watching the coverage curve after a distil, and scrubbing captures before they leave the boundary.
Own the corpus as an organisational asset: where it lives, who may add to it, how it is versioned and reviewed given that every entry is code you will execute, and what its ongoing upkeep costs against the defects the campaign actually returns.
### The corpus is the fuzzer's memory A **seed corpus** is the set of starting inputs a campaign begins from, and the **working corpus** is the evolving set the fuzzer retains during the run because each entry reached something new. Together they are the most valuable artefact a fuzzing programme produces: a run that starts from a good corpus reaches interesting code in minutes, and a run that starts from an empty corpus may spend hours rediscovering the format's header. ### What makes a seed good - **Real, not synthetic.** Captured production inputs encode the shapes the format actually takes, including the vendor quirks nobody documented. In a hotel booking channel manager, that means real partner availability-and-rate messages, not hand-written examples of the specification. - **Diverse in features, not in volume.** One message per distinct feature beats a thousand near-identical ones: each optional block, each encoding, each rare record type, each locale-dependent numeric or date format. Coverage of *variants* is what buys reach. - **Small.** Fuzzers mutate what they are given, and the cost of an execution scales with input size. A 47 KB message and a 900 byte message may reach the same code; the small one is worth far more per second. - **Fast and valid.** A seed that already fails the schema check teaches the fuzzer nothing about the interior. - **Scrubbed.** Production captures carry credentials, personal data and customer identifiers. Corpora get committed, shared and shipped to shared infrastructure, so treat scrubbing as a precondition, not a cleanup step. - **Including past failures.** Every input that once caused a failure goes back in as a seed. It is both a regression check and a known-interesting starting point. ### Two different things are called minimisation This is the term to disambiguate out loud in an interview, because the two operations solve different problems. **Corpus minimisation (distillation)** operates on the *set*. It selects the smallest subset of the corpus that preserves the coverage the whole set achieved, discarding entries whose reach is already covered by others. A campaign in the channel manager that had accumulated 1,240 captured partner messages might distil to 388 files with no loss of reached edges. The benefit is direct: the fuzzer cycles through the corpus, so a corpus three times larger than necessary means each genuinely interesting entry is mutated a third as often, and the memory and start-up cost rise for nothing. **Test-case reduction (input minimisation)** operates on a *single input*, and usually on a single failing one. It repeatedly removes or simplifies parts of the input while checking that the same failure still occurs, until nothing more can be removed. This is what turns a 47 KB message that crashes the parser into a 31 byte reproducer a human can read and a developer can debug. Notice that the check being preserved here is "the same failure signature still happens" — it is a property of the *failure*, not of coverage. Both are automated and both should be routine. Confusing them in an interview is a common stumble; naming which one you mean is a cheap way to sound like you have run a campaign. ### Operating the corpus over time A corpus is a living asset and needs the same care as any other: - **Persist it between runs.** A campaign that starts cold every night throws away the search progress of every previous night. Store the working corpus and restore it at start-up. - **Version it, and treat additions as reviewable.** New entries arriving from a run are fine; new entries arriving from an unknown source are inputs you are about to execute. - **Re-distil periodically**, not on every run — distillation itself costs a full pass over the corpus. - **Re-seed after a format change.** When the message format gains a new record type, no existing seed exercises it, and the fuzzer has no way to invent a valid one from nothing. Add examples deliberately. - **Watch the coverage curve after distillation.** If reached edges drop after a distil, the distillation criterion was too coarse for the target. ### Why interviewers ask this Seeding is where fuzzing stops being "point the tool at it" and becomes engineering judgement. A candidate who says "I would collect real inputs, cut them down to one per feature, scrub them, distil the set against coverage, and keep every past crasher as a permanent seed" has clearly run something. A candidate who treats the corpus as a folder of random files, or who cannot say what minimisation is preserving, has not.
- Corpus minimisation and test-case reduction are both called minimisation. What does each one preserve?Corpus minimisation operates on the set and preserves *coverage*: it keeps the smallest subset of files that still reaches everything the whole corpus reached. Test-case reduction operates on a single failing input and preserves the *failure*: it strips and simplifies while the same signature still reproduces, until a human-readable reproducer remains. Different inputs, different objects, different preserved property.
- The message format gains a new record type. Why will the existing corpus not cover it?No seed contains an example of it, and mutation of unrelated bytes is unlikely to synthesise a valid new record type from nothing — especially behind a schema check. The coverage curve will simply flatten short of that code. The fix is deliberate re-seeding: add captured or hand-built examples exercising the new type, then let the search build from there.
- What is the risk in taking seeds straight from production capture?Real traffic carries credentials, personal data and customer identifiers, and corpora get committed to repositories, shared between teams and uploaded to shared fuzzing infrastructure. Scrubbing is a precondition, not a cleanup step. Seeds also arrive oversized, so trimming them down is part of the same intake pass.
saying these in an interview costs you the question
- Thinks more seed files always means a better campaign
- Cannot say what corpus minimisation actually preserves
- Confuses shrinking the corpus with shrinking one failing input
- Commits production captures without scrubbing sensitive fields
- Discards the working corpus between runs
- Never re-seeds after the input format changes