skip to content

A weekend fuzzing campaign produced 4,117 failing inputs. How do you turn that into filed bugs?

level: seniorimportance: should knowfreq 44%

answer

  1. Crashers are a queue, not a count
  2. Fingerprint how it failed, not what went in
  3. Normalised top frames, harness frames dropped
  4. The heuristic both splits and merges wrongly
  5. Reduce, reproduce, rank, file, re-seed

basics

~20 s

Group the failing inputs by crash signature — failure type plus normalised top stack frames — reduce one representative per group to a minimal input, confirm it reproduces, rank severity apart from priority, and file each with a regression seed.

solid answer

~50 s

A pile of crashers is not a defect count: one hot bug is hit thousands of times from different mutations. First **deduplicate by crash signature** — the failure type plus a hash of the top few stack frames, normalised for addresses and stripped of harness and allocator frames. Expect the heuristic to over-split one bug reached by several paths and to over-merge two bugs failing inside a shared helper, so treat the group count as an estimate. Then **reduce** one representative per group until it is readable, and confirm it reproduces; a non-reproducing crasher usually means missing state — layout randomisation, a hash seed, a locale setting, the run's own seed — so record the environment with every finding. Finally rank technical **severity** apart from business **priority**, and file with the reduced input, the build settings, and that input added to the corpus as a permanent regression case.

go deeper

for a junior

Be ready to say that many failing inputs usually mean far fewer real defects, and that findings are grouped by how the program failed before anyone files anything. Knowing the pipeline exists is enough at this level.

for a middle

Explain the mechanics of a crash signature — failure type plus normalised top frames, harness and allocator frames excluded — and why the input is reduced before filing. Be able to name both failure modes of signature grouping.

for a senior

Show that you have run the pipeline: tuning frame depth after over-splitting, chasing a non-reproducing crasher back to layout randomisation or a locale setting, keeping severity and priority as separate rankings, and attaching a regression seed to every filed bug.

for a principal

Own the programme design: which surfaces carry a standing campaign, what gates a merge versus what runs out of band, how alerting avoids re-reporting known findings, and how a flat coverage curve is read as a signal to change the harness rather than to buy more hours.

### A pile of crashers is not a list of bugs A long campaign produces failing inputs, not defects. One real bug in a hot code path is hit again and again from different mutations, so the raw output of a weekend run on a hotel booking channel manager's rate-message parser might be 4,117 failing inputs standing for a couple of dozen actual problems. Triage is the pipeline that turns the first number into the second, and it is the part of fuzzing an interviewer can use to tell an operator from a tourist. ### Step 1 — deduplicate by crash signature A **crash signature** is a normalised fingerprint of *how* the program failed, computed so that two inputs hitting the same bug produce the same value. The usual ingredients: - the failure type (fatal fault, abort, assertion, timeout, detector report class); - the top *N* frames of the call stack, hashed after normalising away addresses, and after dropping frames that belong to the harness, the allocator, or the failure handler itself; - for an instrumented-build report, the reported location plus the site where the memory was allocated or freed, which separates two different lifetime bugs that fail at the same read. Both failure modes of this scheme should be stated out loud. **Over-splitting:** the same bug is reached through several call paths, so stack hashing yields a dozen signatures for one defect — mitigated by truncating to a small *N* and by normalising aggressively. **Over-merging:** two genuinely different bugs both fail inside one shared helper, so a shallow signature collapses them into one — mitigated by keeping a couple of frames of context and by re-checking after the bug is fixed to see whether the signature still appears. Deduplication is a heuristic; treat the signature count as a working estimate, not a truth. In the channel manager run, 4,117 inputs might collapse to 63 signatures, which after reduction and review turn out to be 11 distinct defects. ### Step 2 — reduce and confirm For each signature, take one representative input and reduce it: repeatedly strip and simplify while the same signature still reproduces, until you have something small enough to read. Then confirm reproducibility on a clean build. A crasher that does not reproduce is not a non-bug — it usually means the failure depends on state the input alone does not capture: address layout, hash-ordering randomisation, wall-clock or locale settings, an accumulated in-process state from the previous execution, or the run's own random seed. Record the seed and the environment with every finding, and re-run in the same configuration before declaring anything unreproducible. In this leaf's example, a rate field in a locale-dependent numeric format reproduced only on hosts whose default locale used a comma as the decimal separator — the input was fine, the environment was the missing variable. ### Step 3 — rank, and keep two rankings apart Two words get muddled here and an interviewer notices. **Severity** is the technical impact of the defect — what it corrupts, whether it terminates the process, whether it can be reached from untrusted input. **Priority** is business urgency — how soon it gets fixed, given who is exposed and what else is in flight. A hang in a parser reachable from a partner feed can be low severity and top priority; a memory fault deep in an internal tool can be the reverse. Whether a given crash also constitutes a security issue is a judgement owned by the security discipline, not by the fuzzing harness; the triage job is to route it there with a reproducer attached. ### Step 4 — file with everything the fixer needs A filed bug carries the reduced input, the signature, the exact build and instrumentation settings, the seed and environment needed to reproduce, and the reduced input added to the seed corpus as a permanent regression case. Without that last step the same defect is rediscovered next month. ### Continuous fuzzing as a regression asset Fuzzing pays off when it stops being a one-off exercise: - **Persist the working corpus** between runs so each night resumes the search instead of restarting it. - **Run a short, deterministic replay in the pipeline.** Every committed crasher is replayed on every change; this is fast, stable and is the part that belongs in a merge gate. - **Run the open-ended search out of band**, on a schedule or continuously, on its own hardware — an open-ended search has no fixed duration and must never gate a merge. - **Alert only on new signatures.** Re-reporting known findings is how teams learn to ignore the fuzzer. - **Track the coverage curve and the corpus size as health metrics.** A curve that has been flat for weeks means the campaign has converged and needs new seeds, a new harness, or structure knowledge — not more hours. Framed that way, the corpus and the committed crasher set become a durable regression asset that keeps paying long after the campaign that produced them.

  • Your signature scheme reports sixty-three groups but the fixes suggest only eleven defects. What went wrong and how do you tune it?
    Over-splitting: one bug is reached through many call paths, so hashing a deep stack yields a different fingerprint per path. Truncate to a smaller number of top frames, normalise addresses and inlined frames away, and drop frames belonging to the harness, the allocator and the failure handler. Then re-check after a fix lands — if several groups disappear together, they were one defect.
  • A crasher from the campaign does not reproduce on a colleague's machine. What do you investigate?
    State the input does not carry. Address-layout randomisation, a per-process hash seed, the run's own random seed, wall-clock or locale settings, accumulated in-process state from prior executions in the same worker, and the build's instrumentation flags all qualify. Re-run in the recorded configuration before calling it unreproducible — in one case a parsing failure appeared only where the host's default locale used a comma as the decimal separator.
  • How does fuzzing become a regression asset rather than a one-off exercise?
    Persist the working corpus between runs so each session resumes the search. Replay every committed crasher deterministically on each change — that part is fast and stable enough to gate a merge. Run the open-ended search out of band on its own schedule, since it has no fixed duration. Alert only on new signatures, and track the coverage curve and corpus size as health metrics.
  • Why is an open-ended fuzzing run a poor fit for a merge gate?
    It has no bounded duration and no deterministic outcome: the same code can pass tonight and fail tomorrow purely because the search wandered somewhere new. That makes it useless as a pass/fail signal on a change and guarantees it will be ignored or bypassed. Gate on the deterministic replay of known crashers; let the open-ended search run continuously and report findings as they arrive.

saying these in an interview costs you the question

  • Reports the raw crasher count as the number of defects
  • Hashes the whole stack, producing one signature per call path
  • Files unreduced multi-kilobyte inputs as reproducers
  • Calls a non-reproducing crasher a false positive immediately
  • Uses technical severity and business priority interchangeably
  • Puts an open-ended fuzzing run inside a merge gate

context