skip to content

Generated Test Cases

Techniques that machine-generate cases instead of enumerating them by hand, and the oracles that judge the results. Interviewers use them to tell a senior tester from a writer of examples.

on this pageshow

questions

13

What is fuzzing, and what counts as a failure when there is no expected output to compare against?

level: juniorimportance: must knowfreq 62%

answer

  1. Input nobody wrote by hand
  2. No expected value can exist
  3. The program itself supplies the verdict
  4. Fault, abort, hang, detector report
  5. Generic oracle, catastrophic failures only

basics

~20 s

Fuzzing drives a program with large volumes of generated, malformed or mutated input. Because no per-input expected result exists, the oracle is the program's own bad behaviour: crashes, hangs, assertion failures, and reports from an instrumented build.

solid answer

~50 s

Fuzzing is automated testing that feeds a target millions of generated inputs — random, mutated from real samples, or built from a format description — and watches for misbehaviour rather than for a specific answer. Nobody can write an expected result for an input nobody has seen, so fuzzing uses a **generic oracle** that must hold for every input: the process must not die on a fault, must not abort or trip an internal assertion, must not exceed a time or memory budget, and must not trigger a report from a memory-error-detecting build. A clean rejection of malformed input is correct behaviour, not a finding — that distinction keeps the noise down. The crash oracle is free but weak: it only catches behaviour bad enough that the runtime objects, so teams strengthen it with assertions, timeouts, resource caps or a comparison against a second implementation.

code

pseudocode · 14 lines
pseudocode
function fuzz_one(input_bytes):
    start = now()
    result = parse_rate_message(input_bytes)   # may abort, fault, or hang
    elapsed = now() - start

    if elapsed > 250ms:
        report_finding("timeout", input_bytes)

    if result.rejected:
        return                                  # clean rejection is correct

    assert result.nightly_rate >= 0
    assert result.nightly_rate < 1000000
    assert reencode(result) parses back to result

go deeper

for a junior

Be ready to define fuzzing in one sentence and to name the failure signals: abnormal termination, abort or assertion failure, a hang past a time budget, and a report from an instrumented build. Say explicitly that a clean rejection of bad input is not a finding.

for a middle

An interviewer at this level expects the mechanics of the oracle: why no per-input expected value can exist, how time and memory budgets are set rather than guessed, and how assertions or a differential comparison against a second implementation strengthen a weak crash oracle.

for a senior

Show production judgement about oracle strength. Explain which classes of defect the bare crash oracle silently misses in your own system, what you added to catch them, and how you keep the false-finding rate low enough that people still read the reports.

for a principal

Own the strategy question: where fuzzing belongs in a testing portfolio, which input-handling surfaces justify a standing campaign, and how you would argue the spend when a clean run is weak evidence of safety rather than a certificate of it.

### What fuzzing actually is **Fuzzing** is automated testing that drives the code under test with a large volume of *generated* input — random bytes, small mutations of real examples, or structures built from a description of the format — and watches for the program to misbehave, rather than checking each input against a particular expected answer. A campaign is measured in millions of executions, so nobody looks at the inputs individually. That single fact is what makes fuzzing a different discipline from example-based testing, and it is what the interview question is really probing. ### The oracle problem Every test needs an **oracle**: the thing that decides whether the observed behaviour is wrong. An example-based test carries its oracle inline — you wrote the input, so you also wrote the expected result. A fuzzer generates inputs nobody has ever seen, so a per-input expected value cannot exist. Fuzzing escapes this by using a **generic oracle**: a set of properties that must hold for *every* input, whatever it contains. The classic bundle is called the *crash oracle*, though "crash" is shorthand for a wider family of signals. ### The signals that make up the crash oracle - **Abnormal termination.** The process dies on a hardware fault, a fatal runtime error, or a deliberate abort. Distinguish this sharply from a *clean rejection*: a parser that returns a well-formed "malformed input" error is behaving correctly and must not be scored as a finding, or the campaign drowns in noise. - **Assertion and invariant failures compiled into the build.** The cheapest way to strengthen a weak oracle. An assertion that a decoded amount is non-negative turns a silent wrong value into a loud failure. - **Hangs and timeouts.** An input that takes wildly longer than normal indicates an unbounded loop, quadratic blow-up, or pathological backtracking. This needs an explicit budget, not a feeling. In a hotel booking channel manager whose rate-message parser has a 92nd-percentile parse time of 1.8 ms, a 250 ms per-input ceiling is generous and still catches the blow-up. - **Reports from an instrumented build.** A memory-error detector linked into the fuzz build reports out-of-bounds access, use of uninitialised memory, use-after-free and leaks at the moment they happen instead of whenever corruption eventually surfaces. In environments where the runtime already traps these, the equivalent lever is enabling every internal consistency check. - **Resource ceilings.** A per-input memory cap turns "allocates a gigabyte from a two-byte length field" into a reportable failure. - **Differential comparison.** Run two independent implementations of the same format on the same input and report any disagreement. This gives a strong oracle without anyone stating a property by hand. ### Oracle strength decides what you can find The bare "did the process die" oracle is free and universal, but it is weak: it only notices behaviour bad enough that the runtime itself objects. Everything wrong-but-survivable passes. Take the channel manager again: a mutated rate field arrives in a locale-dependent numeric format, the parser reads the separator according to the host's locale rather than the message's declared one, and a nightly rate of 1.234,50 is ingested as 1234.50 or as 1.23. No process dies, no memory is corrupted, and the crash oracle reports a clean run — while the system has just published a wrong price to a partner. The fix is to strengthen the oracle: add a range assertion, or compare against a second implementation, or re-encode the parsed value and compare it with the input. This is the honest contrast to draw in an interview. A **crash oracle** is generic, costs nothing to state, works on code you do not understand, and catches only catastrophic behaviour. A **stated property** must be written specifically for the code under test by someone who understands it, and in exchange it catches wrong answers, not just fatal ones. Real campaigns start with the free oracle because it is free, and strengthen it exactly where a wrong-but-alive result would be expensive. ### What fuzzing is bad at Fuzzing finds robustness failures on the input-handling surface. It does not know what the product is supposed to do, so it finds no missing features and no business-logic mistakes unless you encode them as assertions. It struggles with input that must satisfy a checksum or signature before anything interesting runs — the usual answer is to recompute or disable that check in the fuzz build. It struggles with multi-step protocol state unless the harness knows how to build a session. And a campaign that reports nothing is evidence about the oracle and the reach of the corpus at least as much as it is evidence about the code. ### How a finding travels A raw finding is not yet a bug. The output of a run is a pile of inputs that made something bad happen; those inputs still have to be deduplicated by signature, reduced to something a human can read, confirmed to reproduce, and only then filed. Treat "the fuzzer found 200 crashes" as an unprocessed queue, never as a defect count.

  • Why is a clean error return on malformed input not counted as a fuzzing finding?
    Rejecting bad input with a well-formed error is the parser doing its job. Almost every generated input is invalid, so scoring rejections as failures would bury the run in noise and hide the handful of real defects. The oracle must fire only on behaviour the program never promises: faults, aborts, hangs, resource blow-ups and detector reports.
  • A campaign runs for two days and reports nothing. What does that tell you?
    Less than it seems. It is evidence about the oracle, the reach of the corpus and the harness at least as much as about the code. Check whether the coverage curve is still rising, what share of generated inputs get past the format check, and whether the oracle is strong enough to notice anything short of a fatal fault. A silent run with a flat curve means the fuzzer explored nothing.
  • How would you catch a wrong-but-not-fatal result during a fuzz run?
    Strengthen the oracle beyond the crash signals: compile in assertions over the parsed values, cap time and memory per input, re-encode the parsed result and compare it against what was parsed, or run a second independent implementation of the same format and report any disagreement. Each of these turns a survivable wrong answer into a reportable failure.

It is like shaking a piece of furniture instead of measuring it: you cannot say what the correct wobble is, but anything that falls apart in your hands is definitely wrong.

saying these in an interview costs you the question

  • Claims every fuzzed input needs its own expected output
  • Counts a clean malformed-input error as a crash
  • Treats a silent campaign as proof the parser is safe
  • Ignores hangs because nothing terminated abnormally
  • Says fuzzing replaces example-based tests entirely
  • Thinks fuzzing finds business-logic mistakes on its own

context

open as a page

What is model-based testing, and where do the executable test cases come from?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Model-based testing builds an explicit model of intended behaviour, usually a state machine, and derives cases from it automatically. A generator walks paths through the model; each walk becomes a test, and the model predicts the expected result.

open as a page

What is a round-trip property in property-based testing, and what bug can it miss?

level: juniorimportance: must knowfreq 63%

basics

~20 s

A round-trip property asserts that transforming a value and then reversing the transformation returns the original value, for every generated input. It misses paired defects: when both directions are wrong in mirror-image ways, the trip still returns the input unchanged.

open as a page

How do dumb mutation, grammar-aware generation and coverage-guided fuzzing differ, and when is each right?

level: middleimportance: must knowfreq 55%

basics

~20 s

Dumb mutation flips bytes in existing inputs and knows nothing about the format. Grammar-aware generation builds well-formed inputs from a description of the format. Coverage-guided fuzzing uses an instrumented build to keep inputs that reach new code and mutate those further.

open as a page

How do you state a property for a function that has no inverse to round-trip against?

level: middleimportance: must knowfreq 56%

basics

~10 s

Assert a rule the output must satisfy rather than a specific value: an invariant or postcondition, idempotence, agreement with a simpler reference implementation, or a metamorphic relation between two runs on related inputs.

open as a page

What makes a good seed corpus for a fuzzing campaign, and what does corpus minimisation preserve?

level: middleimportance: should knowfreq 41%

basics

~20 s

A good seed corpus is real, small, fast, scrubbed of sensitive data, and spans one input per distinct feature rather than many near-duplicates. Corpus minimisation keeps the smallest subset that preserves the coverage the whole set reached.

open as a page

How do all-states and all-transitions criteria differ when a generator emits paths from a behaviour model?

level: middleimportance: should knowfreq 44%

basics

~20 s

All-states requires the generated paths to visit every state in the model at least once. All-transitions requires every legal action edge to be taken at least once. All-transitions subsumes all-states and typically costs several times more generated steps.

open as a page

A weekend fuzzing campaign produced 4,117 failing inputs. How do you turn that into filed bugs?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Group the failing inputs by crash signature — failure type plus normalised top stack frames — reduce one representative per group to a minimal input, confirm it reproduces, rank severity apart from priority, and file each with a regression seed.

open as a page

How does a behaviour model act as the oracle, and what does a passing generated run actually prove?

level: seniorimportance: should knowfreq 39%

basics

~20 s

At each step the model predicts the next state; the harness compares that prediction with what the system actually did. A pass proves only agreement with the model, so an incomplete or wrong model goes confidently green.

open as a page

A property test fails on a generated input in nightly CI. How do you turn it into a durable regression test?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Shrink the failure to the smallest input that still fails, decide whether the code or the property is wrong, fix it, and commit the minimal case as a named example test so the defect is caught deterministically instead of waiting for the generator to rediscover it.

open as a page

Why does a naive random generator explore almost none of a structured input space?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Because independently randomised fields almost never assemble into the interesting shapes: most values are rejected at the first validation check, or are small and unrelated, so the states that carry the defects - collisions, duplicates, boundaries - are never reached.

open as a page

When is the upkeep of a behaviour-model test suite no longer worth it for a small team?

level: principalimportance: should knowfreq 31%

basics

~20 s

Stop when the model has drifted into a stale second specification, when most generated failures are model or adapter noise rather than defects, or when one person alone can edit it. Judge by defects found per maintenance hour.

open as a page

How would you decide which parts of a regression pack should become property-based tests?

level: principalimportance: should knowfreq 40%

basics

~20 s

Convert behaviour that is expressible as a general rule with an independent, cheap oracle - codecs, normalisers, calculators, queue and ordering logic. Keep example tests where the expected value is a business decision, where a case is a communication artefact, or where each run is expensive.

open as a page