skip to content

questions

4

What is mutation testing, and what does a surviving mutant tell you about a suite?

level: juniorimportance: must knowfreq 45%

answer

  1. Break the code on purpose
  2. Many copies, one small edit each
  3. Killed versus survived
  4. A survivor names a missing assertion
  5. Measures checking, not executing

basics

~20 s

Mutation testing seeds small deliberate faults into the code and reruns the tests. A fault the tests catch is a killed mutant; one no test notices survives, showing that code is executed but never actually asserted on.

solid answer

~40 s

A mutation-testing runner makes many copies of the code under test, each with one small deliberate edit — a flipped comparison, a changed return value, a deleted call. Each copy is a **mutant**. It reruns the tests against each mutant: if a test fails, the mutant is **killed**; if every test still passes, the mutant **survived**. The **mutation score** is the share of mutants killed. A survivor is the valuable output, not the score: it is a concrete, reproducible statement that this exact change to this exact line caused no test to fail. That usually means a missing assertion, a missing input case, or an oracle too coarse to notice. It is a strictly stronger signal than structural coverage, which only records that a line ran, not that anything checked what it produced.

code

pseudocode · 11 lines
pseudocode
// original
if reviewerClearance >= application.requiredClearance:
    allow()

// mutant A: conditional boundary
if reviewerClearance > application.requiredClearance:
    allow()

// mutant B: conditional forcing
if true:
    allow()

go deeper

for a junior

Be ready to define the technique in two sentences and to say which outcome is the useful one. Remember that the tool edits the production code, not the tests, and that a survivor points at a weak or missing assertion.

for a middle

An interviewer expects you to explain the run mechanically: how mutants are generated, why only the tests touching a mutated line need to run, and how killed, survived, timed-out and no-coverage outcomes differ in what they tell you.

for a senior

Show you have read a real survivor list. Talk about triaging survivors into missing assertions, missing input cases and coarse oracles, and about which parts of a codebase justify the run cost.

for a principal

Own the argument for whether the signal is worth its price at all: where it beats a cheap coverage number, what it costs in machine time and human triage, and what it should and should not be allowed to block.

### The idea in one line Mutation testing asks a question that no execution-based measurement can answer: *if the code were wrong, would this suite say so?* It answers it empirically. A runner takes the code under test, produces many slightly-broken copies of it — each copy differing from the original by one small deliberate edit — and reruns the relevant tests against each copy. Each broken copy is called a **mutant**. If some test fails, the mutant is **killed**. If every test still passes, the mutant **survived**, and the suite has just demonstrated, on a concrete example, that it cannot tell correct code from incorrect code at that spot. The **mutation score** is the fraction of generated mutants that were killed. Unlike a structural coverage percentage, which counts what the tests *ran*, the score counts what the tests *would notice*. That difference is the whole point of the technique. ### Why a surviving mutant is the interesting output The score is a headline; the surviving mutants are the deliverable. Each survivor is a specific, reproducible statement of the form "this exact change to line N causes no test to fail." That narrows the diagnosis to one of four things: 1. **A missing assertion.** The test exercises the line and then asserts on something else — or on nothing. This is the classic case and the reason the technique exists. 2. **A missing case.** No test supplies input that makes the mutated expression behave differently from the original, so the fault is never *triggered* even though the line is executed. 3. **A weak oracle.** The test asserts on a coarse artefact — a status flag, a non-empty result — that the injected fault does not disturb, while the value the user actually cares about is unchecked. 4. **Dead or equivalent code.** The mutated behaviour is genuinely indistinguishable from the original, or the code path cannot matter. This last family is the technique's known limitation and it puts a ceiling on the achievable score. The first three are actionable defects in the suite. Triaging survivors is therefore a code review of your tests, driven by evidence rather than opinion. ### A worked example A grant-application review queue routes each submitted application to a reviewer and enforces one rule: a reviewer may not approve an application they themselves submitted or co-signed. The rule lives in a guard that compares the acting reviewer's clearance against the application's required clearance and checks the submitter identity. A nightly run over that service generated 1,438 mutants and killed 1,102 of them — a score of about 76.6%, against a line coverage figure in the high nineties. Among the survivors was a mutant that changed the clearance comparison from "at least" to "strictly greater than", and another that replaced the submitter-identity check with a constant that always passes. Both survived. The suite had many tests for the happy path — application routed, reviewer sees it, decision recorded — and one test for rejection, but nothing asserted that a self-submitted application was *refused*. In other words the suite could not have detected a permission escalation in which a reviewer approves their own application. Line coverage said the guard was covered; it was executed, never interrogated. Two new tests killed both mutants. ### How a run is structured A typical runner works in three stages. First it builds the mutants by applying **mutation operators** — small, mechanical, single-point edits — to the compiled or parsed form of the code. Second, it establishes which tests touch which parts of the code, so that for each mutant it can run only the tests that could plausibly observe it rather than the whole suite. Third, it executes those tests per mutant and records the outcome. Outcomes are usually classified as **killed** (a test failed), **survived** (all tests passed), **timed out** (the mutant made execution run far longer than the baseline — commonly an altered loop bound or exit condition — so the runner aborted it), and **no coverage** (no test executes the line at all, so nothing could have detected it). Timeouts are conventionally counted as killed, on the argument that a real change producing a hang would be caught in practice; treating them as survivors is defensible too, and either way a run full of timeouts is a signal to look at the operators being applied. "No coverage" mutants are not a test-quality finding at all — they are ordinary uncovered code, and it is more honest to report them separately than to bury them in the score. ### Where it sits relative to other signals Coverage is cheap, continuous and shallow; mutation score is expensive, discrete and deep. They are not substitutes and they answer different questions: coverage tells you what was reached, mutation score tells you what was checked. A suite can sit at 100% line coverage with a mutation score near zero if it has no assertions, and that fact is the single most quoted argument for the technique. The costs are real: the run multiplies test execution by the mutant count, the survivor list needs human triage, and some survivors will turn out to be unkillable by construction. Those costs shape how teams adopt it — usually incrementally on changed code rather than as a whole-repository gate.

  • How does a mutation score differ from a structural coverage percentage?
    Coverage counts what the tests executed; a mutation score counts what they would have noticed. A suite with no assertions at all can reach 100% line coverage and score near zero on mutants, because every seeded fault passes unremarked. The score is therefore evidence about assertions and oracles, while coverage is evidence about reach.
  • What does a runner do with a mutant that makes the code run forever?
    It applies a time budget derived from the baseline test run and aborts the mutant, recording it as **timed out**. This usually comes from an edit to a loop bound or exit condition. Timeouts are conventionally counted as killed, on the argument that such a change would be noticed in practice, though counting them separately is also defensible.
  • Why should a mutant with no covering test be reported separately from a survivor?
    If no test executes the mutated line, nothing could possibly have detected the fault, so the outcome says nothing about assertion strength — it is ordinary uncovered code. Folding it into the survivor list mixes two different findings and misdirects the fix: one needs a new test at all, the other needs a stronger assertion in a test that already exists.

A fire alarm that has never been tested might work. Setting off a small controlled smoke source in each room tells you which alarms actually sound — and a silent room is the finding, not the percentage.

saying these in an interview costs you the question

  • Says mutation testing changes the tests rather than the code
  • Claims a killed mutant means a bug was found in production code
  • Treats mutation score as just another coverage percentage
  • Believes 100% mutation score is always achievable
  • Ignores survivors and reports only the headline score

context

open as a page

What is a mutation operator, and why is each mutant a single small edit?

level: middleimportance: should knowfreq 34%

basics

~20 s

A mutation operator is a rule that rewrites one small piece of code - shifting a comparison boundary, negating a condition, swapping an arithmetic operator, replacing a return value. One operator at one site makes one mutant.

open as a page

A full mutation run takes six hours nightly — how do you make it pay for itself?

level: principalimportance: should knowfreq 28%

basics

~20 s

Split the two goods it produces. Run changed-lines-only analysis in review for per-change evidence in minutes, and keep a periodic incremental full run as the planning signal. Cut cost with per-mutant test selection, caching and parallelism.

open as a page

Why do equivalent mutants put a ceiling on any mutation score?

level: seniorimportance: nice to knowfreq 20%

basics

~20 s

An equivalent mutant behaves identically to the original for every possible input, so no test can ever kill it. Those mutants sit permanently among the survivors, which makes a score below 100% the best any suite can reach.

open as a page