Why do equivalent mutants put a ceiling on any mutation score?
answer
- Changed text, unchanged behaviour
- No input can tell them apart
- Detecting them is undecidable
- Permanent residents of the survivor list
- Judge against the achievable ceiling
basics
~20 sAn equivalent mutant behaves identically to the original for every possible input, so no test can ever kill it. Those mutants sit permanently among the survivors, which makes a score below 100% the best any suite can reach.
solid answer
~40 sEquivalence means the source changed but the observable behaviour did not — a redundant bound, an unreachable branch, a write nothing reads, an arithmetic edit that folds to the same result. No input distinguishes the mutant, so no test can fail against it, and it stays in the survivor list forever. Detecting equivalence automatically is undecidable in general, so tools only apply heuristics and the residue must be triaged by a human, at real cost per mutant. Practically this means never targeting 100%: judge the score against the achievable ceiling, or better, drop the absolute target and require that new survivors on changed code are triaged. Record each equivalence verdict with a written justification and revisit it when surrounding code changes, because equivalence is a property of context, not of the edit alone.
code
pseudocode · 8 lines// urgency() is defined to return a value in 0..100
urgency = computeUrgency(application)
// original
stored = min(urgency, 100)
// mutant: cannot be killed, urgency never exceeds 100
stored = min(urgency, 101)go deeper
It is enough to know that some seeded faults change nothing observable, so a perfect score is not a realistic goal and not every survivor is a missing test.
Be able to give a concrete example — a redundant bound or a value nothing reads — and to explain why such a mutant stays in the survivor list no matter how many tests you add.
Show you have triaged a real survivor list: name the undecidability behind automatic detection, describe suppression with justification, and explain why 'I cannot think of a killing test' differs from 'no test can kill it'.
Own how the ceiling shapes policy: why an absolute score target is the wrong instrument, how triage effort is budgeted, and how stale suppressions are audited so they do not quietly hide real weaknesses.
### The definition An **equivalent mutant** is a mutant whose behaviour is indistinguishable from the original program for every possible input. The source text changed; the observable behaviour did not. Because no input can make it behave differently, no test can ever make a test fail on it, and it therefore sits permanently in the survivor list. That is the ceiling. If a run generates 1,438 mutants and 92 of them are equivalent, the highest score any suite could ever reach is roughly 93.6% — and no amount of test writing will close the remaining gap, because the gap is not a defect in the suite. A team that treats 100% as the target will burn effort trying to kill unkillable things, or worse, will contort production code to remove them. ### Why they cannot be detected automatically Deciding whether two programs are behaviourally equivalent is undecidable in general — it reduces to the halting problem. There is no algorithm that takes a mutant and its original and always answers correctly. Tools can therefore only apply heuristics: compiler-style equivalence checks, constraint solving over the changed expression, detecting mutants whose changed code is provably unreachable, or refusing to generate operator/site combinations that are known to produce equivalents often. All of these reduce the population; none eliminates it. The residue is triaged by a human, and that triage is the technique's main hidden cost — reported effort in the literature runs to minutes per mutant, and studies typically find equivalents to be a small but non-trivial single-digit-percentage share of generated mutants. ### The shapes you will actually meet - **Redundant conditions.** A bound that is already guaranteed by an earlier check. Changing `<=` to `<` on a value that can never equal the bound changes nothing. - **Unreachable-in-practice branches.** Defensive code that a caller contract already excludes. - **Optimisation-equivalent arithmetic.** An edit that a compiler or the runtime would fold to the same result. - **Unobservable state.** A mutated field write that nothing ever reads, or a mutated value that is overwritten before use. - **Ordering that does not matter.** A change to the order of operations over a commutative combination, or to iteration order where the result is a set. ### A worked example In the grant-application review queue, applications are held in a priority order and a background sweep re-scores them nightly. One helper clamps a computed urgency score into a range before storing it. The score is produced by a function whose codomain is already inside that range, so a boundary mutant that changed the upper clamp from "at most 100" to "below 100" produced a mutant no input could distinguish. It survived the 6-hour nightly run, as it had survived every run for months. Contrast that with a survivor in the same run that *looked* equivalent and was not: a mutant in the reviewer-eligibility guard that removed a redundant-seeming re-check of the submitter identity. A reviewer following the ordinary path could never reach it with a self-submitted application, so the mutant appeared harmless — but a second entry point, the bulk-decision endpoint, reached the guard without the earlier check. The mutant modelled a real permission escalation. The lesson is the one that makes this a senior topic: "I can't think of a test that kills it" is not the same claim as "no test can kill it", and the first is what an engineer under time pressure actually means. ### Living with the ceiling The practical discipline is to stop treating the raw score as the target and start treating the survivor list as a work queue. 1. **Never set the goal to 100%.** Set it relative to the achievable ceiling, or better, drop the absolute target and require that new survivors introduced by a change are triaged. 2. **Record the verdict, don't re-derive it.** Once a survivor is judged equivalent, suppress it with a durable marker and a written justification, so the next run does not re-present it. An unjustified blanket suppression is how a real gap gets hidden, so the justification is the artefact that matters, not the suppression. 3. **Re-open suppressions when the code around them changes.** Equivalence is a property of the surrounding context, and the clamp that was equivalent yesterday stops being equivalent the moment the scoring function's range widens. 4. **Prefer trends to absolutes.** Equivalents are roughly stable for stable code, so movement in the score is meaningful even when its absolute value is not. 5. **Consider deleting the redundant code.** Many equivalent mutants sit on genuinely dead defensive checks. Removing them raises the ceiling and simplifies the code — though a check kept deliberately as a safety net is worth keeping and suppressing instead. ### Why interviewers ask it It is a differentiator rather than a gate. A candidate who knows only that "mutation testing kills mutants" will assert that a low score always means weak tests. A candidate who names equivalent mutants, explains the undecidability behind them, and describes triage as an ongoing cost has clearly run the technique on a real codebase rather than read about it.
- Why can no tool reliably decide whether a mutant is equivalent?Program equivalence is undecidable in general — it reduces to the halting problem — so no algorithm always answers correctly. Tools use heuristics: constraint solving over the changed expression, reachability analysis, compiler-style folding, and refusing operator/site combinations known to produce equivalents. These shrink the population; the remainder needs human judgement.
- How should a team record a survivor it has judged equivalent?With a durable suppression carried alongside the code plus a written justification saying why no input can distinguish it, so the next run does not re-present it. The justification is the artefact that matters — an unjustified blanket suppression is indistinguishable from hiding a real gap, and it is how a genuine weakness slips through review.
- When does a previously equivalent mutant stop being equivalent?When the surrounding context changes: the range of an upstream function widens, a new caller reaches a branch that was unreachable, or a field that nothing read acquires a reader. Equivalence is contextual, so suppressions must be revisited whenever nearby code moves, and a stale suppression can mask a real defect.
A locked door in a wall with no room behind it will never be opened by any key. Counting it among the doors you failed to open makes your key ring look worse than it is.
saying these in an interview costs you the question
- Assumes any surviving mutant means the tests are weak
- Targets a 100% mutation score as achievable
- Claims a tool can detect equivalent mutants reliably
- Suppresses survivors without recording a justification
- Treats an equivalence verdict as permanent regardless of code changes
- Confuses an equivalent mutant with an uncovered one