skip to content

Coverage & Test Quality

Measuring how much of the code the tests execute, and reasoning honestly about what that number does and does not say about correctness. Interviewers ask because coverage targets are common and widely misunderstood.

on this pageshow

questions

14

What is an approval test, and how does it differ from a test with hand-written assertions?

level: juniorimportance: must knowfreq 46%

answer

  1. Two files, not one typed expression
  2. The expectation is a stored artefact
  3. A human blesses the output once
  4. Received compared against approved
  5. A failing run leaves the received file

basics

~20 s

An approval test runs the code, writes the output to a received artefact, and compares it against a previously approved artefact. Nobody types the expected value: a human reviews the output once and approves it, and the test then guards that exact output.

solid answer

~50 s

An approval test replaces a hand-written expected value with a stored artefact. The test runs the code, writes what it produced to a **received** file, and compares that against an **approved** file committed alongside the test. Identical means pass; different means fail, and the failing run leaves the received file behind so a human can diff it. Making it pass again is a deliberate act: someone reads the difference and promotes received to approved. So the claim an approval test makes is *unchanged since a human approved it*, not *correct* — the oracle is that reviewer's one-time judgement. It earns its place where the output is too large or too detailed to hand-assert: a rendered receipt, a generated report, a wide serialised record. Where a rule is small and expressible, a hand-written assertion states intent better.

code

pseudocode · 13 lines
pseudocode
received = render_receipt(order_case_id)
write_file("receipt_" + order_case_id + ".received.txt", received)

approved_path = "receipt_" + order_case_id + ".approved.txt"
if not file_exists(approved_path):
    fail_test("no approved artefact yet; review the received file and approve it")

if received == read_file(approved_path):
    delete_file("receipt_" + order_case_id + ".received.txt")
    pass_test()
else:
    show_diff(approved_path, "receipt_" + order_case_id + ".received.txt")
    fail_test("output differs from the approved artefact")

go deeper

for a junior

Be ready to name the two artefacts — received and approved — and to say plainly that the test passes when they are identical. Mention that a human approves the output once instead of typing an expected value.

for a middle

Explain the run loop and the re-approval step: a failing run leaves the received artefact, someone diffs it, and promoting it to approved is a deliberate reviewed act. Be able to say what the assertion really claims.

for a senior

Show judgement about where the technique pays — wide rendered or generated output — and where a written assertion is stronger. An interviewer expects you to name determinism and review decay as the running costs.

for a principal

Own the framing that approved artefacts are a reviewed contract the team maintains, not a cache of last week's output. Be ready to say when you would introduce the technique and when you would refuse it.

### The problem it solves A conventional test states its expectation in code: you write the value you expect, run the code, and the framework compares the two. That works while the expectation is small enough to type and stable enough to be worth typing. It stops working when the output is a rendered receipt, a generated report, a forty-field serialised record, or the text a generator emits. Hand-writing those expectations is slow, the test becomes a wall of quoted text, and every intended change means editing the wall by hand. Approval testing inverts the order of the work. Instead of writing the expectation and then running the code, you run the code, look at what it produced, decide whether it is right, and **store that artefact as the expectation**. From then on the test's only claim is: today's output is identical to the artefact a human approved. ### The two artefacts Every approval test case owns a pair of files: - the **approved** artefact — committed next to the test, reviewed like source code, and the only expectation the test has; - the **received** artefact — written on every run, transient, and excluded from version control. The run loop is small: produce output, write it to received, compare received with approved. On a match the received file is discarded and the test passes. On a mismatch the test fails and *leaves the received file on disk*, usually launching a diff so the difference is visible immediately rather than as a truncated string comparison in a log. ### What the assertion actually says This is the point candidates most often miss. A green approval test says **"unchanged since approval"**. It does not say "correct". Two consequences follow. First, an approval test can never be stronger than the review that created it. If the reviewer approved a receipt with a wrong tax line, the suite now defends the wrong tax line and will go red the day someone fixes it. Second, it detects *any* change, including the ones you wanted. An intended behaviour change turns the test red by design; you re-run, read the diff, confirm it is the change you meant, and re-approve. That re-approval step is the discipline. Approving without reading converts the whole technique into an expensive way of recording whatever the code currently does. ### Where it earns its place Approval tests are strongest on output that is wide, structured and human-judgeable but not concisely expressible: - rendered documents — receipts, invoices, statements, printable summaries; - generated text — emitted source, configuration, templated messages; - serialised records with many fields, where a hand-written expectation would repeat the whole structure; - error and message catalogues, where the exact wording is the behaviour; - combinatorial output, where one artefact per input combination covers a wide surface for the cost of one review each. They are also the cheapest wide net you can throw over behaviour you are about to change: lock the current output, change the code, and every unintended difference shows up as a diff. ### A worked shape Consider the receipt renderer behind an online bookstore checkout. A **340-case regression pack** holds one order shape per case — single item, mixed physical and downloadable items, gift wrapping, a partially refunded order, a multi-currency edge. Hand-asserting each rendered receipt would be thousands of lines of quoted text. As approval tests, each case is a few lines that render the order and compare against its approved receipt. When a loyalty-discount line is added, 118 of the 340 artefacts change, and the review is a read of 118 diffs rather than an edit of 118 expectations. ### What it costs Three costs are real and are what interviewers probe. **Determinism** — anything that varies between runs (a clock reading, a generated identifier, an unspecified collection order) must be normalised or injected, or the test fails on unchanged code. **Granularity** — approving too much makes diffs unreviewable and couples many artefacts to one change; approving too little rebuilds hand-written assertions with less intent. **Review decay** — a team that re-approves reflexively still has green tests and no oracle at all. ### Relationship to assertion-based tests They complement rather than replace each other. A rule such as "the total equals the sum of the line amounts minus the discount" belongs in a written assertion: it states intent, survives an unrelated whitespace change, and explains itself when it fails. An approval test on the same receipt would go red for a changed column width and would never tell you *why* the output should look that way. Sensible suites carry both: written assertions for the invariants that matter, approval tests for the shape and detail nobody wants to type.

  • If nobody hand-wrote an expected value, what is the test's oracle?
    The reviewer at approval time. Their judgement is frozen into the approved artefact, and the test only detects departures from it. That means the test defends exactly as much correctness as that one review contained — approve output nobody read, and the test asserts nothing but self-consistency.
  • When would you prefer a hand-written assertion over an approval test?
    When the expectation is a short, expressible rule — a total, a status transition, a boundary. A written assertion states intent, fails with a message that explains the rule, and survives cosmetic changes to surrounding output. An approval test on the same behaviour goes red for a changed label and says only 'this differs'.
  • What happens the very first time an approval test runs, before any artefact exists?
    It fails by design. There is nothing to compare against, so the run produces a received artefact and stops. You read that output, decide whether it is the behaviour you want, and approve it. A framework that silently created the approved file on first run would approve unreviewed output, which is the failure mode the workflow exists to prevent.

It is the difference between describing a photograph in words and keeping a signed print in the file: the test only checks that today's print matches the one someone signed off.

saying these in an interview costs you the question

  • Says an approval test has no expected value at all
  • Treats promoting the received file as a formality
  • Claims a green approval test proves the output is correct
  • Thinks approval tests replace assertion-based unit tests
  • Commits the received artefact instead of the approved one
  • Assumes the framework should auto-approve on first run

context

open as a page

What is the difference between line coverage and branch coverage in a test run?

level: juniorimportance: must knowfreq 84%

basics

~20 s

Line coverage counts source lines executed at least once. Branch coverage counts which outcomes of each decision were taken, the true side and the false side. A line can execute while one of its two outcomes never runs.

open as a page

Can a test suite with no assertions reach 100% line coverage, and what does that prove?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Yes. Coverage instrumentation records which code the tests executed, not whether anything checked the result. A suite that calls every function and asserts nothing still reports 100%, and proves only that the code runs without crashing.

open as a page

What is mutation testing, and what does a surviving mutant tell you about a suite?

level: juniorimportance: must knowfreq 45%

basics

~20 s

Mutation testing seeds small deliberate faults into the code and reruns the tests. A fault the tests catch is a killed mutant; one no test notices survives, showing that code is executed but never actually asserted on.

open as a page

Why must an approval test normalise timestamps and generated ids before comparing output?

level: middleimportance: must knowfreq 41%

basics

~20 s

Because values that change every run make the received artefact differ from the approved one even when the code has not changed. The test then fails nondeterministically, so teams re-approve reflexively and the approval stops meaning anything.

open as a page

How does a coverage tool record which lines and branches actually executed?

level: middleimportance: should knowfreq 48%

basics

~20 s

A coverage tool inserts probes at the entry of each basic block, rewriting either the source before compilation or the compiled artefact at build or load time. Each run dumps the probe hits, which a report step maps back to source lines.

open as a page

Why gate a merge on new-code (diff) coverage instead of the whole repository's global percentage?

level: middleimportance: should knowfreq 54%

basics

~20 s

A global percentage is a ratio over the whole codebase, so one change is too small to move it and the gate cannot detect an untested change. Diff coverage measures only the changed lines, so the signal is proportional and the fix is local.

open as a page

What is a mutation operator, and why is each mutant a single small edit?

level: middleimportance: should knowfreq 34%

basics

~20 s

A mutation operator is a rule that rewrites one small piece of code - shifting a comparison boundary, negating a condition, swapping an arithmetic operator, replacing a return value. One operator at one site makes one mutant.

open as a page

How can an approval test's normalisation hide a real regression, such as a stale-cache read serving an out-of-date price?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Every normalisation removes information from the comparison. A pattern broad enough to tame a varying value also erases everything else it matches, so a wrong price behind a masked amount, or a wrong order behind a sort, no longer produces any difference to fail on.

open as a page

A team meets its 85% coverage target every sprint yet boundary defects keep shipping — how do you investigate?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Read the cases rather than the number: look for missing, tautological or collaborator-only assertions, check whether the coverage sits on risky code or on trivia, and remember that a boundary defect lives inside a fully covered line by construction.

open as a page

What granularity of output should an approval test approve, and how do you decide it?

level: principalimportance: should knowfreq 25%

basics

~20 s

Approve the smallest artefact a reviewer can judge in one sitting that still has one reason to change. Too coarse and one intended change forces hundreds of unreadable re-approvals; too fine and you have rebuilt hand-written assertions with less stated intent.

open as a page

A full mutation run takes six hours nightly — how do you make it pay for itself?

level: principalimportance: should knowfreq 28%

basics

~20 s

Split the two goods it produces. Run changed-lines-only analysis in review for per-change evidence in minutes, and keep a periodic incremental full run as the planning signal. Cut cost with per-mutant test selection, caching and parallelism.

open as a page

Why is full path coverage impractical, and what criteria are used instead?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

Path coverage requires exercising every route through a function, and routes multiply exponentially with sequential decisions and become unbounded once loops are involved. Suites use weaker criteria instead: decision coverage, or modified condition/decision coverage where each condition must independently flip the outcome.

open as a page

Why do equivalent mutants put a ceiling on any mutation score?

level: seniorimportance: nice to knowfreq 20%

basics

~20 s

An equivalent mutant behaves identically to the original for every possible input, so no test can ever kill it. Those mutants sit permanently among the survivors, which makes a score below 100% the best any suite can reach.

open as a page