What granularity of output should an approval test approve, and how do you decide it?
answer
- The decision sets the cost of a re-approval
- One artefact, one reason to change
- A reviewer must judge it in one sitting
- Approve a projection you designed
- Count re-approvals per intended change
basics
~20 sApprove the smallest artefact a reviewer can judge in one sitting that still has one reason to change. Too coarse and one intended change forces hundreds of unreadable re-approvals; too fine and you have rebuilt hand-written assertions with less stated intent.
solid answer
~50 sGranularity is the main design decision in approval testing, because it sets what a re-approval costs and therefore whether reviews stay real. Two failure directions bound it. Approve too much — a whole rendered page, chrome and all — and every artefact is coupled to every change, diffs stop being readable, and re-approval degrades into a rubber stamp. Approve too little — one artefact per field — and you have hand-written assertions with worse readability and none of their stated intent. Between them, the working rules are: an artefact should have **one reason to change**; it should be judgeable by a reviewer in one screen; and it should carry the payload under test, not the surrounding chrome, which usually means approving a deliberate projection of the output rather than the raw form. The health metric is blast radius — how many artefacts one intended behaviour change forces you to re-approve.
go deeper
Recall the two directions of failure: an artefact so large nobody reads its diff, and an artefact so small it is just a written assertion with extra ceremony. Aim for something a person can judge.
Be ready to explain coupling concretely: why an artefact containing unrelated content changes for unrelated reasons, and why approving a narrower projection of the output reduces both diff noise and normalisation.
Show that you choose depth by risk, keep written assertions on the invariants that must not drift, and can describe splitting an over-coarse artefact along its reasons to change.
Own the economics. Manage re-approvals per intended change as a metric, set ownership and a deletion policy for approved artefacts, and be able to argue where in the suite the technique should not be used at all.
### Why granularity is the decision that matters Everything else in approval testing is mechanics. Granularity is the design choice, because it fixes two quantities a lead actually has to manage: how much human attention each intended change costs, and how much of the system a single artefact is speaking for. Get it wrong in either direction and the discipline degrades — in one direction into rubber-stamping, in the other into an expensive re-implementation of ordinary assertions. ### The coarse failure Approving the whole output of a wide operation — a complete rendered page including navigation, a full response with every environment-dependent header, a concatenation of all cases into one artefact — looks efficient. One test, enormous coverage. What it buys is coupling. Every artefact now contains everything, so every artefact changes when anything changes. Three specific harms follow. *Unreviewable diffs.* A reviewer facing a 1,200-line difference where 9 lines matter will skim, and skimming is how a wrong value gets approved. *Blast radius.* In the receipt renderer behind an online bookstore checkout, a **340-case regression pack** of full rendered pages meant that changing the footer's legal wording turned all 340 red. The change was correct and trivial; the review was 340 diffs, and what actually happened was a bulk approval with nobody reading any of them. Had a real regression ridden along in that batch, it was approved too. *No localisation of meaning.* When an artefact covers everything, a failure tells you nothing about which behaviour moved. ### The fine failure The opposite extreme — one approved artefact per scalar field — removes every advantage the technique had. You now maintain a file and a review workflow to assert what one line of written assertion would have said better, and the written assertion would at least state the rule. Fine granularity also multiplies artefacts faster than anyone can own them. ### The rules that actually decide it **One reason to change.** The most useful test of an artefact's boundaries: list the behaviour changes that would alter it. If that list spans unrelated concerns — pricing rules *and* footer wording *and* navigation — the artefact is too coarse and should be split along those lines. **Reviewability in one sitting.** An artefact should be small enough that a competent reviewer can read the whole diff and form a judgement. This is a human constraint, not a byte count, and it is the constraint that keeps approvals meaningful. **Approve a projection, not the raw output.** The strongest move available is to design what gets approved: extract the payload under test into a stable, deliberately ordered, human-readable rendering, and approve that. Chrome, environment detail and incidental formatting stay outside the comparison by construction rather than by scrubbing — which is also the cheapest way to avoid the masking problems that broad normalisation rules cause. **Risk-weight the depth.** Behaviour that costs money or is hard to reverse deserves a deep, detailed artefact plus written assertions on its invariants. Peripheral output deserves a shallow one or none. **Layer the suite.** A workable shape is a small number of deep end-to-end artefacts that prove the whole pipeline still assembles, plus many narrow per-behaviour artefacts that carry the detail. The deep ones are allowed to be expensive because there are few of them. ### The metric to manage Track **re-approvals per intended change**. If a single, well-scoped behaviour change forces a three-digit re-approval count, granularity and coupling are wrong, and it should be treated as a defect in the test design rather than as a busy afternoon. The number is also the honest predictor of whether reviews are real: past roughly a screen's worth of diffs per person per change, they are not. ### Ownership and lifecycle Approved artefacts are source. They need an owner, they belong in review like any other change, and they need a deletion policy: an artefact nobody can judge — because the behaviour it captures is no longer understood, or because it is mostly placeholder text — is a liability that will be approved blindly the next time it moves. Being willing to delete such an artefact and replace it with a narrower one, or with a written assertion, is part of owning the strategy. ### Where the technique should not be extended Not every output deserves an approval. When the expected behaviour is expressible as a rule, a written assertion is better documentation and survives cosmetic change. Approval tests are for output too detailed to state and too valuable to leave unguarded — and deciding which outputs those are, across a suite, is exactly the granularity question.
- What number would tell you the granularity is wrong?Re-approvals per intended change. If one well-scoped behaviour change reddens a three-digit number of artefacts, they are all carrying content unrelated to their purpose. Treat that as a design defect: split along the reasons to change, or approve a narrower projection so that a footer edit touches the artefacts about footers and nothing else.
- You inherit a suite of huge approved artefacts nobody reads. What do you do first?Stop the bleeding before rewriting. Find which artefacts change most often per unrelated change and split those first, approving a projection of the payload instead of the raw output. In parallel add written assertions for the invariants that must never drift, so the suite keeps some real oracle while the artefacts are being reshaped.
- Does approving a projection weaken the test compared with approving the raw output?It narrows it deliberately, which is different from weakening it by accident. The raw form still needs some coverage, so keep a small number of deep artefacts for the assembled output. The projection buys reviewability and decoupling, and unlike broad scrubbing it removes content by an explicit design decision a reader can see.
- How do you decide an approved artefact should be deleted rather than re-approved?When nobody can judge it. If the behaviour it captures is no longer understood, or the file is mostly placeholder text after successive normalisations, its next approval will be blind. Replace it with a narrower artefact covering behaviour someone owns, or with written assertions on the invariants, and delete the old one rather than carrying it.
saying these in an interview costs you the question
- Approves whole rendered pages including unrelated chrome
- Treats a large re-approval count as normal work
- Creates one approved artefact per individual field
- Never deletes artefacts nobody can judge any more
- Assumes bigger artefacts always mean stronger tests
- Ignores whether a reviewer can read the diff at all