skip to content

How would you decide where test-first is worth mandating, given its published evidence is contested?

level: principalimportance: nice to knowfreq 24%

answer

  1. Do not pretend the studies agree
  2. Reason from mechanism, not from a number
  3. Value scales with cost of late change
  4. Ordering is invisible in the artefact
  5. Coverage is the wrong adoption metric

basics

~20 s

Mandate by mechanism, not by faith: apply test-first where its payoff — early requirements feedback and caller-driven interface design — is largest, such as new interfaces with unclear rules, and leave it optional where a clear oracle and a stable interface already exist.

solid answer

~50 s

Start by refusing the framing that this is a belief question. The published studies disagree about defect and effort outcomes, and much of the reported benefit may come from writing small tests at all rather than from the ordering, so no number settles it. What is not contested is the mechanism: writing the test first buys a requirements decision and an interface decision at the moment they are cheapest. So mandate it where those decisions are expensive to get wrong — new behaviour, unclear rules, interfaces many callers will adopt — and leave it to judgement where the oracle is obvious and the interface is fixed. Then measure the mechanism, not the ritual: how often a blocked assertion produced a real requirements decision, how many interfaces changed shape before implementation. Measuring adoption by coverage percentage buys the compliance and loses the point.

go deeper

for a junior

You are unlikely to set policy here, but know the honest position: the practice has a clear mechanism and a contested evidence base, so confident claims about measured defect reduction should be treated with care.

for a middle

Be able to say where the practice pays most — unclear rules, widely used interfaces, defect repair — and least, and explain that the payoff comes from early requirements and interface decisions rather than from coverage.

for a senior

Show you can argue for the practice without overstating the research, and that you would enforce it through checkable artefacts rather than through an ordering nobody can observe in a finished change.

for a principal

Own the differentiated policy and its cost: where it is mandatory, where it is encouraged, which mechanism metrics you would report, why you would refuse coverage as the adoption proxy, and how you would roll it out where teams have no slack.

## Why this is a judgement question and not a research question A principal-level answer starts by being honest about the evidence. Controlled studies and industrial case studies of test-first versus test-after report mixed results: some find lower defect density, some find no significant difference, several find higher up-front effort, and reviews note that studies vary so much in participant experience, task size and what the control group actually did that the results do not aggregate cleanly. There is also a well-known confound — test-first groups usually write more, smaller tests, so it is genuinely hard to separate the effect of the *ordering* from the effect of the *tests*. Anyone who quotes a settled percentage is overstating the literature. That does not leave you with nothing. It leaves you deciding on mechanism and cost, which is the normal condition for engineering practice decisions. ## Decide by where the mechanism pays Test-first delivers two things reliably, independent of any study: a requirements decision forced at the moment the assertion is written, and an interface decided by its first caller before any code depends on it. Both have value proportional to the cost of getting them wrong later. **Highest value:** - New behaviour whose rules are genuinely unsettled — the blocked assertion is worth more than the test. - Interfaces that many callers will adopt, where a shape change later is expensive across a codebase. - Rules with awkward edges: rounding, ordering, duplicates, boundary conditions in domain calculations. - Defect repair, where the report already supplies the example and the failing test proves the reproduction. **Lower value:** - Behaviour with an obvious oracle behind a fixed interface, where the test is genuinely a check and the design decisions are already made. - Thin wiring whose only real risk lives in the integration, not the shape. - Work whose purpose is to learn rather than to deliver, where the useful output is knowledge and the code is discarded. ## Mandate the outcome, not the keystroke order The practical difficulty with any mandate is that ordering is unobservable in the artefact. A finished change looks the same whichever order produced it, so a rule of the form "tests must be written first" is unenforceable and therefore becomes theatre. What you can actually require is checkable: that new behaviour arrives with tests in the same change, that a repaired defect arrives with a test that fails without the repair, that no interface with many intended callers ships without evidence someone used it before it existed. Those requirements are satisfied most easily by working test-first, which is the point. ## Measure the mechanism, and beware the proxy If you must show the practice is paying, measure the things the mechanism claims to produce, not the ritual: - How often did writing a test first surface a requirements question that reached a decision-maker? A team that never gets blocked is either working on fully specified problems or guessing. - How often did an interface change shape before it had an implementation? - Did defect reports arrive with reproducing tests, and did the same defect return? Avoid coverage percentage as the adoption metric. It is the most available number and the worst one here: it is reachable by test-after, it is reachable by tests with weak assertions, and making it the target reliably produces assertion-free tests that execute lines. Choosing a metric is choosing a behaviour, and this one selects for exactly the practice you were trying to move away from. ## Cost, and who pays it Be explicit that the practice has a real cost: it is slower per unit of visible progress at first, and it is a skill that takes weeks to become fluent, so a mandate imposed on a team under delivery pressure without slack produces compliance without benefit. The realistic rollout is narrow and voluntary first — one area with genuinely unclear rules, people who want to try it, and a specific claim to evaluate — then widened on the strength of what that area reports rather than on the strength of the literature. ## What a strong answer sounds like Say plainly that the evidence is mixed and that you would not pretend otherwise to a team. Then reason from mechanism to a differentiated policy: mandatory where requirements ambiguity and interface reach make early decisions valuable, encouraged elsewhere, and enforced through checkable artefacts rather than an unobservable ordering. Weak answers either declare test-first universally mandatory as a matter of professionalism, or dismiss it because a study found no effect. Both replace a judgement with a slogan.

  • A mandate cannot see the order in which files were written. How do you make the policy enforceable?
    Require checkable artefacts that the practice produces: new behaviour shipping with its tests in the same change, and a repaired defect shipping with a test that fails without the repair. Those are visible in a diff and in a pipeline. Working test-first is the cheapest way to satisfy them, so the policy pulls in the right direction without pretending to observe keystroke order.
  • Your leadership asks for a single number showing the practice is working. What do you offer?
    Offer a small set tied to the mechanism — requirements questions raised and decided before implementation, interfaces reshaped before they had callers, repaired defects arriving with a failing test, and recurrence of those defects. Decline coverage percentage explicitly, and say why: it is achievable without the practice and rewards tests that execute lines without asserting anything.
  • What would genuinely change your mind about a mandate in a given area?
    Evidence from that area rather than from the literature: if teams there report that assertions are almost never blocked and interfaces almost never change shape before implementation, the mechanism has nothing to pay for and the mandate is costing time for a benefit that does not arise. That argues for encouraging rather than requiring it there, while keeping the same-change test requirement.

It is like deciding where to require a surveyor rather than requiring one on every job: you send them where a wrong boundary is expensive to discover late, not everywhere the ground looks flat.

saying these in an interview costs you the question

  • Quotes a precise defect-reduction figure as settled fact
  • Mandates test-first everywhere as a matter of professionalism
  • Dismisses the practice because one study found no effect
  • Uses coverage percentage as the adoption metric
  • Ignores the learning cost when rolling it out under deadline

context