skip to content

Test-Driven Development

Driving code from tests you write first: the failing-test cycle, tiny increments, and design that arrives while restructuring. Asked because most candidates who claim TDD describe test-after.

on this pageshow

questions

25

What does "emergent design" mean in test-driven development, and where does the structure come from?

level: juniorimportance: must knowfreq 62%

answer

  1. Design is an output, not an input
  2. Which of the three steps designs
  3. Nothing fails without it, nothing justifies it
  4. Extracted from working code, never predicted
  5. Duplication and painful arrangement are signals

basics

~20 s

Emergent design means the structure of the code is an output of the cycle, not a plan made before it. Each restructuring step reshapes working code, so the design arrives gradually instead of being guessed up front.

solid answer

~50 s

Emergent design is the idea that the units, names and boundaries in the code come out of the loop rather than being decided ahead of it. The failing test and the make-it-pass step are deliberately unambitious — a conditional or a copied branch is fine. The design happens in the restructuring step, on code that already runs and is already pinned by tests, which is what makes reshaping cheap and reversible. A second rule does the negative work: only a failing test justifies new production code, so an abstraction has to be *extracted* from working duplication rather than written on a hunch. The cycle also hands you signals — duplication showing up in a third similar case, or an arrangement block that keeps growing. It is not "no design thinking"; the judgements are simply made later and on evidence.

code

pseudocode · 27 lines
pseudocode
function acceptBid(bid):
    auction = store.load(bid.auctionId)
    if now() > auction.closesAt: reject("closed")
    audit.append(bid.id, "accepted")
    return place(bid)

function acceptReserveBid(bid):
    auction = store.load(bid.auctionId)
    if now() > auction.closesAt: reject("closed")
    if bid.amount < auction.reserve: reject("under reserve")
    audit.append(bid.id, "accepted")
    return place(bid)

function acceptLateBid(bid):
    auction = store.load(bid.auctionId)
    if now() > auction.closesAt: reject("closed")
    auction.closesAt = extend(auction.closesAt)
    audit.append(bid.id, "accepted-extended")
    return place(bid)

# after the third case, extract what is now demonstrably shared
function acceptAnyBid(bid, rules):
    clock = auctionClock(store.load(bid.auctionId))
    if clock.isClosed(): reject("closed")
    outcome = rules.apply(bid, clock)
    audit.append(bid.id, outcome.label)
    return outcome

go deeper

for a junior

Be ready to state it in one sentence: the structure comes out of the restructuring step, repeated many times, not from a plan drawn before any code exists. Know the companion rule that new production code needs a failing test to justify it.

for a middle

Explain the mechanics: why restructuring covered, running code is cheap and reversible, and how an extracted abstraction differs from one written on a hunch. Name the two signals the cycle gives you — repeated duplication and a growing arrangement block.

for a senior

An interviewer expects you to defend this as a design practice with real limits: describe how you keep it working over months, what you do when restructuring stops paying, and where a green suite tells you nothing about the design.

for a principal

Own the boundary of the claim. Say which decision classes the restructuring step can reach, which it cannot, and what preconditions a team needs — continuous refactoring and a shared design bar — before "let it emerge" is responsible advice.

### The claim **Emergent design** is the claim that under test-driven development the structure of the code — its units, their names, the boundaries between them — is an *output* of the loop rather than an input to it. You do not draw the object graph first and then fill it in. You make one small behaviour work, reshape the working code, and repeat; after many repetitions the accumulated reshapings *are* the design. ### Where the structure actually arrives The loop has three beats: a test that fails, the smallest change that makes it pass, then restructuring while the suite stays green. The first two beats are deliberately unambitious. They are allowed to produce a conditional, a copied branch, even a hardcoded return. Almost all of the design work happens in the third beat — and it happens on code that already runs and is already pinned by tests. That is the whole trick. Reorganising behaviour you can execute is cheap and reversible; a structure guessed before any behaviour exists is a bet you cannot price, because you have no feedback on whether the split you invented is the split the problem has. So the honest one-line answer to "where does the structure come from" is: **from the third step, performed hundreds of times.** Not from writing tests as such — from restructuring under them. ### The rule that keeps undriven structure out The discipline carries a second rule that does the negative work: *only a failing test justifies new production code.* An abstraction therefore has to be pulled into existence by something that does not work without it. This distinguishes two very different ways a thing enters a codebase: * **Extracted** — it already exists as working, duplicated code, and you give it a name and one home. The tests that covered the duplication now cover the extraction. * **Undriven** — it is written because the author expects to need it. Nothing fails without it, so nothing tells you whether its shape is right, and nothing will tell you later either. Emergent design allows the first and has no route to produce the second. That is the mechanism, not a slogan. ### Reading the signals the cycle hands you Two signals do most of the work. **Duplication.** One occurrence is information. Two may be coincidence. By the third similar case the shape is usually real, and extracting it is a response to evidence rather than a prediction. Practitioners disagree about the exact threshold, and extracting too early is a genuine cost — an abstraction fitted to two cases often has to be torn out when the third case does not fit it — so treat "wait for the third" as a heuristic about *evidence*, not a law. **Pain in the test.** A test that needs a long arrangement before it can assert, or a class whose constructor takes a dozen collaborators, is reporting a fact about the production code: this unit depends on that many things. Because the test is written from the outside, it feels that cost before any caller does. ### A worked example A marketplace bidding engine. The first test says a higher bid beats a lower one. The second adds a reserve price. The third adds an auto-extend rule that pushes the close time out when a bid lands near the end. By that third case, all three paths independently load the auction, compare a timestamp against the close time, and append an audit line. The duplication is concrete, so extraction is warranted, and what comes out of it — an auction-clock concept that owns "is this bid in time", and a bid-outcome record that owns the audit line — are two ideas nobody would have named on day one. Contrast the undriven version: on day one someone writes a bid-strategy registry and a loader for the single strategy that exists. It passes every test, adds two indirections, and encodes a guess about how bid types will vary. When the auto-extend rule arrives it turns out to vary along a different axis, and the registry is now a thing to work around. ### What emergent design is not It is not "no design thinking". A practitioner makes design judgements every few minutes; they are simply made *later*, on evidence, and in small reversible moves. It is not permission to skip the restructuring beat — skip it and nothing emerges but sediment. And it is not a claim that every decision is equally reversible: some decisions sit outside anything a refactor step can reach, and those still get made deliberately. ### Where the evidence stands Be honest in an interview: the empirical research on whether test-first work improves *design quality* is mixed rather than settled, and several studies attribute part of the measured effect to working in small increments rather than to test-first ordering specifically. Argue emergent design from its mechanism — feedback, reversibility, and the rule that undriven code cannot get written — not from a claimed number.

  • If the design emerges from restructuring, what stops a developer from adding a useful-looking abstraction anyway?
    Nothing mechanical stops them — the discipline does. The rule is that new production code needs a test that fails without it, so an abstraction is either extracted from duplication that already exists and is already covered, or it is undriven. Undriven structure gets no feedback on its shape, which is why the rule exists; the moment a team relaxes it, the design stops being driven by evidence and goes back to being guessed.
  • Does emergent design mean the team never discusses design before writing code?
    No. Teams still discuss the shape of a feature, sketch a boundary, and argue about naming — the difference is that those discussions are treated as starting hypotheses rather than commitments, and the code is allowed to contradict them. What emergent design rejects is a detailed structure specified before any behaviour exists, because that structure has had no chance to be corrected by feedback.
  • How strong is the evidence that this actually produces better designs?
    It is contested. Studies on test-first development report mixed results on internal design quality, and some attribute part of the benefit to working in small increments rather than to test-first ordering specifically. In an interview it is better to argue from the mechanism — reversible changes on covered code, and the rule that undriven code cannot get written — than to claim a measured improvement.

It is closer to laying a footpath where people have already worn a line in the grass than to drawing the path on a map before anyone has walked anywhere.

saying these in an interview costs you the question

  • Says emergent design means doing no design thinking at all
  • Believes the design comes from writing tests, not from restructuring
  • Writes the abstraction first and adds a test to cover it
  • Treats it as licence to skip the restructuring step entirely
  • Claims the cycle guarantees a good design automatically
  • Says every design decision can safely be deferred

context

open as a page

In TDD, what is the difference between outside-in and inside-out development?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Outside-in TDD starts at an outer boundary and invents collaborator roles as they are needed, standing them in until they are built. Inside-out TDD starts with domain behaviour built from real objects and wires the outer layers in last.

open as a page

Why must a test-driven test be watched failing before you write the code that passes it?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Watching the test fail proves the assertion can actually detect the behavior's absence. A test that has never been red may be mis-wired, never executed, or trivially true, and would report green forever while guarding nothing.

open as a page

Why run the test suite after each small refactoring move instead of once at the end?

level: juniorimportance: must knowfreq 66%

basics

~20 s

Because a green run after every move tells you exactly which move changed behaviour. Run once at the end and a red suite only says something in the last hour was wrong, so you debug instead of undoing.

open as a page

What is test-first development, and how does it differ from writing tests after the code?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Test-first means writing an executable test for behaviour that does not exist yet, then writing code to satisfy it. Test-after writes tests against finished code, so the cases are shaped by what the code already does.

open as a page

Why do unit tests that assert on a class's internals turn red during a behaviour-preserving restructuring?

level: middleimportance: must knowfreq 68%

basics

~20 s

Because those assertions describe how the code is built, not what it does. Rename an internal helper or replace a private data structure and the observable behaviour is unchanged, but the assertions no longer match, so the suite fails.

open as a page

While refactoring production code under a green suite, may you edit the tests in the same step?

level: middleimportance: must knowfreq 58%

basics

~20 s

No. While structure changes, the tests are the fixed oracle that says behaviour did not. Change both at once and neither is checking the other, so any drift you introduce can be absorbed by the edited assertion.

open as a page

What is triangulation in TDD, and what does the second test case force?

level: middleimportance: must knowfreq 56%

basics

~20 s

Triangulation is deliberately adding a second, differing example so a hardcoded or over-specific implementation can no longer pass. The two cases pin down what varies, forcing you to replace the constant with the general rule the examples share.

open as a page

How does writing the test first shape the interface of code that does not exist yet?

level: middleimportance: must knowfreq 64%

basics

~20 s

The test is the interface's first caller, so it must invent the operation's name, the arguments handed in, the shape of the result and how failures are reported — all judged from the outside, before any implementation can bias them.

open as a page

In test-driven development, what goes wrong when a team always skips the refactor step?

level: juniorimportance: should knowfreq 58%

basics

~20 s

You keep the regression protection but lose the design payoff. The code becomes whatever the quickest path to green left behind: duplicated blocks, growing setup and a branch per case, until changing it costs more than the tests save.

open as a page

What is "fake it till you make it" in TDD, and when do you stop faking?

level: juniorimportance: should knowfreq 50%

basics

~20 s

Fake it till you make it means passing a new failing test with the smallest code possible, often a hardcoded constant, then generalising once further tests demand it. The fake proves the test really discriminates before any design work begins.

open as a page

A unit test needs a 60-line setup block and its class takes eleven constructor parameters — what is that telling you about the design?

level: middleimportance: should knowfreq 57%

basics

~20 s

Test pain is design feedback. A long arrangement block and a wide constructor both report that the unit depends on too many things, so the fix belongs in the production code, not in a bigger setup helper.

open as a page

Why does outside-in TDD leave more test doubles behind than inside-out TDD?

level: middleimportance: should knowfreq 54%

basics

~20 s

Outside-in starts before the collaborators exist, so each invented role needs a stand-in for the test to run — and nothing forces a later swap for the real object. Inside-out builds collaborators first and uses them directly.

open as a page

A new unit test passes on its first run, before the code it drives exists. How do you diagnose that?

level: middleimportance: should knowfreq 46%

basics

~20 s

Treat an unexpected green as a broken test. Force it red by changing the expected value to something certainly wrong; if it still passes, the case is not executed, the assertion is unreachable, or the comparison is vacuous.

open as a page

When starting a feature with TDD, how do you choose outside-in or inside-out?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Start where the uncertainty is. If the collaboration and the boundary are unclear, drive outside-in and invent roles at the moment of need; if the boundary is agreed and a domain rule is the hard part, drive inside-out.

open as a page

When is test-driven development the wrong tool for the work in front of you?

level: seniorimportance: should knowfreq 47%

basics

~20 s

When you cannot state the expected result before writing the code, or cannot get an answer in seconds. Exploratory spikes, probing an unfamiliar external interface, visual judgement and threshold-based performance work all break one of those two preconditions.

open as a page

Your refactoring left every test green. Why is that not proof that behaviour is unchanged?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A suite pins only what it exercises and only what it asserts on. Paths it never reaches, wire-level shapes it never inspects, and timing, ordering or output it never measures can all drift while every test passes.

open as a page

A TDD increment now needs hundreds of lines to go green. How do you split it into smaller steps?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Split along the behaviour's degrees of freedom: pick one input dimension or one outcome, write a case that varies only that, and let the rest stay unimplemented or degenerate. Each step should be a single edit you can reason about between two green runs.

open as a page

You cannot write the assertion for a new feature's first test. What does that tell you, and what do you do?

level: seniorimportance: should knowfreq 46%

basics

~20 s

An assertion you cannot write means the requirement has no agreed oracle for that case yet. Treat it as a discovered ambiguity: stop, get a decision on the expected outcome from whoever owns the behaviour, then encode that decision as the assertion.

open as a page

In emergent design, which decisions must still be made deliberately up front, and how do you decide?

level: principalimportance: should knowfreq 44%

basics

~20 s

Decide by reversibility and reachability. Anything a later restructuring can reach inside code you own — unit shapes, module internals, names — let it emerge. Stored data shapes, process boundaries, contracts other teams consume and cross-cutting authorities get decided deliberately.

open as a page

As a lead, how do you keep the red-green-refactor cycle honest once the feedback loop it depends on has degraded?

level: principalimportance: should knowfreq 38%

basics

~20 s

Fix the loop rather than police the ritual. The micro-cycle needs seconds-scale feedback, so tier the suite into a fast local subset and a slower full pack, and hold the fast tier to a runtime budget.

open as a page

Where does an automated rename stop guaranteeing that behaviour is preserved?

level: middleimportance: nice to knowfreq 24%

basics

~20 s

The guarantee covers references the tool's analyser can resolve in the sources it indexes. Names reached from strings, configuration, templates, generated code, stored data or another repository are outside it, and a rename silently breaks them.

open as a page

How do you tell whether an emerging design is converging, and what do you do when repeated refactoring stops improving it?

level: seniorimportance: nice to knowfreq 21%

basics

~20 s

Convergence shows as restructurings getting smaller and rarer while changes stay local. When the same duplication keeps returning and small features touch many files, the missing structure is bigger than a restructuring step can reach and needs a deliberate decision.

open as a page

How long do you let one red-green-refactor cycle run before discarding the edit and returning to the last green commit?

level: seniorimportance: nice to knowfreq 24%

basics

~10 s

Minutes, not hours. Once you are debugging your own uncommitted edit and can no longer say which change caused which failure, discard it, return to the last green commit, and restart smaller.

open as a page

How would you decide where test-first is worth mandating, given its published evidence is contested?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Mandate by mechanism, not by faith: apply test-first where its payoff — early requirements feedback and caller-driven interface design — is largest, such as new interfaces with unclear rules, and leave it optional where a clear oracle and a stable interface already exist.

open as a page