skip to content

Describe the red-green-refactor cycle of test-driven development, and explain exactly what work belongs in each of its three steps.

level: juniorimportance: must knowfreq 74%

answer

  1. Red = failing test, watch it fail
  2. Green = simplest thing, ugly allowed
  3. Refactor = clean under green, no new behavior
  4. Fake it / obvious / triangulate
  5. Refactor step is the one teams skip

basics

~20 s

Red: write one small failing test for behavior you want. Green: write the simplest code that makes it pass. Refactor: clean up the code and the test while all tests stay passing. Then repeat with the next small behavior.

solid answer

~50 s

Red-green-refactor is the core loop of test-driven development. **Red**: write one small test for a behavior that does not exist yet and watch it fail for the expected reason — a test you never saw fail proves nothing. **Green**: do the simplest thing that makes it pass, even something ugly like a hardcoded return; speed matters more than elegance here because you want the safety net live as soon as possible. **Refactor**: with a green bar, remove duplication, name things properly, and reshape the design — running tests after each small step, never adding behavior. Tests themselves get refactored too. Then loop. The value is that design improvement happens under continuous verification, and each cycle is minutes long, so a mistake costs a revert rather than a debugging session. The most-skipped step is the third: teams stop at green, and design debt accumulates one "we'll clean it later" at a time.

code

pseudocode · 14 lines
pseudocode
// RED — fails: no such function yet
test("empty basket costs zero") { assertEquals(0, price([])) }

// GREEN — simplest thing that passes (fake it)
function price(items) { return 0 }

// RED again — forces generalisation (triangulation)
test("one item at 5 costs 5") { assertEquals(5, price([item(5)])) }

// GREEN
function price(items) { var t = 0; for (i in items) t = t + i.price; return t }

// REFACTOR — same behavior, clearer expression; tests stay green, untouched
function price(items) { return sum(items, i -> i.price) }

go deeper

for a junior

Name the three steps in order and say what happens in each; give a tiny concrete example and mention that tests must fail first.

for a middle

Add the strategies for going green (fake it, obvious implementation, triangulation), stress that refactoring happens only on green, and that tests get refactored too.

for a senior

Connect the cycle to the two-hats rule, commit-at-green discipline, and why coupling tests to implementation details kills the refactor step for the whole team.

for a principal

Frame it as the smallest instance of a general principle — change under continuous verification — and discuss what makes it scale or fail organisationally: suite speed, trunk-green policy, flaky tests destroying the signal, and how large refactorings need staged equivalents (parallel change, branch by abstraction) because a single red step across a shared codebase is not affordable.

## The loop **Test-driven development (TDD)** drives implementation from tests. Its inner loop has three states, conventionally named after the colour a test runner shows: ``` ┌──────► RED (write a small failing test) │ │ │ ▼ │ GREEN (make it pass the simplest way) │ │ │ ▼ └────── REFACTOR (improve design, tests stay green) ``` ### Step 1 — Red Write **one** test describing a behavior that does not exist yet, then run it and **watch it fail**. - Watching the failure is not ceremony. A test that passes before the feature exists is testing nothing — a typo in the assertion, a mis-wired fixture, or a filter that excludes the file all produce silently useless tests. - Check *why* it fails. "Expected 5, got 0" is the right red. "NullPointerException in setup" or "class not found" is a different red — fix the harness first. - Keep the step small. A test that requires four new classes to go green is too big; shrink the scope until the green step is minutes, not hours. - The test is also a design act: writing the call before the implementation forces you to choose a name, an argument list, and a return shape from the *caller's* point of view. ### Step 2 — Green Make the test pass **as quickly as possible**. Kent Beck names three strategies: 1. **Fake it** — return a constant. Legitimate, because the next test will make the constant untenable. 2. **Obvious implementation** — if you can see the real code and it's small, just type it. 3. **Triangulation** — write a second test with different data to force the constant into a real generalisation. Ugly code is *allowed* here. This is the counter-intuitive part for newcomers: the point is to get back to a verified state fast. Cleanliness is the next step's job, and you will do it better with tests protecting you. ### Step 3 — Refactor Now, and only now, improve the design — with all tests green and **without adding any behavior**: - Remove the duplication that the green step introduced (Beck: "the duplication between the test data and the code is the driver of generalisation"). - Extract functions, rename to intention-revealing names, move responsibilities to the class that owns the data, replace conditional chains with polymorphism or lookup, collapse special cases. - **Refactor the tests too**: extract builders, remove copy-pasted setup, improve test names. Test code is production code for the team's velocity. - Run the suite after each small step; commit at green points. If a step goes wrong, revert rather than debug. Then return to red with the next behavior. ## Why the ordering matters - **Refactoring before green is unsafe** — you have no net. - **Refactoring after green is cheap** — the net catches you within seconds, so you take structural risks you would otherwise avoid. - **Behavior work and structure work never overlap**, which is the *two hats* rule: at any moment you are either adding a test-driven behavior or refactoring, never both. If, while refactoring, you spot a missing behavior, write it on a to-do list and continue; switch hats deliberately. ## Common failure modes | Failure | Symptom | Fix | |---|---|---| | Skipping refactor | Green tests, sprawling design, duplication everywhere | Treat the third step as non-optional; time-box it if needed | | Steps too big | Long red periods, debugging inside the cycle | Shrink the test; delete and re-approach | | Writing tests after the code | Tests shaped to the implementation, rarely fail meaningfully | Write the test first, watch it fail | | Testing implementation details | Every refactoring breaks tests, so nobody refactors | Assert on observable behavior through a stable seam | | Never watching red | Tests that can never fail ship silently | Always run the failing test once | ## Relationship to refactoring in general Red-green-refactor is one *delivery mechanism* for refactoring, not the only one. Refactoring also happens as **preparatory refactoring** ("make the change easy, then make the easy change" — Beck), as **opportunistic cleanup** under the Boy Scout rule (leave the code better than you found it), and as **planned campaigns** on legacy code where characterization tests must be written before any structural move is safe. What all of these share with TDD's third step is the precondition: *a green, trusted suite before you start moving things*. ## What "green throughout" means in practice - Commit only at green. A repository history where every commit compiles and passes lets you bisect, revert, and release at any point. - Keep the suite fast. If the loop takes 20 minutes, developers stop running it, and the cycle silently degenerates into "code, then test later". - Never leave the tree red for others. On a shared mainline, a broken build blocks everyone; that is why large refactorings are staged (parallel change, feature toggles, branch by abstraction) rather than done as one giant red step.

  • Why insist on watching the test fail before writing the implementation?
    Because a test that has never failed has never been shown to be capable of failing. Typos in assertions, tests excluded by a naming filter, fixtures that swallow errors, or asserting on the wrong object all produce tests that pass unconditionally. Seeing the expected failure message also confirms you are testing the thing you think you are.
  • During the refactor step you notice a missing edge case. What do you do?
    Write it on a to-do list and finish the current refactoring step, then switch hats: go back to red with a test for that edge case. Mixing a behavior fix into a refactoring step means a failing test can no longer tell you whether the structure move or the new logic was wrong.
  • Teams often say "we do TDD but we skip the refactor step when we're busy." What is the consequence?
    You get the test coverage without the design benefit. The green step deliberately produces expedient code — duplication, poor names, procedural clumps — on the promise that step three cleans it. Skipping it institutionalises that expedient code, and because the suite is green, nothing signals the accumulating design debt until change becomes expensive.

Climbing with protection: clip a bolt (red→green gives you a verified anchor), then move (refactor) knowing a slip costs a short fall, not the ground. Nobody makes the bold move before clipping in.

saying these in an interview costs you the question

  • Writing tests after the implementation and calling it TDD
  • Never observing a red bar, so unconditionally-passing tests ship unnoticed
  • Treating the refactor step as optional "if there's time"
  • Refactoring while the suite is red or while adding behavior in the same step
  • Steps so large that the green phase takes hours and requires debugging
  • Tests bound to private methods and call sequences, so any refactoring breaks them and the team stops refactoring

context