skip to content

Red-Green-Refactor Cycle

The micro-cycle: a failing test that states the desired behavior, the minimum code to pass it, then restructuring with the test as a safety net. Asked because the point of each distinct phase — especially why the test must fail first — is what separates practice from ritual.

on this pageshow

questions

4

Why must a test-driven test be watched failing before you write the code that passes it?

level: juniorimportance: must knowfreq 78%

answer

  1. A test has two jobs, not one
  2. Green from birth proves nothing
  3. The missing code is a free defect
  4. Read the message, not just the colour
  5. Vacuous, unreached, uncollected, or self-comparing

basics

~20 s

Watching the test fail proves the assertion can actually detect the behavior's absence. A test that has never been red may be mis-wired, never executed, or trivially true, and would report green forever while guarding nothing.

solid answer

~50 s

A test does two jobs: it states the intended behavior, and it detects that behavior going missing. Only a failing run demonstrates the second job. If you write the code first and the test is green on its first run, you have no evidence the assertion is reached, that the runner even collected the case, or that the comparison is meaningful rather than a value checked against itself. The red step is a one-off experiment where the missing implementation is the known defect, and the test is the instrument being calibrated against it. The discipline goes further than "see red": read the failure message and confirm it fails for the reason you intended, naming the expected and actual values you meant to compare. A red caused by a setup error or a missing symbol has not yet calibrated anything.

code

pseudocode · 4 lines
pseudocode
test "review screen shows the deduction total":
    expected = formatCurrency(1247.35)
    actual   = renderReviewLine(deduction = 1247.35)
    assert actual contains expected

go deeper

for a junior

Be ready to say plainly that a test which has never failed has not been shown to work. Recall two or three ways a green test can guard nothing: it was never run, the assertion is never reached, or it compares a value to itself.

for a middle

An interviewer expects the mechanics: what the failure output should name, why a setup error or a missing symbol is a weaker red than an assertion mismatch, and how you calibrate a test added around code that already works.

for a senior

Show that you treat the first red as a chance to inspect the failure message a future on-call reader will see, and that you can list the concrete ways real suites go silently vacuous — derived oracles, over-wide tolerances, doubles that answer the assertion directly.

for a principal

Own the scaling question: on a large suite you cannot personally watch every test fail, so argue for the mechanisms that substitute — deliberate-defect checks in review, mutation-style analysis on critical modules, and alarms on suites whose pass rate has never varied.

## The red step is a calibration, not a ceremony A test is an instrument. It has two responsibilities, and they are not the same: 1. **State the intended behavior** — a reader can see, from the test alone, what the system is supposed to do. 2. **Detect the absence of that behavior** — when the behavior breaks or disappears, the test turns red. Writing the test carefully gets you the first responsibility. Nothing about writing it gets you the second. The only evidence that a test can detect a defect is having seen it detect one. In the red step you have a guaranteed defect available for free — the implementation does not exist yet — so you point the instrument at it and confirm the needle moves. Skip that, and the first time you learn whether the test works is the day it was supposed to catch a regression and did not. A useful way to think about it: watching the test fail is a single-case mutation experiment, where the "mutant" is the absence of the production code. The test either kills that mutant or it does not. ## What a never-red test can be hiding All of these report green and all of them guard nothing: - **The assertion is never reached.** An early return, a conditional the test does not satisfy, or an exception swallowed inside the test body means the run ends before any comparison happens. - **The case is never executed at all.** It was not collected by the runner because of a naming or discovery rule, it carries a skip or ignore marker left over from an earlier day, or a tag filter excludes it from the command that is actually being run. - **The comparison is vacuous.** A value compared to itself; a literal compared to the same literal; an assertion that checks a value is "not empty" when it never could be. - **The oracle is derived from the system under test.** The test computes its expected value by calling the same helper the production code calls, so the two agree no matter what either does. - **The setup supplies the answer.** The fixture writes the value the assertion later reads back, so the assertion measures the fixture, not the behavior. - **A stand-in returns exactly what the assertion wants.** A configured double satisfies the check while the real collaborator is never involved. - **The tolerance swallows the defect.** A numeric comparison with a wide margin passes over the very error the test was written for. Every one of these is caught in seconds by the red step, and by nothing else that is cheap. ## "Fail for the right reason" Seeing any red is weaker than seeing the red you predicted. Before you write the implementation, read the failure output and ask three questions: does it name the assertion you wrote; are the expected and actual values the ones you meant to compare; and is the failure about behavior rather than about the test's own plumbing? A run that dies in setup, or that reports an unresolved reference, or that errors on a fixture that never loaded, tells you the test infrastructure is broken — not that the assertion works. The failure message is also the message a future teammate will read at the worst possible moment. The red step is the one time you are guaranteed to see it. If it says only that two values differ without saying which case or which field, improve it now, while the cost is zero. ## A worked example A team building a tax-filing wizard adds a step that renders the taxpayer's deduction total on a review screen. The developer writes a test asserting the rendered line matches an expected string, and builds that expected string by calling the same currency-formatting helper the screen uses. Written test-after, it passes immediately and gets committed. Months later a change makes the helper emit a locale-dependent format, so the review screen shows a decimal comma where the filing schema requires a decimal point, and the test still reports green — because both sides of the comparison moved together. Written test-first, the very first run would have shown green before the screen existed, and the developer would have caught the self-referential oracle in under a minute. ## When there is no natural red Adding a test to code that already works — a characterization test around existing behavior — cannot be red on its own. The discipline still applies, inverted: break the production code deliberately once (change a constant, invert a condition, return early), confirm the new test turns red, then restore the code. That is the same calibration, done by hand. It is worth being honest about scope. The published evidence on whether test-driven development lowers defect density overall is mixed and study-dependent, and a candidate who cites a hard number is over-claiming. The argument for watching a test fail does not need that evidence: it is a local, verifiable claim about one test on one run, and it costs seconds.

  • The first run fails because the function the test calls does not exist yet. Is that a satisfactory red?
    It is a real red, but a weak one: it proves the test names something that is missing, not that the assertion discriminates. Once the function exists as an empty shell, run again and check that the failure now reports the expected and actual values you intended. That second red is the one that calibrates the assertion.
  • How do you get the same guarantee for a test you are adding to code that already behaves correctly?
    Introduce a deliberate, temporary defect in the production code — flip a condition, change a constant, return early — confirm the new test goes red, then restore the code and confirm it is green again. That is the same calibration performed by hand, and it takes under a minute.
  • You watched it fail, so the test is trustworthy forever?
    No. It is calibrated against one defect at one moment. Later edits to the test, its fixtures or its doubles can quietly neutralise it, and a test whose assertions are removed or whose setup starts supplying the answer will still show green. Re-establishing confidence means seeing it fail again, which is why deliberate-defect checks and mutation-style analysis exist.

A smoke alarm you have never tested is a plastic box on the ceiling. Pressing the button while you can still hear the beep is the only thing that turns it into a smoke alarm.

saying these in an interview costs you the question

  • Treats the red step as bureaucracy to be skipped when busy
  • Says any red counts, without reading the failure message
  • Thinks a green first run means the feature already works
  • Cannot name a way a passing test can guard nothing
  • Builds the expected value with the same code under test
  • Believes writing a test carefully proves it can detect a defect

context

open as a page

A new unit test passes on its first run, before the code it drives exists. How do you diagnose that?

level: middleimportance: should knowfreq 46%

basics

~20 s

Treat an unexpected green as a broken test. Force it red by changing the expected value to something certainly wrong; if it still passes, the case is not executed, the assertion is unreachable, or the comparison is vacuous.

open as a page

As a lead, how do you keep the red-green-refactor cycle honest once the feedback loop it depends on has degraded?

level: principalimportance: should knowfreq 38%

basics

~20 s

Fix the loop rather than police the ritual. The micro-cycle needs seconds-scale feedback, so tier the suite into a fast local subset and a slower full pack, and hold the fast tier to a runtime budget.

open as a page

How long do you let one red-green-refactor cycle run before discarding the edit and returning to the last green commit?

level: seniorimportance: nice to knowfreq 24%

basics

~10 s

Minutes, not hours. Once you are debugging your own uncommitted edit and can no longer say which change caused which failure, discard it, return to the last green commit, and restart smaller.

open as a page