As a lead, how do you keep the red-green-refactor cycle honest once the feedback loop it depends on has degraded?
answer
- Cost decides the practice, not belief
- Fix the loop, do not police the ritual
- Two tiers, one budget
- A small minority owns the runtime
- Proxies: change size, feedback time
basics
~20 sFix the loop rather than police the ritual. The micro-cycle needs seconds-scale feedback, so tier the suite into a fast local subset and a slower full pack, and hold the fast tier to a runtime budget.
solid answer
~50 sThe cycle is an economic mechanism: it works when running the relevant tests costs less attention than switching context, and it quietly dies when it does not. So the lead's job is the loop, not the ceremony. Split the suite into tiers — a subset that runs in seconds against the area being changed, and the full pack on the pipeline — and treat the fast tier's runtime as a budget that fails the build when it is exceeded. Attack the slow minority first, because a handful of cases usually dominate; remove per-case infrastructure startup, run selectively by impacted area, and parallelise. Then watch proxies rather than mandating behavior: local feedback time, change size per green commit, and how often people run tests before pushing. Be honest that the published evidence on defect outcomes is mixed, and argue from feedback speed, which is measurable here.
code
pseudocode · 5 linesfastTier = suite.filter(tag == "fast")
elapsed = run(fastTier)
if elapsed > seconds(45):
fail("fast tier budget exceeded: " + elapsed)go deeper
Understand the basic dependency: the cycle only works when running the relevant tests is quick. If a run costs minutes, people stop doing it, whatever the team says it practises.
Be able to explain the tiering — a fast in-process subset for the cycle, the slower levels on the pipeline — and to name the usual causes of a slow suite, such as per-case infrastructure startup and real waiting.
Show that you measure before optimising, know that a small minority of cases usually dominates runtime, and can describe the risks of the fixes themselves, such as nondeterminism introduced by parallel runs over shared state.
Own the tradeoffs: what share of engineering time suite speed deserves, whether a runtime budget should fail the build, which proxies you would report upward, and how to argue for the practice honestly when the published evidence on its outcomes is mixed.
## The cycle is an economic mechanism Red-green-refactor is not sustained by belief. It is sustained by cost. When running the tests that matter takes a couple of seconds, a developer runs them constantly because it is cheaper than reasoning; when it takes long enough to lose the thread, they batch verification to the end, and what remains is test-after wearing the cycle's vocabulary. Nobody announces the switch. It shows up as larger commits, longer reds, and "we do TDD" in a retro alongside a suite nobody runs locally. So the lever a lead actually has is the cost of a run, and the failure mode to watch for is a team that still says it practises the cycle while its behavior has quietly changed. ## Tier the suite deliberately One suite cannot serve both a developer mid-edit and a release gate. Give the cycle its own tier: - **A fast tier** — in-process, no external infrastructure, no network, no real clock — targeted at the area being changed and expected to complete in seconds. This is the tier the micro-cycle runs against. - **A full pack** — integration, contract and end-to-end levels — run on the pipeline on every push, not on every keystroke. The tiering only holds if it is enforced. Make the fast tier's runtime a budget with a threshold that fails the build when exceeded, so the drift back to slowness is caught at the moment it is introduced rather than a year later. ## Attack the slow minority, and the structural causes Suite runtime is nearly always dominated by a small fraction of cases. Measure per-case durations before optimising anything, because the intuition about which cases are slow is usually wrong. Then, in rough order of return: - **Move infrastructure out of the per-case path.** Standing up a container, a browser, or a full application context per case is the usual culprit; share one instance across the tier, or replace it with an in-process stand-in for cases that do not exercise the integration. - **Remove real waiting.** Fixed sleeps and real timers inflate a suite without adding coverage; inject the clock instead. - **Select by impact.** Run the cases whose covered code the change touches, keeping the full pack on the pipeline as the safety net. - **Parallelise**, but only after isolating shared state, or you trade slow for nondeterministic. - **Delete duplication across levels.** The same rule asserted at three levels costs three runtimes and pays once. ## A worked example On a tax-filing wizard, the suite reaches 27 minutes and developers stop running anything locally. Per-case timings show 61 cases — under a tenth of the suite — accounting for roughly two-thirds of the wall time, nearly all of them starting a full application context to assert one validation rule about a locale-dependent amount format. Rewriting those against the validator directly, and keeping a handful at the integration level for wiring, brings the fast tier to about 40 seconds. The full pack still runs on the pipeline. Within a few weeks the observable behavior changes: commits get smaller and more frequent, and the median time a developer sits red drops — which is the actual objective, and it is measurable, unlike "the team does TDD". ## Measure proxies, do not police the ritual You cannot observe whether someone wrote the test first, and trying to is corrosive. Observable proxies that correlate with a living cycle: - **Local feedback time** for the fast tier, tracked as a number that has a budget. - **Change size per commit** and how often commits land on a green state. - **How long people sit red**, if your tooling can see it, or self-reported in retros. - **Cases added in the same change as the behavior**, visible in review without asking anyone to prove ordering. What not to use: a coverage percentage as a proxy for the cycle. It measures execution, not discrimination, and treating it as a target reliably produces cases written to raise it. ## Be honest about the evidence A principal-level answer should be candid that the research literature on whether test-driven development improves defect density or productivity is mixed, with results varying by study design, team and what counted as the practice. Selling the cycle on a claimed defect-reduction multiple is over-claiming and invites a fair challenge. The defensible argument is local and measurable: a fast feedback loop shortens the distance between a mistake and its discovery, and short cycles keep changes small, which is independently valuable for review and for recovery. Those you can demonstrate on your own team's numbers. ## What you own that a senior does not The tradeoff calls: how much engineering time to spend on suite speed versus features this quarter; whether to hold a runtime budget as a build failure and accept the friction; whether a team with a genuinely slow domain should adopt a different cadence rather than pretend; and how to introduce the practice without turning it into a compliance exercise that produces tests written after the fact and reordered in the commit.
- A team insists it practises the cycle, but the suite takes half an hour. What do you expect to find?That the practice has become test-after in vocabulary only: verification batched to the end, larger commits, long reds, and local runs replaced by pushing and waiting. I would check per-case timings, commit sizes and how often the fast path is actually run before treating the claim as accurate.
- Why not use a coverage percentage as the metric for whether the discipline is healthy?Because it measures which lines were executed, not whether any assertion would notice them behaving wrongly. It rises with cases that assert nothing, and once it is a target people write to the number. Prefer proxies tied to the loop itself — fast-tier runtime, change size per green commit — and reserve deeper analysis for critical modules.
- How do you introduce the practice on a team that has never worked this way, without it becoming compliance theatre?Make the fast path the easy path first — a seconds-scale tier and a clean way to run it — then teach the loop on real work in pairs rather than by mandate. Judge it by observable outcomes such as smaller changes and shorter reds, and never by asking people to prove they wrote the test first.
saying these in an interview costs you the question
- Mandates the ritual without fixing the feedback loop
- Cites a hard defect-reduction number as settled evidence
- Uses a coverage percentage as the health metric
- Optimises runtime by guesswork instead of per-case timings
- Parallelises before isolating shared state
- Treats one undifferentiated suite as serving every purpose