When is test-driven development the wrong tool for the work in front of you?
answer
- Start from preconditions, not a banned list
- Can you state the result first?
- Is the feedback measured in seconds?
- Spikes produce knowledge, not code
- Timebox, discard, then drive it
basics
~20 sWhen you cannot state the expected result before writing the code, or cannot get an answer in seconds. Exploratory spikes, probing an unfamiliar external interface, visual judgement and threshold-based performance work all break one of those two preconditions.
solid answer
~50 sFrame it as preconditions rather than a banned list. The cycle needs an expected result you can state in advance and feedback fast enough to keep steps small. Exploratory spikes fail the first: the output is knowledge, not code, so you timebox it, throw the code away, and then drive the real implementation with the assertion you can finally write. Probing an unfamiliar external interface is the same shape. Visual and wording work has no mechanical oracle. Performance work fails it differently — the expected result is a distribution, so a hard threshold assertion flips without any code change and trains people to ignore failures. Say plainly that a real exception is named, timeboxed and has a re-entry point; a standing exemption for a whole subsystem is the abuse of this answer, and areas merely described as hard are usually where a pre-written assertion pays most.
go deeper
Know that a spike is throwaway exploratory code with a timebox, and that its purpose is to answer a question. Be ready to say why you cannot write an assertion for an answer you do not have yet.
Explain the two preconditions — a statable expected result and fast feedback — and map at least two concrete cases onto them, including why a hard threshold assertion on a measured duration is unstable.
Show the judgement: deciding on a real piece of work which part is a spike and which part is specifiable, timeboxing it, discarding the draft, and being candid that the evidence base for the practice is mixed.
Own where the exception line sits for an organisation, so it stays a named and expiring decision rather than a subsystem-wide exemption, and so nobody has to argue the case from scratch on every piece of work.
### The preconditions a cycle needs A red-green-refactor cycle needs two things, and both are easy to state: 1. **You can name the expected result before you write the code.** The failing test is a specification. If you cannot write the assertion, you do not yet have a specification — you have a question. 2. **You get an answer in seconds.** The cycle is a feedback loop; at ten minutes per run it stops being a loop and becomes a batch process, and people stop taking the steps small. Almost every "TDD does not fit here" case is one of these two preconditions failing. That framing is more useful in an interview than a list, because it tells you what to do next: restore the missing precondition, or accept that a different technique owns this piece of work. ### Where the first precondition fails **Exploratory spikes.** You are not building a feature, you are answering a question — can this library do the thing, what shape does that response actually have, is this approach fast enough. The output of a spike is *knowledge*, not code. Driving it with tests is expensive and pointless, because you cannot assert on an answer you are trying to discover. The discipline is: timebox it, write it as roughly as you like, and **throw it away**. Then drive the real implementation with tests, using what you learned to write the assertion you could not write before. The failure mode is not doing the spike; it is keeping the spike code once it works, at which point untested exploratory code has become the feature. **Unfamiliar external interfaces.** You do not know what the other side returns, so any test you write encodes your guess. Probe first, then encode the observed behaviour as a test of your own adapter around it. **Work whose oracle is a human.** Visual layout, wording, the feel of an interaction. There is no assertion that says "this looks right". Automate the parts that do have oracles and leave the judgement to a person. ### Where the second precondition fails **Performance and capacity work.** The expected result is a distribution, not a value. A test that asserts a hard threshold on a measured duration flips between pass and fail without any code change, and a suite that cries wolf is worse than no suite. Measure with the tools built for measurement, hold the numbers as trend data, and reserve assertions for the coarse regressions you would actually act on. **Long feedback loops.** Anything where the smallest honest verification takes minutes — a full environment provisioning, an end-to-end flow through several systems. TDD is not banned there, but the cycle you run is a different, slower one, and the fine-grained driving happens at a level where the loop is fast. ### A worked example A four-person team on a document e-signing flow needed to know whether a third-party signing service returned per-signer completion times or only an envelope-level timestamp. That is a spike: two hours timeboxed, one throwaway script, question answered. What they then drove with tests was their own fee-splitting logic — where the expected result was perfectly statable in advance, and where a currency-rounding drift of two hundredths per envelope was exactly the kind of defect a before-the-code assertion catches. Mixing those two pieces of work under one policy would have been wrong in both directions: tests around the spike would have tested their own guesses, and a spike around the fee logic would have shipped arithmetic nobody had specified. ### The honest caveats Say plainly that the empirical evidence for the practice is contested: studies disagree on whether it reduces defect density, and the design-quality claims are largely experience reports rather than settled measurement. That does not make it a bad default; it makes confident numbers a red flag. Say equally plainly that "this does not fit" is the most abused sentence in this area. The distinguishing move is that a real exception is **named, timeboxed and has a re-entry point** — "we spike until Thursday, then the real implementation is driven normally" — while an abused one is a standing exemption for a whole subsystem. Concurrency, data transformation and anything described as "too dynamic" are usually not exceptions; they are areas where the assertion is harder to write and therefore more valuable once written. ### What good sounds like A strong answer starts from preconditions rather than a list of forbidden categories, gives one or two concrete cases with the discipline attached (timebox, throw away, then drive), refuses to turn an exception into a policy, and is candid that the evidence base is mixed. A weak answer either insists the practice applies everywhere without exception or treats one awkward area as licence to abandon it wholesale.
- After a spike answers your question, what exactly do you keep?The knowledge, and occasionally a short note or a single test that encodes what you observed about the other side. The code goes. Keeping it is the failure mode: exploratory code written without an expected result becomes the feature, and nobody ever goes back to specify it. If part of the spike is genuinely worth keeping, rewrite that part deliberately with tests rather than promoting the draft.
- How do you stop "this does not fit here" becoming a permanent exemption for a subsystem?Require every exception to be named, timeboxed and to have a re-entry point — spike until a stated moment, then the real implementation is driven normally. Exemptions that apply to a whole area indefinitely are policy, not judgement. Be suspicious of areas excused for being hard: concurrency, data transformation and anything called too dynamic are usually where a pre-written assertion is worth most.
- Someone claims the practice cuts defects by a specific percentage. How do you respond?Treat the number as unsupported. The empirical literature disagrees on whether the practice reduces defect density, study designs vary widely, and the design-quality claims are largely experience reports rather than settled measurement. The honest position is that the mechanism — a specification before the code, and a safe moment to restructure — is demonstrable, while headline percentages are not.
saying these in an interview costs you the question
- Claims the practice applies to every task without exception
- Uses "this is exploratory" as a standing exemption
- Keeps spike code in production once it works
- Asserts a hard timing threshold and blames the flakiness on luck
- Quotes a firm defect-reduction percentage as settled fact
- Says concurrency is simply untestable