Why does outside-in TDD leave more test doubles behind than inside-out TDD?
answer
- It follows from the order things get built
- The collaborator does not exist yet
- Nobody comes back to swap it out
- A stand-in freezes an imagined contract
basics
~20 sOutside-in starts before the collaborators exist, so each invented role needs a stand-in for the test to run — and nothing forces a later swap for the real object. Inside-out builds collaborators first and uses them directly.
solid answer
~40 sIt falls out of the build order. An outside-in test is written at a boundary whose collaborators have not been built, so every invented role needs a stand-in just to make the test runnable, and each step inward adds another layer of them. The outer tests keep those stand-ins after the real collaborator exists, because nothing in the cycle prompts a swap and the real object is often slower or needs infrastructure. Inside-out reverses the order: collaborators are already built and tested when their caller is written, so stand-ins appear only where reality is genuinely unusable in a test. The cost of a large stand-in population is twofold — tests become coupled to call structure, and each stand-in freezes a contract you imagined, which can drift from the one the real collaborator honours.
code
pseudocode · 11 linestest "handler stores a six decimal reading":
store = standIn(ReadingStore)
handler = IngestHandler(store)
handler.ingest(reading(kwh = 3421.874416))
verify store.save(reading(kwh = 3421.874416)) was called
class RealReadingStore implements ReadingStore:
save(reading):
writeColumn(numeric(precision = 10, scale = 2), reading.kwh)go deeper
Recall the mechanical reason first: at the boundary the collaborators have not been written yet, so something has to take their place for the test to run at all. That single fact explains most of the difference.
Explain the recursion — each step inward adds another layer of stand-ins — and why the outer tests keep theirs afterwards. Be ready to name the two costs: coupling to call structure, and a frozen contract that can drift.
Show how you keep a large stand-in population honest in practice: narrow roles, re-pointing cheap tests at real collaborators, and at least one thing that exercises both sides of every seam so two green suites cannot silently disagree.
Own the policy question. Decide what the team standardises — which seams must be verified against the real implementation, and what evidence a release needs beyond unit-level green — and be able to justify the run-time cost that buys.
### Why the stand-ins pile up The accumulation is not sloppiness — it is a direct mechanical consequence of the order in which outside-in TDD builds things. Outside-in starts at an outer boundary and writes a test for behaviour visible there. That test needs the object under test to be constructible and runnable *now*, before its collaborators exist. The only way to get there is to invent each collaborator's role and put a stand-in in its place. Each collaborator you invent is one stand-in in that test. Then you move inward: the role you invented becomes the next subject, and it in turn needs collaborators that do not exist yet, so it gets its own stand-ins. Every layer you descend multiplies rather than replaces the previous layer's scaffolding. Crucially, the outer test **keeps** its stand-ins after the real collaborator is built. Nothing in the cycle forces you back to swap them out, and there are usually good reasons not to: the real collaborator may be slow, may need infrastructure, or may drag half the object graph into a test that was meant to be about one behaviour. So the count only grows. Inside-out inverts the order. The core is built first with real objects, and each outer object is written over collaborators that already exist and already pass their own tests. A stand-in is introduced only where reality is genuinely unusable in a test — a network call, the wall clock, a device, a random source — so the steady-state population is small and every one of them is justified by an environmental constraint rather than by build order. ### What the accumulation costs **Coupling to call structure.** With no real collaborator to inspect afterwards, an outside-in test frequently asserts on the calls made rather than on a resulting value or state. Those assertions describe the internal shape of the design. Rename the role, merge two collaborators, push a decision one level down — behaviour identical, suite red. The suite then resists exactly the restructuring the third step of the cycle is supposed to make cheap. **Contract drift, and the failure it hides.** A stand-in agrees with the contract you imagined when you invented the role. Nothing keeps that in step with the contract the real collaborator ends up honouring. On the smart-meter ingest path above, the handler was driven outside-in against a stood-in reading store, and every test asserts the handler passes a reading of `3421.874416` kWh through unchanged. The real store, written later by someone else on the eleven-person team, persists that column at two decimal places. Both sides have green tests. The integrated system silently rounds every reading, and the corruption is invisible until a monthly reconciliation is short by a fraction of a percent. No test failed, because no test ever ran the two together. **Comprehension cost.** Reading a test whose arrangement is six stand-ins and their programmed responses tells you what the object *calls*, not what it *does*. New joiners learn the wiring before they learn the behaviour. ### How practitioners keep it under control - **Re-point the outer tests once the real collaborator exists**, at least for the cheap ones. The stand-in was scaffolding for a build order, not a permanent fixture. - **Keep every invented role narrow.** A role with one method and one caller produces a stand-in that is trivial to keep honest; a wide role produces a fiction. - **Verify each seam against the real implementation.** Whatever the mechanism — an integrated test through the real chain, or a paired test that checks the stand-in's assumptions against the real object's behaviour — the point is that *something* exercises both sides of a contract that two green suites otherwise never compare. - **Assert on outcomes wherever an outcome exists.** Reach for a call-level assertion only when the effect is genuinely invisible — a message sent, a record deleted. - **Never stand in something you own and can build cheaply.** A pure value object or a small calculation should be the real thing even in an outside-in test. The honest summary for an interview: outside-in trades a larger permanent population of stand-ins, and tests that know more about structure, for early feedback on the outer design and for seams that let several people work in parallel. Whether that trade is worth it is a per-context judgement, not a rule.
- What would you do to stop a stood-in seam from drifting away from the real collaborator?Make something exercise both sides. Either run at least one test through the real chain for that seam, or pair every set of programmed responses with a test that checks those same assumptions against the real object. The failure this prevents is two independently green suites that disagree about a contract neither of them ever compared.
- Is it acceptable to leave a stand-in in place permanently?Yes, when the real collaborator is genuinely unusable in the test — remote calls, the clock, a device, randomness — or when swapping it in would drag a large object graph into a test about one behaviour. What is not acceptable is leaving it there by inertia while the real object sits in the same codebase, unexercised by any test that also runs the caller.
- Why do outside-in tests tend to assert on calls rather than on values?Because at the moment the test is written there is no real collaborator whose state you could inspect afterwards, and often no return value either — the effect is something that happened to another object. So the only observable left is the call itself. Where an outcome does exist, assert on the outcome instead.
saying these in an interview costs you the question
- Says the extra doubles are just undisciplined test writing
- Claims a stand-in cannot disagree with the real collaborator
- Thinks more doubles automatically means better isolation
- Cannot name a legitimate permanent stand-in, such as a clock
- Says integrated tests are unnecessary once every unit test is green