A unit test replaces every collaborator with a stand-in and still passes while the feature is broken. What went wrong?
answer
- Ask what this test could catch
- Every answer was authored by the test
- Wiring confirmed, behaviour untouched
- Break the logic and watch for red
- Move the boundary to the process edge
basics
~20 sThe test is tautological: every value the code received was authored by the test, so it confirms only that the code calls the stand-ins it configured. How the real components combine was never exercised, and the defect lives exactly there.
solid answer
~50 sWhen every collaborator is replaced, the only propositions a pass supports are that the code called what the test told it to call and returned what the stand-ins were told to give it. Nothing about the assembled system was exercised, so the defect can sit untouched in the joins: mismatched assumptions between two components, ordering, caching, state carried across calls. The diagnostic is to ask whether the code could be wrong in a plausible way and this test still pass; if yes, the test has no oracle. The fix is to move the substitution boundary outward, keeping stand-ins for the process edge and nondeterministic sources while letting in-process collaborators the team owns run for real, and to add a level where those real collaborators meet. Then prove it by deliberately breaking the logic and confirming a test goes red.
code
pseudocode · 24 lines# Tautological: the answer is supplied, then asserted
test "summary uses current brackets":
cache = standIn(RulesCache)
calculator = standIn(BracketCalculator)
cache.current() returns RuleSet(revision = 7)
calculator.liabilityFor(Amount("41_320.75"), RuleSet(revision = 7))
returns Amount("5_764.15")
assert FilingWizard(cache, calculator).summarise(Amount("41_320.75")).liability
== Amount("5_764.15") # cannot fail for the right reason
# Boundary at the process edge and the clock only
test "republished rules invalidate the cached snapshot":
clock = AdvanceableClock("2026-02-17T09:00")
source = standIn(RulesSnapshotSource) # out-of-process
cache = RulesCache(source, clock, ttl = minutes(45)) # real
wizard = FilingWizard(cache, BracketCalculator()) # real
source.forTaxYear(2025) returns RuleSet(revision = 6)
wizard.summarise(Amount("41_320.75"))
source.forTaxYear(2025) returns RuleSet(revision = 7)
publishNotice(year = 2025)
assert wizard.summarise(Amount("41_320.75")).rulesRevision == 7go deeper
Recall the core idea: if the test supplied every answer the code received, the test cannot discover that the code is wrong. Get used to asking, before you commit a test, what would have to break for it to go red.
Explain the mechanism and the drift into it: one stand-in per constructor parameter, substitution used as a shortcut for object construction, and a long arrangement that nobody wants to lengthen further. Describe moving the boundary out to the process edge.
Diagnose a real suite out loud. Point at the arrangement-to-assertion ratio, name the defect classes that concentrate in the joins, and demonstrate the deliberate-break check. Be able to say which stand-ins you would keep and why.
Own the systemic version: a default that produces this pattern makes a green pipeline meaningless, and the metrics teams watch tend to reward it. Talk about how you would make the boundary a reviewed decision and how you would measure escaped defects by seam.
### The failure has a name and a mechanism A test in which every collaborator has been replaced is **tautological**: every value the code under test receives was authored by the test itself, so the only propositions the test can verify are "the code called the things I told it to call" and "the code returned what I told the stand-ins to give it". The production behaviour — how the pieces actually combine — was never present in the run. A green result therefore proves that the wiring in the test matches the wiring in the code, which is a statement about the test, not about the system. The diagnostic question is short: **could the code be wrong in a plausible way and this test still pass?** If yes, the test has no *oracle* — nothing in it can tell correct behaviour from incorrect behaviour — and the pass carries no information. ### Why suites end up here Rarely by decision. It accumulates: - **One stand-in per constructor parameter**, applied mechanically, because that is the arrangement the last test used. - **Substitution as a shortcut for construction.** The real collaborator needs setup, and replacing it is one line, so it gets replaced even though it is fast, deterministic and in-process. - **Fear of arrangement.** Once a test file's setup is long, nobody wants to make it longer with a real object graph, so the next collaborator is doubled too. - **A suite judged by the wrong signal.** Test counts, pass rates and executed-line percentages all go *up* as substitution increases, so the metric that gets watched rewards the defect. ### What it costs False confidence is the headline cost, but not the only one. A fully doubled test is coupled to the call structure of the code, so a refactor that preserves behaviour breaks tests — which teaches the team that tests obstruct change. And escaped defects concentrate in exactly the joins the suite stopped exercising: mismatched assumptions between two components, ordering, caching, state carried across calls. ### The incident that makes it concrete A **tax-filing wizard** had 214 unit tests, all green, and a filing-summary component whose in-process collaborators — a bracket calculator and a rules-snapshot cache — were both replaced in every one of them. The cache's stand-in returned a fresh snapshot on demand, always. The real cache kept an entry until its own eviction rule expired it, and that rule did not consider a mid-year republication of the rules. In production the wizard performed a **stale-cache read** and computed liabilities from the previous year's brackets for about 3 days before a user noticed. The fix was one line in the cache's invalidation. The interesting part is that no unit test could have failed: the component that read the cache and the cache that served it never met in the suite. The defect then rode a **3-week release train**, so the interval between writing the bug and hearing about it was measured in weeks. The test that would have caught it is unremarkable — the summary component running against the *real* cache with a controlled clock, republishing the rule set mid-scenario. Only the clock and the out-of-process rules service needed replacing. ### How to fix it, and how to check the fix 1. **Move the boundary outward.** Substitute at the process edge and for nondeterministic sources; let in-process collaborators the team owns run for real. 2. **Add a level where real collaborators meet.** If the fast tests must keep their doubles, then some test somewhere has to exercise the assembled components; otherwise no test does. 3. **Prove the oracle by breaking the code.** Introduce a deliberate fault — invert a condition, return a stale value on purpose — and confirm a test goes red. A suite that stays green under a deliberate break is telling you precisely what it is worth. 4. **Read the arrangement-to-assertion ratio.** A test with thirty-odd lines of setup and two lines of assertion, where the assertions restate the setup, is the signature of this defect. ### The balance to strike, out loud The conclusion is not that stand-ins are bad. A suite with no substitution at all is slow, flaky and dependent on other teams' uptime, and the clock and the network genuinely have to go. The claim is narrower and worth stating in an interview in one sentence: **a stand-in is a hole in the evidence, so put the holes where you cannot afford the real thing and nowhere else** — and be able to say, for any test you have written, what would have to be wrong in the system for it to fail.
- How do you prove to yourself that a test actually has an oracle?Break the code on purpose and watch the test. Invert a condition, return a stale value, drop a rounding step, then run the test that claims to cover it. If it stays green, the test is confirming wiring rather than behaviour, and you have learned exactly what it is worth. Doing this deliberately for a handful of representative tests tells you more about a suite than any executed-line percentage.
- The network client genuinely has to stay substituted. Where do you add a test that would catch this defect class?At the level where the real in-process collaborators meet, with only the process edge and the clock replaced. Assemble the components as production does, drive a scenario that crosses them, and assert on observable behaviour rather than on calls. You keep the speed benefit of substituting the remote dependency while restoring the evidence about how your own pieces combine, which is where these defects live.
- Is the answer simply to use fewer stand-ins everywhere?No, and saying so is a trap. A suite with no substitution is slow, depends on other teams' uptime, and cannot reach failure paths or control time, so the clock and the network genuinely have to go. The claim is about placement, not quantity: a stand-in is a hole in the evidence, so put the holes where the real thing is unaffordable, and be able to say for any test what would have to be wrong for it to fail.
It is a fire drill where the alarm, the smoke and the exits are all played by volunteers reading from the same script the inspector wrote: everyone performs correctly and the building's actual doors are never tried.
saying these in an interview costs you the question
- Treats a passing test as proof the behaviour is verified
- Responds to an escaped defect by adding more stand-ins
- Cannot say what change would make the test fail
- Judges the suite by test count or executed-line percentage
- Concludes all stand-ins are bad and deletes them wholesale
- Blames the escape on missing manual checks rather than the boundary