How do you pick the cheapest test double that still makes a test fail for the right reason?
answer
- Start from the oracle, not the double
- Cheap means fewer ways to fail
- No value reveals it? Then the call is the oracle
- Consistency across many calls favours a working implementation
- Break the behaviour and confirm red
basics
~20 sStart from the oracle — what decides that the behaviour is wrong. A returned value means a stub suffices; a collaboration that is itself the requirement needs a mock or a spy; many interdependent calls need a fake.
solid answer
~50 sWork backwards from the oracle — the thing that decides the behaviour is wrong — and ask what must be true for this test to fail *only* when the named behaviour breaks. If the verdict rests on a returned value or resulting state, the cheapest double that reaches it is a **stub**; anything richer adds ways to fail without adding coverage. If the behaviour under test *is* a call the code makes, no value can catch its absence, so you need a **mock** or a **spy**. If the code drives a long run of interdependent calls whose consistency matters, a **fake** costs more once and far less per test. And an untouched collaborator takes a **dummy**, which documents its irrelevance. Both extremes bite: over-specified doubles fail on changes that broke nothing, under-specified ones stay green while the behaviour is gone.
go deeper
Learn the default: if your assertion is about a returned value, a stub is enough. Resist adding checks on how many times a collaborator was called simply because the tooling makes it easy.
Be able to walk from a stated requirement to a chosen kind and justify it out loud. Show that you know which choices add ways to fail that have nothing to do with the behaviour under test.
Demonstrate the discipline, not the vocabulary: name the oracle, pick the least capable double that can still go red, then prove it by breaking the behaviour deliberately. Interviewers want the over- and under-specification failure modes described from experience.
Frame it as suite economics. You should be able to say what a codebase of over-specified doubles costs in engineer-hours per refactor, and how you would shift the default without writing a rule nobody can apply in review.
### The question behind the question "Cheapest" is not about typing. It is about the total cost a double imposes: the setup a reader must digest, the ways it can fail that have nothing to do with the behaviour under test, and the maintenance it demands when the code is restructured. "Fails for the right reason" is the constraint that keeps cheapness honest — a test that never fails is free and worthless. So the selection procedure starts at the **oracle**: the thing that decides the observed behaviour is wrong. Name the oracle in one sentence, and the right kind usually falls out. ### The procedure 1. **Is the collaborator involved at all?** If the path under test never touches it, pass a **dummy**. That is not laziness; it is documentation. A configured stub on an unused collaborator tells the next reader it matters, and they will preserve it forever. 2. **Does the oracle rest on a value or on resulting state?** Then a **stub** supplies whatever input drives the path, and the assertion carries the verdict. This is the common case and should be the default. 3. **Is the collaboration itself the requirement?** When the observable effect leaves no trace in any returned value — a record refreshed, a message dispatched, an audit line written to a collaborator you own — the only oracle available is the call. Use a **mock** with an expectation, or a **spy** the test interrogates afterwards. Prefer a spy when you want the failure message and the assertion to sit in the test body; prefer a mock when the expectation genuinely reads better up front. 4. **Does the run depend on state accumulating across many calls?** Ten chained canned answers that must stay mutually consistent are a bug factory. A **fake** with a real, if simplistic, implementation is more work once and much less work per test thereafter. It also survives reordering, which a positional chain of answers does not. ### Working the parcel-tracking example The rule: a status request that exceeds the team's 92nd-percentile budget of 340 ms falls back to the last cached status, and any parcel whose cached status is older than 45 minutes is refreshed before it is reported. - *"An intermittent timeout falls back to cache."* The oracle is the returned status. A **stub** that raises a timeout on the call is exactly sufficient. Adding an expectation that the gateway was called once buys nothing — the fallback could not have happened without the call — and buys a failure the day someone adds a retry. - *"A cache entry older than 45 minutes is refreshed first."* No return value distinguishes a refreshed entry from a stale one that happened to hold the same status. Here the call *is* the oracle, so a **spy** the test inspects, or a **mock** holding the expectation, is the cheapest thing that can fail correctly. - *"A batch of 3,180 parcels is reconciled and the ones that changed are re-reported."* Canned answers for 3,180 identifiers, in order, is unmaintainable and fragile. A **fake** gateway holding a dictionary of parcels lets the test seed 12 changed entries and assert on the outcome. - *"The formatter renders a delayed status with an apology line."* The gateway is not touched. Pass a **dummy**. ### The two failure modes to name out loud **Over-specification.** Expectations on calls that were never the requirement turn the current implementation into the spec. The test then fails on caching, batching, retry or extraction changes that broke nothing — and every one of those failures costs an engineer the time to prove it was harmless. A suite that cries wolf gets rewritten to be quiet, usually badly. **Under-specification.** The mirror image: the test only checks a value that would have been produced anyway, so the behaviour it claims to cover could be deleted and the test would still pass. The classic case is asserting a returned summary while the requirement was the dispatch that never happened. Whenever you choose a cheaper double, check this by breaking the behaviour on purpose and confirming the test goes red. If it does not, you saved cost by removing the coverage. ### How to say it in an interview Say that you choose per test, name the oracle first, and default to the least capable double that can still go red for the named reason. Then admit the check: a deliberately broken implementation must fail the test. That last sentence is what separates an answer about vocabulary from an answer about engineering.
- How do you check that a cheaper double did not quietly remove coverage?Break the behaviour on purpose and re-run: change the implementation so the named requirement is violated and confirm the test goes red for that reason. If it stays green, the double is under-specified and the test is decorative. Doing this once when the test is written costs a minute and is the only direct evidence the test can fail correctly.
- When would you accept the cost of building a fake rather than stubbing?When several tests drive long, order-sensitive sequences against the same collaborator and the canned answers have started to encode a script rather than a scenario. The tipping point is usually maintenance: if changing one step forces edits in many unrelated tests, a working in-memory implementation pays back quickly, and it is shared code with a real owner from that point on.
- Is a spy ever preferable to a mock when the collaboration is the requirement?Often, yes. A spy keeps the assertion in the test body next to the other assertions, so the failure reads like every other failure and the reader is not hunting for an expectation declared 30 lines earlier. A mock is preferable when stating the expectation up front genuinely documents the scenario better than a trailing assertion would.
saying these in an interview costs you the question
- Using an expectation on every collaborator by default
- Choosing the double before naming what decides failure
- Treating a fake as always more expensive than stubs
- Never checking the test can actually go red
- Configuring a collaborator the path never calls
- Equating cheap with less code to type