Mockito's own documentation warns that RETURNS_DEEP_STUBS usually means something is wrong with the design. What is the underlying argument, when would you accept deep stubs anyway, and what would you reach for first instead?
answer
- couples the test to graph shape, not behaviour
- every link is fiction — no real contract exercised
- Law of Demeter signal, papered over
- first: inject the leaf, use real value objects
- OK for third-party fluent APIs and legacy characterization
basics
~20 sA deep stub encodes a call chain, so the test depends on the shape of an object graph rather than on behaviour — refactoring the intermediate types breaks tests that never changed meaning. Accept it for third-party fluent APIs and legacy characterization; otherwise inject the leaf collaborator, use real value objects, or hide the chain behind an adapter you own.
solid answer
~60 sThe argument is coupling. `when(a.getB().getC().doIt()).thenReturn(x)` bakes the traversal `a → B → C` into the test. Nothing about that path is behaviour anyone cares about, so any refactor of `B` or `C` breaks tests whose meaning did not change. Worse, every intermediate is a mock: no real contract is exercised, so the test can stay green while the real graph cannot produce that shape at all. It is the Law of Demeter complaint, made concrete by test pain. First alternatives, roughly in order: **inject the leaf** — if the class only needs `C`, pass `C`, not `A`; **use real objects** for the value-ish links, since a `Config` or a DTO is usually cheap to construct; **build a test data builder** for graphs you need often; **wrap third-party chains in a thin adapter you own** and mock the adapter; and for fluent builders that return `this`, `Answers.RETURNS_SELF` is the targeted tool. Accept deep stubs when the chain belongs to something you cannot change — SDK clients, servlet-style APIs, configuration trees — or in characterization tests around legacy code you are about to refactor. Then treat each usage as temporary and reviewable.
go deeper
Say that deep stubs mean the test knows too much about how objects are chained, and that passing the needed object in directly is usually simpler.
Name the Law of Demeter link and the refactor-breakage cost, and give injection and real value objects as the alternatives.
Weigh it: acceptable against third-party or legacy code, with an adapter when the chain recurs, plus the point that mocked intermediates verify nothing real.
Frame it as a design metric and a boundary policy — allowed only for types you do not own, visible in review, trending down as production navigation is fixed.
## The argument against Three distinct costs hide behind the convenience. **Coupling to structure, not behaviour.** A test exists to pin down what a unit does. A deep stub pins down how it *navigates*: `session.getUser().getAccount().getBalance()`. If someone later gives `Session` a `getBalance()` shortcut, or moves the balance to a different type, the test breaks even though the observable behaviour is unchanged. Tests that fail on refactors they should not notice are the reason suites get abandoned. **Everything in the middle is fiction.** Each auto-created link is an empty mock. No constructor ran, no invariant was checked, no null-handling in the real `Account` was exercised. The test asserts that *given a graph that may not be constructible in production*, the unit does something. Deep chains are exactly where that fiction gets thickest, because you never wrote the intermediate objects down and so never noticed how implausible they are. **It hides a design signal.** Needing three links usually means the class asks a collaborator for a collaborator. That is the Law of Demeter violation: the unit depends on the transitive shape of a graph rather than on the one thing it needs. Deep stubs make the pain go away without addressing the dependency, which is precisely why Mockito's javadoc flags it. ## What to do first 1. **Inject the leaf.** If the method only reads a balance, pass `BigDecimal balance` or an `AccountReader`. The test then has one mock, or none. This is the highest-value fix and usually the smallest diff. 2. **Use real objects for data.** Value objects, DTOs, configuration holders and records are cheap. `new Config(new Server("host", 8080))` is clearer than a deep stub, never lies about constructibility, and gives better failure messages. 3. **Introduce a test data builder** when constructing the graph is genuinely tedious but real objects are the right call. One builder amortises across dozens of tests. 4. **Wrap what you do not own.** For a third-party fluent client, write a one-method adapter interface in your code, implement it with the real chain, and mock the adapter everywhere else. The chain then appears in exactly one place and is covered by one integration test. 5. **Use the targeted tool for builders.** If the chain exists only because a builder returns `this`, `Answers.RETURNS_SELF` expresses that directly, without turning the whole return graph into mocks. ## When deep stubs are the right answer - **Third-party fluent APIs you cannot change**, where writing an adapter is disproportionate to a single test — SDK client builders, servlet/JNDI-style request objects, deeply nested configuration trees. - **Characterization tests around legacy code.** You need the existing code under test *before* you refactor it; deep stubs get it into a harness quickly, and the refactor then deletes them. - **Genuinely structural interfaces** whose whole point is navigation — a parsed document tree or an AST — where the chain is the domain and mocking it mirrors reality rather than distorting it. ## Governing it in a codebase At scale the useful move is not a ban but visibility. Treat `RETURNS_DEEP_STUBS` as a review trigger and, if you like metrics, count usages over time: rising counts mean the production code is growing train-wreck navigation, which is a design report, not a test-hygiene report. Pair that with a rule that deep stubs are acceptable only against types you do not own — an easy line to explain and to check. And be honest about the trade: a team drowning in legacy will get more value from deep stubs plus a refactoring plan than from a purity rule that keeps the code untested. The answer an interviewer wants is not "never use them". It is that you can name the cost precisely, you reach for injection and real objects first, and when you do use deep stubs you know why the exception is justified and how it will end.
- How would you remove a deep stub from a test without changing the behaviour under test?Look at what the unit actually consumes at the end of the chain and make that a direct dependency — constructor parameter or method argument — so the unit no longer navigates. Where the intermediate types are plain data, construct them for real instead. Both are behaviour-preserving refactors, and the test usually shrinks to one mock or none.
- When is wrapping a third-party fluent API in your own adapter worth the extra type?When the chain appears in more than one place, or when you want the shape of the third-party API to stop leaking into your unit tests. The adapter is a one-method interface you own, mocked trivially everywhere, with a single integration test covering the real chain. For a chain used once in a single test, the adapter is usually not worth it.
- Is RETURNS_SELF a substitute for deep stubs?Only for fluent builders. RETURNS_SELF returns the mock itself for methods whose return type is compatible with the mocked type, which matches `builder.a().b().c()` patterns. It does not help when the chain walks through genuinely different types, which is what deep stubs are for.
It is like testing a courier by faking the whole city — every street, sign and door. The delivery succeeds in your model even if the city could never be laid out that way.
saying these in an interview costs you the question
- Treating deep stubs as a neutral convenience with no design implication
- Insisting on an absolute ban, even for third-party APIs you cannot change
- Reaching for a deep stub before considering injecting the leaf collaborator
- Claiming a deep-stubbed graph proves the real object graph can produce that state
- Mocking cheap value objects instead of just constructing them