What are the practical signs that a Mockito test is over-mocked — for example a mock whose method returns another mock, or a test that fails on every refactoring even though behaviour did not change — and how do you fix each one?
answer
- mock returning mock = Law of Demeter breach
- deep stubs are a hint, not a tool
- pass the leaf, not the graph
- verify only what has no return value
- verifyNoMoreInteractions everywhere = brittle
basics
~20 sTwo big smells: mocks returning mocks (or RETURNS_DEEP_STUBS), meaning the test knows an object graph the code should not navigate; and mirror-the-implementation tests that stub and verify every call, so any refactoring breaks them. Fix by passing the leaf object the code needs, and by asserting outcomes instead of every interaction.
solid answer
~50 sThe first smell is `given(a.getB()).willReturn(bMock)` chains, or `mock(A.class, RETURNS_DEEP_STUBS)`. It means the code navigates a graph — `a.getB().getC().doThing()` — so the test must know that graph, and any change to the traversal breaks it. Mockito's own documentation treats deep stubs as a hint that something is wrong with the code under test. The fix is in production code: pass the leaf collaborator or the value the method actually needs, or add a method on the direct collaborator that answers the question. The second smell is a test mirroring the implementation: every call stubbed, every call verified, `verifyNoMoreInteractions` everywhere, sometimes `InOrder` for calls whose order is irrelevant. Such a test asserts *how* the code works, so a rename or a reordering fails it while behaviour is intact — and it still will not catch a real behaviour change. The fix is to assert the outcome: the returned value, the observable state, and only those boundary interactions that are themselves the effect (an event published, a mail sent).
code
java · 9 lines// smell: the test knows the object graph
Order order = mock(Order.class, RETURNS_DEEP_STUBS);
given(order.getCustomer().getAddress().getCity()).willReturn("Berlin");
// better: real value objects, zero mocks
Order order = new Order(new Customer(new Address("Berlin")));
// best: the method takes what it actually needs
shipping.quoteFor("Berlin");go deeper
Recognise the two shapes and give the direction of the fix: fewer mocks, assert the result rather than each call.
Explain why each smell breaks under refactoring and name the alternatives: pass the leaf object, build real value objects, verify only side effects.
Diagnose the design cause — graph navigation, orchestration tangled with logic — and describe rewriting the test from a one-sentence statement of the behaviour.
Talk about suite-level economics: brittle interaction tests raise refactoring cost and train teams to update tests reflexively, eroding the signal the suite exists to provide.
## Smell 1: a mock that returns a mock When a test contains `given(order.getCustomer()).willReturn(customerMock)` and then `given(customerMock.getAddress()).willReturn(addressMock)`, or creates `mock(Order.class, RETURNS_DEEP_STUBS)` so that `order.getCustomer().getAddress().getCity()` works without explicit chaining, the test is describing an object graph. Why it hurts: the test now depends on the exact path the production code walks. Change `order.getCustomer().getAddress()` into a delegating `order.shippingCity()` and the test breaks though nothing observable changed. It is also hard to read — three mocks exist purely to deliver one string — and it hides a design problem. The chain is a Law-of-Demeter violation: the class reaches through collaborators to their collaborators, coupling itself to types it should never have known about. Mockito documents deep stubs as a feature whose need should make you suspicious of the code, not of the test. Fixes, in order of preference: pass what the method actually needs (`quoteFor(String city)` or `quoteFor(Address address)`) so no graph is required; give the direct collaborator a method that answers the question (`order.shippingCity()`), leaving one stubbing; or, when the graph is made of value objects, build them for real — `new Order(new Customer(new Address("Berlin")))` needs no mocks at all. Deep stubs are a last resort for legacy code you cannot restructure yet. ## Smell 2: the test mirrors the implementation The second pattern is a test whose body reads like a transcript of the method: every collaborator call stubbed, every call verified, often `verifyNoMoreInteractions(...)` on all mocks and an `InOrder` covering calls whose order carries no meaning. Such a test encodes *how* the method is written rather than *what it does*. The cost shows at every refactoring: extracting a helper, replacing two repository calls with one batched call, caching a lookup, or reordering independent statements all fail the test although the system behaves identically. Teams then update tests mechanically to match the code, which destroys the suite's signal — a genuine regression gets 'fixed' the same way as a refactor. The converse cost matters too: over-verification adds no safety. Asserting that a getter was called three times says nothing about correctness, while the actual outcome may go unchecked. Fixes: assert the result of the call and the state the caller can observe. Verify interactions only where the interaction *is* the outcome — an event published, a payment charged, a row saved, a mail sent — because those have no return value to assert. Verify absence deliberately (`verify(gateway, never()).charge(any())`) when "must not happen" is the rule under test. Reserve `verifyNoMoreInteractions` for the rare test whose contract genuinely is "nothing else happened"; used routinely it turns every added log line or extra read into a failure. Use `InOrder` only when order is part of the contract, such as acquiring before releasing a lock. ## The common root Both smells come from the same place: the test was written to satisfy the code rather than the behaviour. A useful rewrite prompt is to state the test's intent in one business sentence — "over the limit, no charge" — and keep only the stubbings and verifications that sentence needs. Everything else is scaffolding. ## Related signals worth naming Other indicators of over-mocking: more mocks than the class has responsibilities; stubbing whose return value nothing consumes (strict stubbing will report it); mocking cheap value objects to shorten setup; a test with no assertion at all, only verifications; and setup blocks duplicated across many tests because they re-describe the same implementation. ## What good looks like A few boundary mocks, real domain objects, one act step, assertions on the returned value or observed state, and one or two verifications of side effects that cannot be observed any other way. Such a test survives refactoring and fails only when behaviour changes — which is the whole point of having it.
- When is verifying an interaction the right assertion rather than a smell?When the interaction is the observable outcome and has no return value to assert: an event published, a message produced, a mail sent, a payment charged, a row deleted. Also when the rule under test is that something must not happen, where verify with never() is the only way to express it.
- Mockito ships RETURNS_DEEP_STUBS — if the library provides it, why treat it as a smell?It exists mainly for legacy code you cannot restructure, and Mockito's own documentation says needing it suggests the code under test violates the Law of Demeter. It couples the test to a traversal path and silently produces mocks nobody configured, so it should be a stopgap rather than a default.
Verifying every internal call is like grading a chef on which drawers they opened rather than on the dish. Rearrange the kitchen and the score collapses while the food is identical.
saying these in an interview costs you the question
- Treating deep stubs as a convenience feature rather than a warning about the design
- Adding verifyNoMoreInteractions to every test 'for completeness'
- Using InOrder for calls whose order has no business meaning
- Tests containing only verifications and no assertion about the result
- Updating mirrored tests mechanically after each refactor instead of removing the coupling