skip to content

Should a team standardize on Mockito's do-family stubbing form everywhere for consistency, or keep when(...).thenReturn(...) as the default and use the do-family only where required? Argue your position.

level: principalimportance: nice to knowfreq 22%

answer

  1. typed thenReturn vs untyped doReturn(Object)
  2. compile failure vs runtime WrongTypeOfReturnValue
  3. natural reading order = when → then
  4. do-family presence signals spy/void/generics
  5. do-family density = spy density = design signal

basics

~20 s

Keep when(...).thenReturn(...) as the default. It is compile-time type-checked and reads in the natural order. Use the do-family only where it is required — spies, void methods, awkward generics, doCallRealMethod — so its presence signals a special case rather than a style choice.

solid answer

~60 s

I would not standardize on the do-family. The consistency argument is real but weak next to what it costs. `when(mock.get()).thenReturn(v)` is typed: the compiler infers the method's return type, so a wrong value or a later return-type change fails the **build**. `doReturn(Object)` cannot check anything; a mismatch becomes a runtime `WrongTypeOfReturnValue`, and a stale stub after a refactor may not fail at all until that test runs. Trading compile-time feedback for visual uniformity is the wrong direction. The reading order matters too: "when this is called, return that" matches how people describe behaviour; the inverted form buries the interesting part at the end of the line. There is also a signalling benefit in *not* unifying. If `doReturn` only appears where it is necessary, seeing it tells a reader "this is a spy, or void, or a generics corner" — free information the reviewer would otherwise have to reconstruct. So: convention by requirement, not by style, and a lot of `doReturn` in a codebase is a spy-density warning worth acting on.

go deeper

for a junior

Know both forms exist and that when(...).thenReturn(...) is the documented default; leave the policy call to others.

for a middle

State the trade concretely — compile-time checking and reading order versus universality — and list the cases that force the do-family.

for a senior

Write the policy: default plus required exceptions, with the absolute rule that spies are never stubbed via when(spy.call()), and explain why enforcement by review beats a blanket lint rule.

for a principal

Treat the syntax ratio as a diagnostic for spy density and mixed responsibilities, be explicit that conventions can be time-boxed responses to a suite's current weakness, and connect the decision to feedback latency across a large test suite.

## Framing the decision This is a test-suite policy question, and the honest answer starts by naming what a convention is supposed to buy: fewer decisions per line, uniform review expectations, and code that reads the same everywhere. Those are real benefits — but they are bought with whatever the chosen uniform style gives up. Here the currency is compile-time type checking, and that makes the trade lopsided. ## The case for standardizing on the do-family Take it seriously before rejecting it: - **One form to learn.** Newcomers hit the spy trap (`when(spy.call())` executing the real method) exactly once and then never again if the codebase never uses `when(...)`. - **No form-switching mid-test.** Mixed syntax within a class can look arbitrary to a reader who does not know why. - **The do-family is universal.** It works for mocks, spies, void methods and generics; `when(...).thenReturn(...)` does not. A single universal form means no line ever has to be rewritten when a mock becomes a spy. - **Lint-ability.** A blanket "never use `Mockito.when`" rule is trivially enforceable; "use `when` unless the target is a spy" is not, because static analysis rarely knows which is which. That last point is the strongest one, and it is why some teams do adopt the blanket rule. It should be acknowledged rather than dismissed. ## Why the default should still be when(...) **Type safety is the decisive asymmetry.** `when(T methodCall)` infers the stubbed method's return type, so `thenReturn` is checked at compile time. Change a method from `long` to `UUID` and every such stub fails the build, immediately, with the method name in the error. `doReturn(Object)` has nothing to check against; the same refactor leaves stubs that compile, and Mockito only reports `WrongTypeOfReturnValue` if and when that test line executes. In a large suite, "fails at build" versus "fails when that test happens to run" is a meaningful difference in feedback latency, and some mismatches (widening, supertypes, erased generics) never fail at all. **Reading order.** `when(mock.findById(1)).thenReturn(order)` follows the sentence people speak. The inverted form leads with the answer and ends with the question, which is measurably harder to skim in a long arrange block. **Signal value.** If `doReturn` appears only where it is necessary, its presence carries information: this collaborator is a spy, or the method is void, or the generics are awkward. Standardizing destroys that signal and, with it, a cheap review heuristic — "why is there a spy here?" is a question worth prompting. **Ecosystem alignment.** `when(...).thenReturn(...)` is the form in Mockito's own documentation, in most examples and in most engineers' muscle memory. Deviating raises the onboarding cost for every new joiner in exchange for removing one gotcha they would learn anyway. ## The policy I would actually write 1. Default: `when(mock.call()).thenReturn(value)` for plain mocks with value-returning methods. 2. Required: the do-family for spies (always), void methods (no alternative), generics that will not compile otherwise, and `doCallRealMethod`. 3. Never stub a spy with `when(spy.call())`, even when the real method looks harmless — a future change to that method would turn a passing test into a mysterious one. 4. Where `doReturn` is used, bind the value to a typed local first so the intended type is stated on the page. 5. Review rule: two or more `doCallRealMethod`/spy stubbings in one test is a design discussion, not a style nit. Item 3 is the part worth being absolutist about, because it is the rule that prevents real damage; the rest is preference with a justification. ## What the metric should be Rather than policing syntax, watch the *ratio*. A codebase where the do-family appears in a handful of places is healthy. One where it dominates is telling you that spies are everywhere, which in turn says classes under test are mixing behaviour that must run with behaviour that must be suppressed. The response is extraction into collaborators, not a linter rule. Framed that way, the syntax question becomes a diagnostic instrument, which is more valuable than the uniformity it would otherwise buy. ## Where I would concede If a team is genuinely burned repeatedly by the spy trap — say a legacy suite that is spy-heavy and cannot be restructured soon — a temporary blanket do-family rule is a defensible stopgap, paired with a plan to reduce spy usage. Conventions are allowed to be time-boxed responses to a current weakness rather than permanent truths, and saying so is a stronger answer than defending the ideal position unconditionally.

  • What is the single rule you would enforce absolutely, regardless of the wider convention?
    Never stub a spy with the when(spy.call()) form. Argument evaluation runs the real method during setup, which can hit a database, throw from the arrange block, or inflate verification counts. That rule prevents concrete damage rather than expressing a preference, so it belongs in review checklists and, where tooling allows, in a custom lint check.
  • How would you detect that the convention is masking a deeper problem?
    Track how often spies and doCallRealMethod appear relative to plain mocks. A rising share means classes under test increasingly need part of themselves suppressed, which points at mixed responsibilities rather than at test style. The remedy is extracting the suppressed behaviour into an injected collaborator so a plain mock suffices; a lint rule would only hide the trend.

saying these in an interview costs you the question

  • Arguing for uniformity without acknowledging the lost compile-time type checking
  • Claiming the two forms are semantically identical in all cases, including on spies
  • Treating the choice purely as style with no diagnostic value
  • Proposing a lint rule as a substitute for reducing spy usage
  • Refusing to concede any situation where a blanket do-family rule could be pragmatic

context