skip to content

You take over a large legacy Java test suite where Mockito mocks are configured leniently and stubbing hygiene has drifted. How would you decide on a strictness policy and roll it out without stalling delivery?

level: principalimportance: nice to knowfreq 20%

answer

  1. destination: strict by default, leniency = tracked exception
  2. WARN to measure, not to live in
  3. module by module, active code first
  4. argument mismatches = possible real bugs
  5. enforce the harness, not just the annotation

basics

~20 s

Target STRICT_STUBS as the default and treat leniency as a tracked exception. Roll out per module: switch to WARN, fix what the hints reveal, then flip to STRICT_STUBS with narrow lenient() escapes where a shared fixture genuinely justifies it. Enforce the harness in review so new tests start strict.

solid answer

~60 s

I would set the destination first — **STRICT_STUBS as the default, leniency as a documented exception** — and then make the path cheap. 1. **Measure.** Run the suite under `Strictness.WARN` to get the inventory of unused stubbings and argument mismatches without a red build. The argument mismatches are the interesting ones: some are real production bugs the lenient suite was hiding. 2. **Land the checks module by module**, newest/most-active first, so the cleanup happens where people are already working. A big-bang flip creates a hundreds-of-file change nobody can review. 3. **Fix, don't silence.** Most unused stubbings should be deleted or moved into the test that needs them; the residue is usually one fat shared fixture, which is the actual defect. 4. **Allow narrow escapes** — `lenient()` on a stubbing or `@Mock(lenient = true)` — with a comment. Ban class-wide LENIENT except as a temporary, ticketed marker. 5. **Stop the bleeding**: new tests use the strict harness by default, and review rejects a class-wide LENIENT added without justification. The honest tradeoff is that this buys long-term diagnosability, not a fixed number of bugs, so I would size it as background work carried by the teams touching those modules.

go deeper

for a junior

Not expected to own this; be able to say strict is the goal and that unused stubbings should be fixed rather than silenced.

for a middle

Describe the WARN-then-STRICT_STUBS path and the narrow escape hatches for a single module.

for a senior

Sequence the rollout against active development, distinguish real bugs found from clutter removed, and name the fixture redesign hiding behind most failures.

for a principal

Lead with the policy statement and its exception rule, justify the investment honestly (diagnosability, not defect counts), and address enforcement across a heterogeneous test estate including code the policy cannot reach.

## Framing the decision This is not a technical question with one right answer; it is a policy question about where a team wants to pay. Lenient tests are cheap to write and expensive to trust: dead stubbings mislead readers, and silent `null` returns from argument mismatches turn a one-line diagnosis into a debugging session. Strict tests are the reverse — occasional friction while writing, sharp diagnostics forever after. For a suite that a team will keep for years, strict wins; for code being deleted next quarter, the migration is not worth funding. ## Step 1 — establish the destination and the exception rule Write it down in one sentence: *the default is `Strictness.STRICT_STUBS`; relaxations must be as narrow as possible and carry a reason.* Without a written destination, every migration stalls at WARN, which is the worst of both worlds — the checks exist but nobody acts on them. The exception ladder, narrow to broad: - `lenient().when(...)` on one stubbing — the normal escape. - `@Mock(lenient = true)` / `withSettings().lenient()` — a fixture collaborator legitimately used by only some tests. - `@MockitoSettings(strictness = LENIENT)` on the class — allowed only as a temporary marker with a ticket. ## Step 2 — measure before you enforce Running the suite at WARN produces the full inventory without breaking anyone. Two categories come out: - **Unused stubbings** — mostly harmless clutter, but the count tells you the size of the job and points at the fixtures that need redesign. - **Argument mismatches** — these deserve individual attention. A mismatch means production code called a collaborator with arguments the test author did not expect, and the test passed anyway on a `null`. A non-trivial share of these are genuine bugs or genuinely stale tests. That second bucket is also the business case: it turns "we want cleaner tests" into "here are N places where the suite was asserting nothing". ## Step 3 — sequence the rollout Big-bang flips fail on review load and merge conflicts. Better sequencing: - **Per module or per package**, starting with the code under active development, so cleanup rides along with work people are already doing. - **New code first**: make the strict harness the default in whatever templates, base classes or archetypes new tests come from, so the problem stops growing on day one. - **Leave the worst module for last** — usually the one with a giant shared fixture — because fixing it is a fixture-design job, not a strictness job, and conflating the two makes the migration look more expensive than it is. ## Step 4 — fix versus silence The discipline that makes or breaks this: when a strict-stubs failure appears, the first question is *why* the stubbing is unused. Typical answers and their correct responses: - The stubbing is genuinely dead → delete it. - Production code changed and no longer calls that collaborator → the test is stale; update the assertions, do not just delete the line. - A `@BeforeEach` stubs everything for a family of scenarios → move stubbing into the tests that need it, or extract small named helpers (`givenUserExists(...)`) that stub on demand. - One collaborator is genuinely shared and optional → mark that mock lenient, with a comment. If the team's default response is `lenient()`, the migration produces no value and a lot of noise. That is worth stating explicitly when you introduce the policy. ## Step 5 — make the harness itself the enforcement point Strictness is only real where a runner, rule, extension or session brackets the test. A suite that mixes `@ExtendWith(MockitoExtension.class)` with bare `Mockito.mock(...)` calls has unenforced pockets no policy statement covers. Practical enforcement options: a lint/ArchUnit-style rule that mock-using test classes must use the extension, a shared base class or composed annotation for the harness, and code review on any newly added class-wide LENIENT. ## Step 6 — accept the costs honestly - Strict stubbing makes some legitimate patterns noisier, especially parameterized tests where each parameter set uses a different subset of a shared fixture. Per-mock leniency is the intended answer there, not a reason to abandon the policy. - Mocks managed by other frameworks (for example Spring-created test doubles) sit outside the Mockito extension's session, so the policy simply does not reach them; do not promise coverage you cannot enforce. - The payoff is diagnosability and reader trust, which shows up as reduced debugging time rather than as a metric on a dashboard. Say that plainly instead of over-claiming a defect reduction. ## What a strong answer sounds like Destination stated in one sentence, WARN used as measurement rather than as a resting place, module-by-module sequencing tied to active work, an explicit narrow-escape ladder, and honesty about the fixture-design work hiding behind most of the failures.

  • How do you handle a parameterized test class where each parameter set uses a different subset of a shared fixture, so strict stubbing fails half the runs?
    Mark the shared fixture mocks lenient at the mock level, or move the stubbing into a per-parameter helper so each run only stubs what it uses. The first is pragmatic and keeps the rest of the class strict; the second is cleaner but costs more setup code. Relaxing the whole class is the wrong answer because it also switches off argument-mismatch detection for mocks that had nothing to do with the problem.
  • A team argues strict stubbing slows them down and asks to keep the suite lenient. What is your response?
    I would separate the two costs: writing friction is real but small and front-loaded, while lenient failures cost debugging time repeatedly and unpredictably. I would show the argument-mismatch findings from the WARN run as concrete evidence that the lenient suite was passing on nulls, and offer the narrow escape hatches so the friction has a cheap outlet. If a module is genuinely short-lived, I would exempt it rather than argue.

saying these in an interview costs you the question

  • Proposing a single big-bang commit that flips the whole suite to STRICT_STUBS.
  • Treating WARN as the end state because the build stays green.
  • Answering every failure with class-wide LENIENT instead of fixing the fixture.
  • Claiming strict stubbing prevents production bugs directly rather than improving diagnosability.
  • Ignoring that unenforced pockets (bare Mockito.mock, framework-managed mocks) make the policy partly fictional.

context