You take over a large legacy Java test suite where Mockito mocks are configured leniently and stubbing hygiene has drifted. How would you decide on a strictness policy and roll it out without stalling delivery?
answer
- destination: strict by default, leniency = tracked exception
- WARN to measure, not to live in
- module by module, active code first
- argument mismatches = possible real bugs
- enforce the harness, not just the annotation
basics
~20 sTarget STRICT_STUBS as the default and treat leniency as a tracked exception. Roll out per module: switch to WARN, fix what the hints reveal, then flip to STRICT_STUBS with narrow lenient() escapes where a shared fixture genuinely justifies it. Enforce the harness in review so new tests start strict.
solid answer
~60 sI would set the destination first — **STRICT_STUBS as the default, leniency as a documented exception** — and then make the path cheap. 1. **Measure.** Run the suite under `Strictness.WARN` to get the inventory of unused stubbings and argument mismatches without a red build. The argument mismatches are the interesting ones: some are real production bugs the lenient suite was hiding. 2. **Land the checks module by module**, newest/most-active first, so the cleanup happens where people are already working. A big-bang flip creates a hundreds-of-file change nobody can review. 3. **Fix, don't silence.** Most unused stubbings should be deleted or moved into the test that needs them; the residue is usually one fat shared fixture, which is the actual defect. 4. **Allow narrow escapes** — `lenient()` on a stubbing or `@Mock(lenient = true)` — with a comment. Ban class-wide LENIENT except as a temporary, ticketed marker. 5. **Stop the bleeding**: new tests use the strict harness by default, and review rejects a class-wide LENIENT added without justification. The honest tradeoff is that this buys long-term diagnosability, not a fixed number of bugs, so I would size it as background work carried by the teams touching those modules.
go deeper
Not expected to own this; be able to say strict is the goal and that unused stubbings should be fixed rather than silenced.
Describe the WARN-then-STRICT_STUBS path and the narrow escape hatches for a single module.
Sequence the rollout against active development, distinguish real bugs found from clutter removed, and name the fixture redesign hiding behind most failures.
Lead with the policy statement and its exception rule, justify the investment honestly (diagnosability, not defect counts), and address enforcement across a heterogeneous test estate including code the policy cannot reach.
## Framing the decision This is not a technical question with one right answer; it is a policy question about where a team wants to pay. Lenient tests are cheap to write and expensive to trust: dead stubbings mislead readers, and silent `null` returns from argument mismatches turn a one-line diagnosis into a debugging session. Strict tests are the reverse — occasional friction while writing, sharp diagnostics forever after. For a suite that a team will keep for years, strict wins; for code being deleted next quarter, the migration is not worth funding. ## Step 1 — establish the destination and the exception rule Write it down in one sentence: *the default is `Strictness.STRICT_STUBS`; relaxations must be as narrow as possible and carry a reason.* Without a written destination, every migration stalls at WARN, which is the worst of both worlds — the checks exist but nobody acts on them. The exception ladder, narrow to broad: - `lenient().when(...)` on one stubbing — the normal escape. - `@Mock(lenient = true)` / `withSettings().lenient()` — a fixture collaborator legitimately used by only some tests. - `@MockitoSettings(strictness = LENIENT)` on the class — allowed only as a temporary marker with a ticket. ## Step 2 — measure before you enforce Running the suite at WARN produces the full inventory without breaking anyone. Two categories come out: - **Unused stubbings** — mostly harmless clutter, but the count tells you the size of the job and points at the fixtures that need redesign. - **Argument mismatches** — these deserve individual attention. A mismatch means production code called a collaborator with arguments the test author did not expect, and the test passed anyway on a `null`. A non-trivial share of these are genuine bugs or genuinely stale tests. That second bucket is also the business case: it turns "we want cleaner tests" into "here are N places where the suite was asserting nothing". ## Step 3 — sequence the rollout Big-bang flips fail on review load and merge conflicts. Better sequencing: - **Per module or per package**, starting with the code under active development, so cleanup rides along with work people are already doing. - **New code first**: make the strict harness the default in whatever templates, base classes or archetypes new tests come from, so the problem stops growing on day one. - **Leave the worst module for last** — usually the one with a giant shared fixture — because fixing it is a fixture-design job, not a strictness job, and conflating the two makes the migration look more expensive than it is. ## Step 4 — fix versus silence The discipline that makes or breaks this: when a strict-stubs failure appears, the first question is *why* the stubbing is unused. Typical answers and their correct responses: - The stubbing is genuinely dead → delete it. - Production code changed and no longer calls that collaborator → the test is stale; update the assertions, do not just delete the line. - A `@BeforeEach` stubs everything for a family of scenarios → move stubbing into the tests that need it, or extract small named helpers (`givenUserExists(...)`) that stub on demand. - One collaborator is genuinely shared and optional → mark that mock lenient, with a comment. If the team's default response is `lenient()`, the migration produces no value and a lot of noise. That is worth stating explicitly when you introduce the policy. ## Step 5 — make the harness itself the enforcement point Strictness is only real where a runner, rule, extension or session brackets the test. A suite that mixes `@ExtendWith(MockitoExtension.class)` with bare `Mockito.mock(...)` calls has unenforced pockets no policy statement covers. Practical enforcement options: a lint/ArchUnit-style rule that mock-using test classes must use the extension, a shared base class or composed annotation for the harness, and code review on any newly added class-wide LENIENT. ## Step 6 — accept the costs honestly - Strict stubbing makes some legitimate patterns noisier, especially parameterized tests where each parameter set uses a different subset of a shared fixture. Per-mock leniency is the intended answer there, not a reason to abandon the policy. - Mocks managed by other frameworks (for example Spring-created test doubles) sit outside the Mockito extension's session, so the policy simply does not reach them; do not promise coverage you cannot enforce. - The payoff is diagnosability and reader trust, which shows up as reduced debugging time rather than as a metric on a dashboard. Say that plainly instead of over-claiming a defect reduction. ## What a strong answer sounds like Destination stated in one sentence, WARN used as measurement rather than as a resting place, module-by-module sequencing tied to active work, an explicit narrow-escape ladder, and honesty about the fixture-design work hiding behind most of the failures.
- How do you handle a parameterized test class where each parameter set uses a different subset of a shared fixture, so strict stubbing fails half the runs?Mark the shared fixture mocks lenient at the mock level, or move the stubbing into a per-parameter helper so each run only stubs what it uses. The first is pragmatic and keeps the rest of the class strict; the second is cleaner but costs more setup code. Relaxing the whole class is the wrong answer because it also switches off argument-mismatch detection for mocks that had nothing to do with the problem.
- A team argues strict stubbing slows them down and asks to keep the suite lenient. What is your response?I would separate the two costs: writing friction is real but small and front-loaded, while lenient failures cost debugging time repeatedly and unpredictably. I would show the argument-mismatch findings from the WARN run as concrete evidence that the lenient suite was passing on nulls, and offer the narrow escape hatches so the friction has a cheap outlet. If a module is genuinely short-lived, I would exempt it rather than argue.
saying these in an interview costs you the question
- Proposing a single big-bang commit that flips the whole suite to STRICT_STUBS.
- Treating WARN as the end state because the build stays green.
- Answering every failure with class-wide LENIENT instead of fixing the fixture.
- Claiming strict stubbing prevents production bugs directly rather than improving diagnosability.
- Ignoring that unenforced pockets (bare Mockito.mock, framework-managed mocks) make the policy partly fictional.