skip to content

When is stubbing a collaborator to return a different value on each successive call the wrong tool for a test, and what would you use instead?

level: principalimportance: nice to knowfreq 26%

answer

  1. sequence as contract vs sequence as workaround
  2. queue = hidden assertion on call count
  3. refactor that caches a read breaks the chain
  4. anonymous values → unreadable in two months
  5. 2-3 meaningful entries fine; 5+ build a fake

basics

~20 s

It is wrong when the sequence encodes how many times the code calls the collaborator rather than what it does. That couples the test to the implementation's call pattern. Prefer a small stateful fake, or state-based assertions, when the collaborator has real behaviour over time.

solid answer

~60 s

Consecutive stubbing is right when the *sequence itself* is the specification — retry after transient failure, poll until ready, a cursor that yields elements. In those cases the ordering is the behaviour under test. It is wrong when it is only a workaround for a collaborator that has state. A queue of five answers silently asserts "the code calls this exactly in this order", so an innocuous refactor — caching a lookup, reading a value once instead of twice — breaks a test that has nothing to do with the change. Long chains also read badly: nothing in `thenReturn(a, b, c, d, e)` says what a, b, c mean. The alternative is a small hand-written fake: an in-memory implementation of the collaborator's interface holding real state, so the test arranges data and asserts on outcomes rather than on call ordering. That is more robust and usually shorter once more than two or three tests need it. Rule of thumb: sequences of two or three that model a named scenario are fine; anything longer is a design signal.

go deeper

for a junior

Know that consecutive stubbing exists and that very long chains are hard to read; defer the design judgment.

for a middle

Explain the coupling: the chain encodes call count, so a caching refactor breaks a test that should not care.

for a senior

Contrast sequence-as-contract with sequence-as-workaround, and be able to introduce a hand-written fake with the state the collaborator really has.

for a principal

Frame it as design feedback — chatty stateful collaborator protocols and over-orchestrating units produce long chains; treat their growth in the test suite as a signal to decompose rather than as a mocking problem.

## Two very different reasons to reach for a sequence When you write `thenReturn(a, b, c)`, you may be doing one of two things, and they have opposite consequences. **Specifying a temporal contract.** The collaborator legitimately behaves differently over time and that is precisely what the test exists to pin down: a call fails once and succeeds on retry; a status poll returns `PENDING` then `DONE`; an iterator yields elements then reports exhaustion. Here the sequence *is* the requirement. The test is not coupled to an accident of implementation — it is asserting the contract. **Papering over collaborator state.** The collaborator is really a store, a queue or a session, and rather than model that, the test enumerates the values it happens to hand back in the order the current implementation happens to ask for them. Now the queue is an implicit, unstated assertion about call ordering and call count. ## Why the second case hurts The damage shows up as refactor-hostility. Suppose production code changes from calling `config.get("timeout")` twice to caching it once. Nothing observable changed, but the second answer now surfaces where the third used to, and every downstream assertion drifts. The test fails, the failure message points at an assertion far from the cause, and the person debugging has to reconstruct a call-count model in their head. The reverse happens too: an added call — a new log line that reads the same value, a defensive re-read — consumes an extra answer and shifts everything. And because the last answer repeats forever, some of these mistakes do *not* fail; they silently feed the same value everywhere and the test goes green while proving less than you think. Legibility suffers as well. `thenReturn(3, 3, 7, 0, 0)` carries no meaning. Two months later nobody can say which call each number was for, so nobody dares change it, so the test ossifies. ## What to use instead **A hand-written fake.** Implement the collaborator's interface with an in-memory version holding real state — a `Map`-backed repository, a list-backed queue, an `InMemoryClock` you advance explicitly. Arrange by putting data in; assert by looking at outcomes or the fake's final state. Order and count stop mattering, so refactors that only change *how often* the code reads a value no longer break the test. The cost is a class to maintain; the payoff is real once three or four tests need the same collaborator, and it forces the interface to be small enough to fake — itself useful design feedback. **A stateful answer.** Where a fake is overkill, one `Answer` that computes from real state (a counter, a queue you drain, the invocation's own arguments) is more honest than a hardcoded list, because it expresses the *rule* rather than a transcript. **Redesign the collaborator.** Frequently the sequence is a symptom: the unit under test is orchestrating too much, or the collaborator interface leaks a stateful protocol (`open`, `next`, `hasNext`, `close`) where a single call returning a collection or a stream would do. Replacing the chatty protocol with one call removes the sequencing problem instead of testing around it. **State-based assertions.** If the sequence exists only so the test can verify "it kept going until done", assert the end state — records written, message published, status final — rather than the intermediate reads. ## When sequences remain the best tool Do not over-correct. Consecutive stubbing stays right for: - **Transient-failure and retry tests**, where the whole point is that attempt 1 fails and attempt 2 succeeds. - **Poll-until-condition loops**, where two entries (`PENDING`, `DONE`) express the contract exactly. - **Exhaustion**, where a single trailing entry that repeats forever models "stays broken" with one line. - **Boundary probes**, such as an ID generator returning two different IDs to prove they are not reused. In each case the chain is short, each entry is meaningful, and the test name states the scenario the sequence encodes. ## A usable heuristic Ask: *if a reviewer deleted one entry from this chain, would the test's stated behaviour change in a way the test name mentions?* If yes, the sequence is specification. If the entry exists only to satisfy an extra call the implementation happens to make, it is coupling. As a practical threshold, two or three meaningful entries in a scenario-named test are healthy; five or more anonymous values are a signal to build a fake or simplify the collaborator. ## Team-level framing Treat long stub sequences the way you treat long mock setups generally — as a measure of how much of the design leaked into the tests. Their growth over time is a useful metric: when a test file's arrange blocks are dominated by ordering, the module underneath usually wants decomposition, not a better mocking trick.

  • What does a hand-written fake give you that a long stub chain does not?
    It holds real state, so the test arranges data and asserts on outcomes instead of on the order and count of calls. Refactors that change how often the code reads something no longer break it, the setup reads as domain data rather than an anonymous value list, and the effort to write the fake pressures the interface toward a small, coherent shape. The cost is one class to maintain, which pays off once several tests share the collaborator.
  • Give a case where a sequence is clearly the right choice.
    Retry-on-transient-failure. The requirement literally is 'the first attempt fails and the second succeeds', so thenThrow(...).thenReturn(...) is the most direct statement of the contract, and paired with verify(times(2)) it pins the retry budget too. Polling loops that return PENDING then DONE are the same shape: two meaningful entries, each named by the test.

A short sequence is a stage cue — 'the phone rings, then it doesn't'. A long one is a transcript of every line the other actor said, which means the scene breaks the moment anyone rephrases a sentence.

saying these in an interview costs you the question

  • Treating any consecutive stubbing as a smell, including retry and polling tests where it is the clearest expression
  • Building long anonymous chains and calling the test thorough
  • Believing a fake is always heavier than stubbing, regardless of how many tests share the collaborator
  • Not recognising that a queue silently asserts call count and ordering
  • Fixing a broken sequence by adding another entry rather than asking why the call pattern changed

context