skip to content

A test needs a collaborator that fails on the first two attempts and succeeds afterwards. Compare encoding that as a MockK `andThen` chain against re-stubbing the same call between phases of the test — how do you choose, and what breaks later?

level: principalimportance: should knowfreq 25%

answer

  1. chain keyed on call ordinal; re-stub keyed on test phase
  2. loop inside one SUT call ⇒ chain (no other option)
  3. extra collaborator call shifts the whole chain
  4. surplus calls absorbed silently — always verify(exactly = n)
  5. argument-driven ⇒ matchers; state-driven ⇒ fake

basics

~20 s

A chain ties behavior to call position, so it is compact but breaks whenever the number of calls changes. Re-stubbing ties behavior to a point in the test's narrative, so it survives extra calls but only works if the test can pause between phases. Choose by whether the test's story is about counting calls or about a state change.

solid answer

~1 min

Both are legitimate; they encode different things. **`andThen` chain** — behavior is a function of *how many times* the call has happened. It is the right choice when the production loop is the thing under test (retry, poll, paginate) and the call count is part of the contract. Costs: any extra invocation — a new log line, a metric, a defensive re-read — shifts every later answer and fails the test far from the change. And the sequence never complains about surplus calls, since the last answer repeats. **Re-stubbing between phases** — behavior is a function of *where you are in the test*. It fits tests that exercise the system in stages: broker down, assert degradation, re-stub broker up, assert recovery. Costs: it only works if the test can interleave with the code under test, and mid-test mutation of a mock makes a long test hard to read. My rule: chain when the loop is inside one call to the system under test; re-stub when the test itself drives the phases. Either way assert the attempt count explicitly, and if the chain grows past three or four links, that is the signal the collaborator wants a small hand-written fake instead.

code

kotlin · 4 lines
kotlin
every { client.get(any()) } throwsMany listOf(IOException(), IOException()) andThen ok

assertEquals("ok", fetcher.fetch("/x"))
verify(exactly = 3) { client.get(any()) }   // count asserted, not assumed

go deeper

for a junior

Recognise both mechanisms exist and that a chain is what you use when the production code loops on its own.

for a middle

Explain the coupling difference — call ordinal versus test phase — and that a chain breaks when the number of calls changes.

for a senior

Add the operational habits: verify the call count, terminate chains that must not be over-consumed, clear recorded calls between phases, and prefer matcher-scoped stubs for argument-driven answers.

for a principal

Give a decision rule the team can apply without you, including the exit criterion for abandoning stub sequencing in favour of a hand-written double, and justify it by maintenance cost rather than taste.

## Two ways to say "it changes over time" MockK gives you two mechanisms for a collaborator whose answer is not constant, and they differ in what the changing answer is *keyed on*. 1. **Answer sequences** — `returns a andThen b`, `returnsMany`, `throwsMany`, `andThenThrows`. One stub holds an ordered list and a counter; the key is the ordinal of the matching invocation. 2. **Re-stubbing** — call `every` again on the same recorded call partway through the test. The newest matching stub wins from that moment on; the key is the position in the *test's* execution, not the mock's call count. Choosing between them is a design decision about what the test is asserting. ## When the sequence is right Use a chain when the production code performs the repeated calls **inside a single entry point** you invoke once. You cannot re-stub in the middle of `fetcher.fetch()`, so anything driven by a loop or a retry policy inside the system under test must be encoded up front: ```kotlin every { client.get(any()) } throwsMany listOf(IOException(), IOException()) andThen ok assertEquals("ok", fetcher.fetch("/x")) verify(exactly = 3) { client.get(any()) } ``` The chain is compact and reads top-to-bottom like the scenario. It is also the only option for that shape. ## The cost of positional coupling The chain silently assumes the exact call sequence. Three realistic ways that assumption rots: - Someone adds a call to the same collaborator for an unrelated reason — a health check before the request, a header lookup, a second read for logging. Every later answer shifts by one and the test fails with a wrong *value*, not with a message about the extra call. The diagnosis time is disproportionate to the change. - The retry policy gains a jittered extra attempt. The chain does not fail; the surplus attempt just repeats the last answer, so the test passes for the wrong reason. - The chain gets long. Five links encode a whole script, and the reader has to simulate the production loop in their head to know which link corresponds to which attempt. Mitigations, in order of value: (a) always pair the chain with an explicit `verify(exactly = n)` so the call count is asserted rather than assumed; (b) end the chain with `andThenThrows` when a surplus call should be an error; (c) prefer argument matchers over positions when the answer genuinely depends on the argument — `every { repo.find(1) } returns … ; every { repo.find(2) } returns …` is far more robust than a two-link chain that happens to be called in that order. ## When re-stubbing is right Re-stubbing suits tests whose narrative has stages that the *test* controls: ```kotlin every { broker.publish(any()) } throws IOException() service.handle(event1) assertTrue(service.isDegraded()) every { broker.publish(any()) } just Runs // broker comes back service.handle(event2) assertFalse(service.isDegraded()) ``` Here the behavior change is a *fact about the world*, not about the number of calls, and encoding it positionally would be a lie: if `handle` happens to publish twice for one event, the chain version breaks while the re-stub version does not. Remember the two mechanics that matter here: the newest matching stub wins, and re-stubbing does not clear recorded calls. If a phase-two assertion should count only phase-two calls, clear the recorded calls between phases explicitly rather than reasoning about cumulative totals. ## The third option, and when to take it When the collaborator's answer depends on accumulated state rather than either call count or test phase — a store that returns what was written to it, a token cache that expires — both mechanisms fight you. That is the point at which a small purpose-written implementation of the interface is the cheaper artifact: it is readable, it has no positional coupling, and it can be shared across tests. The signal is concrete: chains longer than about four links, or a test whose setup you must trace call-by-call to understand. ## Team-level guidance Write it down, because this is exactly the kind of choice that varies per author and makes a suite inconsistent: - Loop inside one call to the system under test ⇒ chain, plus a call-count verification. - Phases driven by the test ⇒ re-stub, plus a recorded-call clear if counts matter per phase. - Answer depends on the argument ⇒ separate matcher-scoped stubs, not a chain. - Answer depends on accumulated state ⇒ a hand-written double. - Never let a chain silently absorb surplus calls; assert the count or terminate the chain with a throw.

  • What single habit most reduces the fragility of chain-based tests?
    Asserting the call count explicitly alongside the chain. The chain cannot detect surplus calls because its last answer repeats forever, so without a verification the test quietly tolerates a changed call pattern. Adding verify(exactly = n) turns a shifted sequence into a precise, well-located failure.
  • When would you stop using MockK for a particular collaborator entirely?
    When the answers depend on accumulated state rather than call position or test phase — a store whose reads must reflect prior writes, or a cache with expiry. Encoding that in chains or repeated re-stubbing produces setup that is longer and less readable than a twenty-line implementation of the interface, which can also be reused across the suite.
  • If you re-stub between phases and then verify a count, what surprises people?
    Recorded calls are not cleared by re-stubbing, so the verification counts calls from both phases. If a phase-scoped count is wanted, clear the recorded calls explicitly between phases while leaving the answers in place.

saying these in an interview costs you the question

  • Treating a chain as self-checking — assuming an extra call would fail the test.
  • Using a chain for behavior that really depends on the argument rather than call order.
  • Trying to re-stub in the middle of a loop that runs inside a single call to the system under test.
  • Assuming a re-stub also resets recorded calls and verification counts.
  • Growing chains to many links instead of recognising the collaborator needs a purpose-written double.

context