skip to content

How do you stub a collaborator that must return a different value on each call?

level: middleimportance: nice to knowfreq 27%

answer

  1. One answer is not always enough
  2. Change over time versus change by input
  3. An ordered script is consumed one per call
  4. Decide what happens after the last entry

basics

~20 s

Script the stand-in with a sequence of answers consumed one per call, or compute its answer from the arguments it receives. A scripted sequence is simpler but couples the test to how many times the unit calls.

solid answer

~50 s

Two shapes work. The first is a scripted sequence: the stand-in is given an ordered list of answers and hands back the next one on each call, so a unit that polls a redemption status can see PENDING, PENDING, then SETTLED. The second is a computed answer: the stand-in derives its reply from the argument it was given, so asking for member 8,741 always yields that member's balance regardless of call order. Prefer the computed form when the varying answer really depends on the input, and reserve the sequence for genuine change over time - a status that settles, a retry that succeeds on the second attempt. A sequence bakes the exact call count into the test, so an innocuous refactor that adds a cache or reorders two reads breaks it. Always define what the stand-in does once the script runs out, rather than leaving it undefined.

code

pseudocode · 7 lines
pseudocode
statusFeed = stub()
when statusFeed.redemptionState(ref) then returnInOrder(PENDING, PENDING, SETTLED)
afterScript: failWith("unit polled more times than the test scripted")

outcome = SettlementPoller(statusFeed, maxPolls = 5).awaitSettled(ref)

assert outcome == SETTLED

go deeper

for a junior

Know that a stand-in is not limited to one fixed answer: it can be given an ordered list consumed one entry per call, which is how a polling loop is tested.

for a middle

Explain both mechanics - an ordered script versus an answer computed from the argument - and say why the computed form survives a refactor that changes call order while the script does not.

for a senior

Show that you weigh the coupling. Decide explicitly what the stand-in does past the end of a script, choose boundary values so a loop's off-by-one surfaces, and keep the assertion on the outcome rather than on the script being consumed.

for a principal

Own the pattern across a suite: order-coupled scripts are a slow tax that surfaces as unexplained breakage during refactors, so set the team's default to argument-driven answers and treat sequences as a deliberate exception.

### Why one canned answer is sometimes not enough A fixed canned answer assumes the collaborator's reply does not change while the unit runs. Plenty of real units break that assumption. A poller reads a status until it changes. A retry calls once, gets a failure, and calls again. A reconciliation loop walks a ledger in slices until it gets a short one. To exercise any of those honestly, the stand-in has to answer differently across successive calls. ### Shape one: a scripted sequence The stand-in is given an ordered list of answers and returns the next one on each call. For a loyalty-points redemption that settles asynchronously, the script is PENDING, PENDING, SETTLED, and the unit under test is expected to stop polling on the third read and report success. This is the easiest form to read - the change over time is written literally, in order - and it is the form most people mean when they say a stand-in returns different values on successive calls. It has one significant cost: **the test now encodes how many times the unit calls the collaborator, and in what order**. That is a fact about the implementation, not about the behaviour. If someone later adds a short-lived cache in front of the read, or reorders two independent lookups, the script drifts out of alignment with the calls and the test fails without any behaviour having changed. On a 4-person team that is a bad trade to make casually, because the person who breaks the test is rarely the person who wrote the script. There is a second, sharper failure. If the unit calls one more time than the script has entries - the classic off-by-one, a loop bounded by `<=` where the script was written for `<` - the stand-in has no answer to give. Depending on how it was built it may return an empty value, repeat the last entry, or fail with an error that names the stand-in rather than the bug. All three are confusing, and the third at least fails honestly. **Decide explicitly what happens past the end of the script**: either repeat the final answer, or fail with a message saying the unit called more times than the test expected. ### Shape two: an answer computed from the arguments Often the answer is not really varying over *time* - it varies by *input*. If the unit reads balances for several members in one pass, do not script three answers in the order you assume the loop runs. Give the stand-in a small lookup instead: member 8,741 has 2,499 points, member 8,742 has 2,500, member 8,743 has 2,501. Now the answers are correct whatever the iteration order, and the test survives a refactor that reorders or parallelises the reads. This form is strictly more robust and usually just as short. Reach for the sequence only when what you are modelling is genuine change over time on the same input. ### Choosing values that mean something The same discipline that applies to a single canned answer applies here. Pick figures that sit on the rule's edges - 2,499, 2,500 and 2,501 around a 2,500-point Gold threshold - so that an off-by-one in the tiering comparison shows up as a failure rather than passing quietly in the middle of a range. A sequence of placeholder values proves the loop ran; a sequence of boundary values proves the loop was right. ### What the test still asserts As with any stubbing, the varying answers are input. The assertion belongs on the outcome: the final settled result, the tier that was written, the number of points debited. It is tempting, once a script is in place, to assert against the script itself - that the third read happened, that the poll count matched. That is a different technique with its own name and its own coupling costs, and it is not what makes this test valuable. ### A practical checklist - Does the answer really change over time, or only by input? Compute from the argument when it is the latter. - What does the stand-in do on the call after the last scripted one? Make it explicit. - Are the scripted values on the rule's boundaries, or arbitrary? - If someone adds a cache in front of this read tomorrow, does the test fail for a real reason or only because the script no longer lines up?

  • Why is a lookup keyed by the argument usually safer than an ordered script?
    Because it states the relationship rather than the sequence. The answers stay correct if the unit reorders, parallelises or repeats its reads, so the test only fails when the behaviour is wrong. An ordered script freezes the call count and order into the test, which are implementation facts that change for harmless reasons.
  • What should a scripted stand-in do when the unit calls one more time than the script covers?
    Fail with a message saying the unit called more often than the test scripted. Silently returning an empty value or repeating the last answer turns an off-by-one in the loop into a confusing downstream failure, or hides it completely when the repeated answer happens to satisfy the code.
  • When is a fixed single answer still the right choice for a collaborator called several times?
    Whenever the unit's behaviour does not depend on the answer changing. If a rule reads a threshold three times and must treat it identically each time, one fixed answer expresses that and stays stable under refactoring. Script a sequence only to model real change over time.

saying these in an interview costs you the question

  • Scripts a sequence when the answer really varies by argument
  • Leaves the behaviour after the last scripted answer undefined
  • Assumes the unit's call order will never change
  • Uses placeholder values instead of the rule's boundary figures
  • Turns the test into a check that the script was consumed

context