skip to content

Substitution Boundaries

Deciding which collaborators to replace with a stand-in and which to keep real, and where to cut the seam. Interviewers ask because a suite that doubles everything asserts only its own wiring.

on this pageshow

questions

4

Which collaborators in a unit test earn a stand-in, and which should be left real?

level: juniorimportance: must knowfreq 72%

answer

  1. Default is to keep it real
  2. Three properties earn a substitution
  3. Crosses a process boundary
  4. Slow, or nondeterministic on identical input
  5. Value objects and pure logic stay real

basics

~20 s

Replace collaborators that cross a process boundary, are slow, or are nondeterministic, such as remote services, queues and clocks. Leave value objects, pure calculations and cheap in-process helpers real, because replacing those deletes the behaviour the test exists to check.

solid answer

~50 s

The default is to keep a collaborator real and substitute only on evidence. Three properties earn a stand-in: the call leaves the process, the call is slow enough to price out a fast feedback loop, or the collaborator is nondeterministic on identical input, such as a clock or a random source. Two more earn it on the same logic: failure paths you cannot trigger on demand, and side effects you must not cause, like moving money. Everything else stays real, above all value objects, pure calculations and in-process collaborators the team owns. Replacing those is not isolation but deletion: the expected value and the produced value then both come from the test author, so the test can no longer fail for the reason it was written. Ask per collaborator: does keeping this real make the test slow, unreliable, impossible to arrange, or dangerous?

code

pseudocode · 16 lines
pseudocode
test "late filing adds a surcharge":
    rules      = RuleSet(bracket(0, 12_500, 0.0), bracket(12_501, 50_000, 0.2))
    snapshots  = standIn(RulesSnapshotSource)      # out-of-process
    gateway    = standIn(SubmissionGateway)        # irreversible side effect
    clock      = FixedClock("2026-02-17")          # nondeterministic

    snapshots.forTaxYear(2025) returns rules

    wizard = FilingWizard(snapshots, gateway, clock,
                          calculator = BracketCalculator(),  # real: pure
                          money      = Amount.rounding(HALF_UP))  # real: value object

    summary = wizard.summarise(income = Amount("41_320.75"), deadline = "2026-01-31")

    assert summary.liability == Amount("5_764.15")
    assert summary.surcharge == Amount("172.92")

go deeper

for a junior

Be ready to name the three earners in one breath: crosses a process boundary, slow, nondeterministic. Then name what stays real: value objects and pure calculations. Interviewers mostly want to hear that replacing everything is not the goal.

for a middle

Explain why replacing a pure calculator makes the test worthless, not merely wasteful: the expected value and the produced value both originate with the test author. Add the two secondary earners, unreachable failure paths and side effects you must not cause.

for a senior

Show the judgement on the borderline cases: another team's in-process component, a shared cache, a collaborator that is fast but heavy to arrange. Say what you do instead of substituting, such as resetting state per test or building data with a factory.

for a principal

Own the consequence at suite scale. Every stand-in is a hole in the evidence, so a team's default determines what a green pipeline is worth. Be ready to say how you would keep that default visible in review rather than as folklore.

### The decision, stated plainly Every test that exercises a piece of code has to answer one question about each of that code's collaborators: does this one run for real, or does a stand-in take its place? The **substitution boundary** is the line drawn through the collaboration graph — inside the line everything runs as it will in production, outside it the test supplies a controlled replacement. Where you draw that line decides how much of the system the test actually exercises, and therefore how much evidence a green result is worth. The useful default is inverted from what most suites do in practice: **keep the collaborator real unless something forces you not to.** A real collaborator costs nothing to trust; a replacement is a model of that collaborator written by the test author, and a model can be wrong. ### Three properties that earn a stand-in 1. **It crosses a process boundary.** The call leaves your process — a remote service, a broker, a file system, another team's deployed component. The test then depends on that thing being up, reachable, seeded and not concurrently mutated by someone else's test run. 2. **It is slow.** Even in-process work can be slow enough to price a fast feedback loop out of existence: a heavy warm-up, a large computation, an artificial delay. Speed here is measured against the whole suite, not against one test. 3. **It is nondeterministic.** Given identical inputs it produces different answers: a wall-clock or calendar source, a random generator, an identifier factory, anything reading ambient machine state. Nondeterminism does not merely slow a suite down, it makes results untrustworthy. Two further situations earn a stand-in on the same logic even when none of the three strictly applies. **Failure paths you cannot trigger on demand** — you cannot politely ask a real dependency to time out, return a truncated payload or refuse authorization at the moment your test needs it. And **side effects you must not cause** — money moving, messages sent to real people, a record filed with an external authority. ### What stays real - **Value objects.** An amount, a date range, an identifier, a percentage. They are cheap, deterministic, and their behaviour is often exactly what the test is about. - **Pure calculations.** A rules engine, a formatter, a comparator, anything that is input-to-output with no ambient state. - **In-process collaborators the same team owns and the same change set can fix.** Replacing these is the single largest source of tests that pass while the code is wrong. The reason is mechanical, not stylistic. When you replace a calculator with a stand-in that returns the answer, the test's expected value and the code's produced value both come from the test author's head; the calculation itself is no longer under test, and the test cannot fail for the reason it exists. That is not isolation, it is deletion. ### A worked example Take a **tax-filing wizard** that produces a filing summary. Its collaborators are: an amount value type with rounding rules; a bracket calculator that turns income and a rule set into a liability; a rules-snapshot source that fetches the current year's rules from another team's service; a submission gateway that files the return; and a deadline clock. The boundary falls between the last three and the first two. The rules-snapshot source is out-of-process. The submission gateway both crosses a boundary and has an irreversible side effect. The clock is nondeterministic, and the wizard's late-filing branch is unreachable without control of it. The amount type and the bracket calculator stay real — they are the wizard's subject matter. The numbers back the split. That wizard's 214 unit tests finish in about 1.9 seconds with those three replaced. An early version that reached the real rules service in one test took 4.3 seconds for that test alone and failed roughly once a fortnight when the other team redeployed. Meanwhile substituting the bracket calculator in six of those tests hid a rounding defect that nothing in the suite could have caught, because in those six tests the liability figure was whatever the stand-in had been told to return. ### The judgement calls Some collaborators sit on the line. An in-process component owned by *another* team has semantics you do not control, which argues for substitution, but is cheap and deterministic, which argues against. A collaborator that is deterministic but needs a large arrangement is a data-construction problem, not a substitution problem — reach for a builder before reaching for a stand-in. A shared cache is deterministic in isolation and quietly stateful across tests, which usually means keeping it real but resetting it per test rather than replacing it. The interview-grade formulation is a question you ask once per collaborator: *does keeping this real make the test slow, unreliable, impossible to arrange, or dangerous?* If the answer is no on all four counts, it stays real. If you cannot name which of the four applies, the stand-in is habit rather than reasoning, and each such habit shrinks what a green suite proves.

  • A collaborator is deterministic and in-process, but arranging it takes twenty lines of setup. Does that earn a stand-in?
    No, that is a data-construction problem wearing a substitution costume. Reach for a builder or a factory that produces a valid object with sensible defaults and lets the test override the one field it cares about. Replacing the collaborator to avoid its setup throws away the behaviour under test to save typing, and the setup cost usually reappears as configuration of the stand-in anyway.
  • How do you test a branch that only runs when a dependency times out, if the real dependency never times out on demand?
    That is a legitimate earner for a stand-in: substitute the dependency so it raises the timeout on cue, and assert the branch. Keep the substitution narrow, at the seam you own, and make sure some other test still exercises the same dependency on its normal path so the two together cover both. Injecting a fault is one of the few things a real collaborator genuinely cannot do for you.
  • The suite reads the system clock and a handful of tests fail near midnight. Is that a substitution decision?
    Yes. A clock is nondeterministic on identical input, so it is exactly the case a stand-in is for: inject a fixed or advanceable time source and pass it in rather than reading ambient time inside the code. Treating those midnight failures as random noise and rerunning the suite trains the team to ignore red, which costs far more than the injection point.

A flight simulator replaces the weather and the airport because you cannot summon a storm on demand, but it keeps the real flight-control laws, because those are the thing being trained.

saying these in an interview costs you the question

  • Substitutes every constructor argument by default, without a reason
  • Replaces a value object or a pure calculator with a stand-in
  • Cannot name a criterion beyond 'it is a dependency'
  • Thinks speed is the only reason to substitute anything
  • Leaves a clock or random source real, then calls the failures noise

context

open as a page

A unit test replaces every collaborator with a stand-in and still passes while the feature is broken. What went wrong?

level: seniorimportance: must knowfreq 60%

basics

~20 s

The test is tautological: every value the code received was authored by the test, so it confirms only that the code calls the stand-ins it configured. How the real components combine was never exercised, and the defect lives exactly there.

open as a page

Why is substituting at an interface your team owns stronger evidence than substituting at a vendor's own type?

level: middleimportance: should knowfreq 54%

basics

~20 s

A stand-in for a vendor type encodes your beliefs about that vendor, so a pass confirms those beliefs, not real behaviour. An interface you define has a contract you control and collapses vendor assumptions into one adapter you test for real.

open as a page

A suite that doubles at vendor seams everywhere has thin integration coverage. How do you shift the substitution boundary without pausing delivery?

level: principalimportance: should knowfreq 44%

basics

~20 s

Inventory the seams, rank them by where defects actually escaped rather than by count, write a narrow enforceable policy, add adapter tests against the real dependencies before converting anything, then migrate one dependency per release cycle, new code complying immediately.

open as a page