skip to content

How do you decide whether a unit test may exercise real collaborators, and what does that choice cost?

level: principalimportance: should knowfreq 44%

answer

  1. It is a policy, not a per-test whim
  2. Substitute at the process edge only
  3. Triage cost versus refactor cost
  4. Wall-clock budget decides the ties
  5. Watch the slow tail, not the average

basics

~20 s

Decide by what the collaborator costs: keep it real when it is in-process, deterministic and fast; substitute it when it reaches outside the process. A narrow boundary localises failures sharply but breaks on refactoring; a wider one survives change but points at a region, not a line.

solid answer

~50 s

This is a boundary policy, not a per-test whim, so I set a default and let teams deviate with a reason. The default I argue for: leave a collaborator real when it is in-process, deterministic and cheap — value objects, calculations, small formatters — and substitute only at the edges of the process, where the clock, disk, network or another thread would cost time or determinism. The tradeoff is explicit. A narrow boundary means a red test names one type, but the tests bind to the structure, so a behaviour-preserving refactor turns a wide band of them red and people learn to distrust the suite. A wider boundary survives restructuring but a failure names a region, so triage costs more. The hard constraint over both is wall-clock time: the suite has a budget, and a boundary choice that spends it is the one I reject.

go deeper

for a junior

Know that there is a choice at all: some collaborators are left real and some are replaced, and the deciding factor is whether the collaborator is fast and deterministic. Do not assume replacing everything is the rule.

for a middle

Be able to describe both styles and name what each buys and costs — sharp localisation versus resistance to refactoring — and to say which dependencies must always be replaced because they cross the process edge.

for a senior

Show that you apply this deliberately in code you own, including how you notice a boundary is wrong: arrangement code outgrowing assertions, refactors forcing wide test edits, or a slow tail of cases quietly crossing the edge.

for a principal

Own the policy and its economics: a written default, a wall-clock budget with a ratchet, migration only where it pays, and an argument grounded in measurements from your own repository rather than in a contested industry multiplier.

## Why this is a leadership decision at all Left to individuals, the boundary drifts. One team substitutes every collaborator by reflex; another instantiates half a module in each test. Both suites still go green, so nothing forces the question — until the bill arrives, either as a refactor that turns two thousand tests red without a single behaviour changing, or as a suite that has quietly grown past the point where anyone runs it before pushing. Both are boundary decisions cashing in late, which is why the choice belongs in a written default rather than in taste. ## The two styles, stated as costs **Narrow.** The boundary is one type; everything it reaches for is substituted. Benefits: a failure implicates exactly one type, and the tests run as fast as anything can. Costs: the test knows the collaboration structure, so any behaviour-preserving change to that structure is a test-editing project. A team in this mode eventually reports that 'refactoring is expensive here', and the suite is why. **Wider.** The boundary is a cohesive behaviour, with in-process, deterministic collaborators left real. Benefits: the tests assert on promises rather than arrangements, so internal reshaping is free, and the tests double as usable documentation of the behaviour. Costs: a failure names a region, so triage takes longer; setup gets larger as the boundary grows; and the temptation to keep widening it — one more real collaborator, and one more — is how a fast suite becomes a slow one without any single decision to blame. The useful framing is that you are choosing between two kinds of maintenance: **triage cost per failure** versus **edit cost per refactor**. Neither is free, and which one dominates depends on how volatile the code's internal structure is. Volatile internals argue for the wider boundary; stable internals with intricate logic argue for the narrower one. ## The default I would write down 1. Substitute at the **process edge** — the clock, the filesystem, the network, another thread, anything shared. This is non-negotiable, because these buy nondeterminism or milliseconds-to-seconds, and the level cannot afford either. 2. Keep real anything **in-process, deterministic and cheap**. A value object, a calculation, a small mapper — replacing these adds test code and removes coverage without buying speed. 3. Where the two rules conflict — an in-process collaborator that is genuinely expensive to build, or one so large that failures stop localising — the decision is documented per module rather than argued per pull request. 4. Any deviation is a comment in the test or a line in the module's notes. The point is not bureaucracy; it is that the next person can see the choice was made. ## Speed as the constraint that decides ties The budget deserves a number, because arguments about boundaries end quickly once someone times the suite. On a renewal codebase I would want a figure of the shape: 14,200 narrow tests finishing in 6 minutes 40 seconds is already past the useful threshold — the suite is no longer in the edit loop, and its feedback arrives after attention has moved. Under a minute keeps it in the loop; a few seconds keeps it in the keystroke loop. So I ratchet: record the wall-clock time, fail the build if it regresses beyond a margin, and treat a boundary choice that spends seconds as a proposal with a cost attached. This also disciplines the wider style, whose failure mode is exactly incremental slowdown that nobody owns. A related discipline: measure the **slowest** tests, not the average. Slow suites are usually a small tail — a handful of cases that quietly cross the process edge — and the tail is where a boundary violation hides. ## What I would not use to justify the policy The familiar claim that a defect costs an order of magnitude more to fix at each later phase is genuinely contested — the underlying studies are old, the conditions are not those of continuous delivery, and the multiplier is cited far more confidently than the evidence supports. I would not build a policy argument on it. The defensible argument is the one you can measure locally: how long triage takes per failure, how often a behaviour-preserving refactor forces test edits, and how long the suite takes. Those are observable in your own repository this quarter. ## How I would roll it out Not as a rewrite. New tests follow the default; existing tests are migrated only when they are already being touched, or when a specific pain — a refactor blocked by a wall of red, a module whose failures never localise — pays for the change. I would ask for two signals in review: does a failure in this test tell the reader where to look, and does this test have any reason to break other than the behaviour changing? A suite that answers yes and no respectively, inside the time budget, has its boundary in the right place, whatever style the team calls it.

  • How do you tell an over-narrow boundary from a well-factored one before the refactor bill arrives?
    Ask what could make each test red. If a behaviour-preserving change to the internal arrangement would do it, the test is bound to structure and the boundary is too tight. In practice I sample: take a planned restructuring, estimate the test edits it forces, and treat a large number as a measured cost rather than an opinion. The other early signal is arrangement code growing faster than assertion code.
  • What number would you actually put on the unit suite's wall-clock budget, and how do you enforce it?
    A target under a minute for the whole suite, and per-test in the low milliseconds, with the caveat that the right number is whatever keeps it in the edit loop for your team. Enforce it by recording the duration on every run, ratcheting the ceiling downward, failing the build on a regression beyond a margin, and reviewing the slowest twenty cases regularly — the tail is where a boundary breach hides.
  • A team argues that a wider boundary makes their tests better documentation. Is that a legitimate factor?
    Yes, within limits. Tests written against promises rather than arrangements read as statements about behaviour, and that value is real. But documentation quality never outranks localisation and speed: a test nobody can triage quickly and a suite nobody runs before pushing document behaviour that is not being defended. Take the readability where the boundary is already justified on the other two grounds.
  • How do you handle a module where the honest boundary is expensive whichever way you cut it?
    Document the decision at the module rather than relitigating it per change, and be willing to move some checks to a level that is designed to be wide instead of stretching this one. If the logic is intricate and the internals stable, take the narrow boundary and accept refactor cost; if the internals churn, take the wider one and invest in failure messages that recover the localisation you gave up.

It is like choosing how much of a machine to put on the test rig: strip it to one part and every fault is unambiguous but the rig must be rebuilt whenever the parts are rearranged; mount a whole subassembly and it survives redesign, but a fault only tells you which subassembly.

saying these in an interview costs you the question

  • Substitute every collaborator, always; that is the definition
  • Keep everything real; it is more realistic that way
  • Suite duration is the build server's problem
  • Defects cost ten times more per phase, so this is settled
  • A refactor breaking many tests proves the tests were working
  • Each developer should pick the boundary case by case

context