skip to content

Which test-double kinds should a codebase standardise on, and how would you decide as a lead?

level: principalimportance: should knowfreq 42%

answer

  1. Defaults, not bans
  2. Vocabulary loss costs review capacity
  3. One review prompt, answerable in a line
  4. The expensive kind needs a named owner
  5. Say up front how you will measure it

basics

~20 s

Set a default rather than a ban: stubs and dummies freely, expectations only when the collaboration is itself the requirement, and hand-written fakes as owned shared code for the few dependencies that earn one. Then measure whether the default is holding.

solid answer

~50 s

Treat it as a defaults question, not a permissions question. My default hierarchy is: a **dummy** where the collaborator is untouched, a **stub** where the oracle is a value, a **spy or mock** only where a call is the requirement itself, and a **fake** for the handful of dependencies many tests drive through long, stateful sequences. The expensive part is the fake: it becomes shared production-grade code, so it gets a named owner and a review path or it should not exist. I make the policy operable in review with one question — *name the kind you chose and what it lets the test catch* — because the underlying problem is vocabulary: when everything is called a mock, reviewers can no longer ask whether a test checks values or calls. Then I look at outcomes: how often tests fail on changes that broke nothing, and whether new doubles come with a stated reason.

go deeper

for a junior

You are not setting this policy, but you will be reviewed against it. Be ready to say which kind you chose and what it lets your test catch — that one sentence is what most conventions actually ask for.

for a middle

Notice the codebase's prevailing habit and whether it is deliberate. Being able to point out that a team calls everything a mock, and to explain what that costs a reviewer, is a strong signal at this level.

for a senior

Show that you can argue a default within a team rather than just apply one to your own tests. Interviewers look for a rule you can state in one line and a reason it is a default rather than a ban.

for a principal

Own the economics and the rollout. Be ready to name the defaults, the escape clause, who owns a shared working double, why you would not migrate a whole suite, and the signals that tell you six months later whether any of it stuck.

### Why a lead needs a position at all Individually, every double choice is a small local decision. In aggregate they set the price of every future change to the system, because doubles decide when tests break. A codebase that defaults to expectations everywhere charges an engineer-hour toll on every restructuring; a codebase that defaults to the weakest possible stand-in quietly loses the ability to catch whole classes of regression. Neither shows up on a dashboard, which is why the default is a leadership decision and not a style preference. ### Start with vocabulary, not rules The most common state of a real codebase is that all five kinds are called mocks. That is not pedantry — it is a measurable loss of review capacity. If everything is a mock, a reviewer cannot ask "is this test checking the value the code produced or the calls it made?", because the code does not say. Restoring the distinction is the highest-leverage move available, and it costs nothing but a naming convention and a review habit. The cheapest enforcement is a single review prompt: *name the kind and what it lets this test catch*. It is answerable in one line, it is not automatable, and it drags the actual decision into the open. Anything heavier tends to become ceremony that reviewers stamp. ### The defaults I would set - **Dummy** wherever a collaborator is not exercised. Free, and it documents irrelevance. - **Stub** as the default kind. Most tests have a value-shaped oracle and want nothing more. - **Spy or mock** only where the collaboration is the requirement and no value can reveal its absence. This is the clause with teeth: it converts "I mocked it because that is what we do" into an argument someone can lose. - **Fake** for a short, explicit list of dependencies — typically the ones many tests drive through long stateful sequences. A fake is not a test utility, it is shared code with real behaviour, so it needs an owner, a review path, and the same standard as anything else people depend on. If nobody will own it, it should not be written. Note what is *not* in the policy: a ban. Banning a kind produces workarounds that are worse than the thing banned, and it removes the case where an expectation is genuinely the only available oracle. ### Working it on a concrete surface A parcel-tracking gateway used by 41 test classes is the kind of dependency that earns a fake: many tests seed parcels, drive a reconciliation, and assert on outcomes, and a chain of canned answers for each would encode a script rather than a scenario. Meanwhile the audit sink that records "status reported" is used by two tests and has no value-shaped oracle at all — an expectation on the call is right there and nowhere else. The policy has to produce both answers from one rule, which is why it is expressed as a hierarchy of defaults with a stated escape clause rather than as a list of approved kinds. ### Rolling it out without a rewrite Apply the default to new and touched tests only. A sweeping migration of an existing suite spends real budget on tests nobody is currently paying for, and it burns the credibility you need for the part that matters. The exception is a specific hot spot: if one collaborator accounts for a visible share of the churn, converting that one to a fake is a contained project with a legible payoff. ### Knowing whether it worked State up front how you will tell, or the policy is folklore in six months. The signals I would watch: - **Do new doubles arrive with a stated kind and reason?** This is directly observable in review and is the leading indicator. - **How often do tests fail on changes that altered no behaviour?** Engineers can tell you this from memory, and the number moving is the outcome you were buying. - **Do the shared fakes have a live owner?** An unowned fake is a liability whose failure mode is silent. - **Are escaped defects clustering where the doubles are weakest?** If a class of bug reaches production repeatedly in an area covered only by value assertions, the default was too weak there. ### The honest tension to name There is a real disagreement in the field about how much to substitute at all, and a lead should say so rather than pretend the question is settled. Different teams land differently depending on how fast their real dependencies are and how much of their risk lives in collaboration rather than computation. What is not defensible is having no position, because then the default is set by whichever constructor is easiest to type.

  • Why not simply forbid expectations on calls across the codebase?
    Because a real class of requirement has no other oracle — an effect that leaves no trace in any returned value can only be caught by observing the call. A ban pushes engineers into worse workarounds, such as adding a return value purely so a test can assert on it, which distorts the production interface. A default with a stated escape clause gets the same reduction without the damage.
  • How do you keep a shared hand-written fake from becoming an unowned liability?
    Treat it as production code from day one: a named owner, the same review path, and tests of its own. If no team will take it, that is the signal not to build it and to accept the per-test cost instead. The decision to create one is a small architectural commitment, not a test-utility convenience.
  • What would make you conclude your default was set too weak?
    A pattern of escaped defects in an area whose tests only assert values — particularly regressions where an effect on a collaborator silently stopped happening. One incident is noise; a cluster in the same surface says the oracle available to those tests cannot see the risk that actually exists there, and that surface needs a stronger default.

saying these in an interview costs you the question

  • Banning a whole kind instead of setting a default
  • Mandating a suite-wide rewrite of existing tests
  • Treating a shared fake as an unowned test utility
  • Having no way to tell whether the policy held
  • Assuming a naming convention needs no review habit
  • Claiming one substitution style is settled in the field

context