skip to content

questions

4

What are the five kinds of test double in Meszaros's taxonomy, and what separates them?

level: juniorimportance: must knowfreq 74%

answer

  1. Five names, one umbrella term
  2. Two axes: behaviour supplied, verdict owner
  3. One of the five actually works
  4. Only one kind fails the test itself
  5. Dummy fills a slot, never called

basics

~20 s

A dummy is passed but never used. A fake has a working lightweight implementation. A stub returns canned answers. A spy records the calls it received. A mock holds expectations and fails the test itself when they are not met.

solid answer

~50 s

"Test double" is the umbrella term for anything a test puts in place of a real collaborator, and Meszaros splits it five ways. A **dummy** exists only to fill a parameter slot and is never exercised. A **fake** has a real, working implementation that is too simplistic for production, such as an in-memory store. A **stub** answers calls with values the test chose in advance, so the code under test can be driven down a chosen path. A **spy** answers calls too, but also records what it received so the test can inspect it afterwards. A **mock** is configured up front with the calls it must receive and fails the test itself when reality does not match. The separating axes are how much behaviour the double supplies and who owns the verdict — the test's own assertions, or the double.

code

pseudocode · 7 lines
pseudocode
# stub: canned answer, verdict belongs to the test
gateway = stub(TrackingGateway)
gateway.status_of("PX-4417") returns "DELAYED"

result = tracker.summarise("PX-4417", gateway)

assert result.label == "Delayed"

go deeper

for a junior

Be ready to name all five and give a one-line description of each without hesitating. Interviewers use this as a screening question, and confidently saying "a dummy is only passed, never used" already puts you ahead of most candidates.

for a middle

Explain the two axes rather than reciting definitions: how much behaviour each kind supplies, and whether the verdict comes from the test's assertions or from the double itself. Show that the role is a property of how the test uses the object.

for a senior

Show that you pick a kind deliberately per test and can say what the choice costs. An interviewer at this level wants to hear you reject a mock where a stub would do, and explain what an over-specified double does to a suite over time.

for a principal

Own the vocabulary problem. When a codebase calls everything a mock, reviewers stop being able to ask whether a test checks values or calls; be ready to describe how you would restore the distinction without turning it into ceremony.

### Where the vocabulary comes from **Test double** is the umbrella term for any object a test substitutes for a real collaborator. The five-part breakdown — *dummy, fake, stub, spy, mock* — is Gerard Meszaros's, from his xUnit test-patterns work, and it is the vocabulary most interviewers have in mind when they ask "what kind of double is that?". The term itself borrows from the stunt double: something that stands in the real actor's place for one shot, convincing from the camera's angle and not from any other. The reason the question is asked so often is that most candidates call all five a mock. Getting the names right is worth little on its own; what the interviewer is really probing is whether you can say *what each one buys you* and therefore *why a test picked one*. ### The five, with one running collaborator Take a parcel-tracking gateway: the code under test asks it for the current status of a shipment and, when the answer is late, records a warning. - **Dummy** — an object handed over only because a signature demands it, never called. If the constructor of the class under test takes a gateway but the method you are testing never touches it, a dummy gateway is enough. It carries no behaviour and no expectations; passing it says "this collaborator is irrelevant to this test". - **Fake** — a real, working implementation that takes a shortcut no production system would accept. An in-memory tracking gateway backed by a dictionary of 3,180 pre-seeded parcels is a fake: you can store into it, read back what you stored, and it stays self-consistent across an arbitrary number of calls. It is the only kind of double that *works*. - **Stub** — it answers the calls it receives with values fixed in advance by the test, and nothing more. Ask the stub gateway for parcel `PX-4417` and it returns `DELAYED` because the test said so. A stub has no memory the test reads and no opinion about being called. - **Spy** — a stub that also keeps a record: which methods were invoked, how often, with which arguments. The test runs the code, then interrogates the spy afterwards. The verdict still belongs to the test's own assertions. - **Mock** — configured *before* the exercise with the calls it must receive, and it fails the test itself when the run does not match. The expectation is stated up front and checked by the double, not by an assertion the test author writes at the end. ### The two axes that actually separate them | Kind | What the code under test gets back | Who produces the verdict | |---|---|---| | Dummy | nothing — it is never called | nobody | | Fake | genuinely computed answers | the test's assertions | | Stub | canned answers chosen by the test | the test's assertions | | Spy | canned answers, plus a call log | the test's assertions, after the run | | Mock | canned answers, against pre-set expectations | the double itself | The first axis is **how much behaviour is supplied**: none, canned, or genuinely implemented. The second is **where the failure comes from**. Only the mock can fail a test on its own; every other kind is inert scaffolding that lets the test reach the state it wants to assert on. That second axis is the one candidates miss, and it is why "mock" is not a synonym for "double". ### A worked judgement Suppose the requirement is that when the gateway takes longer than the team's 92nd-percentile budget of 340 ms, the caller gives up and returns a cached status. The test wants an intermittent timeout on demand. A stub that raises a timeout on the second call is enough to drive the path, and the assertion sits on the cached status the method returns. Reach for a mock here and you have promoted "the gateway was called exactly twice" into a hard requirement that nobody asked for; reach for a fake and you are building a working gateway to test a code path that never needs one. ### Where the vocabulary is loose Be honest in an interview that usage varies. Many teams, and much tooling, say "mock" for every kind; some writers collapse the five into two groups — those that merely let the test run and the one that can fail it. The taxonomy is a communication tool, not a standard anyone enforces, so the strong move is to name the kind *and* describe its behaviour in one breath: "a stub — it returns a fixed status, and the test asserts on the result". That is unambiguous whatever vocabulary the other person uses.

  • Which of the five is the only one with a working implementation, and what does that buy a test?
    The fake. It computes real answers from real state, so it stays self-consistent across many calls — write a record, read it back, list it later. That makes it the right choice when the code under test drives a long sequence of interdependent calls that a chain of canned answers would model badly. The price is that somebody has to build and maintain it.
  • Can one object play more than one of the five roles in a single test?
    Yes, and that is the usual source of confusion. The kinds describe how a test *uses* a double, not how it was constructed. The same stand-in can supply a canned answer for one method and carry an expectation on another, which makes it a stub and a mock at once. Interviewers like candidates who say the role is a property of the test, not of the object.
  • Why do people insist that "mock" is not a synonym for "test double"?
    Because it erases the distinction that matters most: only a mock can fail a test by itself. Calling every stand-in a mock hides whether a test is checking the value the code produced or the calls the code made — two very different things to assert. Once the words blur, so does the review question "is this test checking the right thing?"

A stunt double stands in for the actor for one shot: convincing from the camera's angle and no other. A test double is the same trade — enough resemblance for this one test, and no more.

saying these in an interview costs you the question

  • Calling all five kinds mocks interchangeably
  • Saying a stub and a mock are the same thing
  • Describing a fake as a stub that returns more values
  • Thinking a dummy returns default values on call
  • Claiming a spy always runs the real implementation
  • Believing the taxonomy is enforced by tooling

context

open as a page

In the test-double taxonomy, how does a stub differ from a mock?

level: middleimportance: must knowfreq 68%

basics

~20 s

A stub supplies canned answers so the code under test can run; it can never fail a test. A mock is configured in advance with the calls it must receive and fails the test itself when the run does not match.

open as a page

How do you pick the cheapest test double that still makes a test fail for the right reason?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Start from the oracle — what decides that the behaviour is wrong. A returned value means a stub suffices; a collaboration that is itself the requirement needs a mock or a spy; many interdependent calls need a fake.

open as a page

Which test-double kinds should a codebase standardise on, and how would you decide as a lead?

level: principalimportance: should knowfreq 42%

basics

~20 s

Set a default rather than a ban: stubs and dummies freely, expectations only when the collaboration is itself the requirement, and hand-written fakes as owned shared code for the few dependencies that earn one. Then measure whether the default is holding.

open as a page