skip to content

How do you choose across a Python codebase between hand-written fakes, unittest.mock.patch, and the real dependency?

level: principalimportance: should knowfreq 35%

answer

  1. Decide by ownership, then by purpose
  2. Fakes for boundaries we own
  3. Patch only the edges we cannot change
  4. A fake is a second implementation
  5. One shared test body, both implementations

basics

~20 s

Decide by ownership and by what the test is for: own the interface, own a fake shipped beside the implementation; patch only edges you cannot change; use the real dependency where fidelity is the point.

solid answer

~50 s

I would set this as policy rather than per test. For a boundary the team owns, the default is a hand-written in-memory fake living next to the real implementation, both typed by the same narrow `typing.Protocol` so a static checker keeps them in step and consumers do not each invent a divergent double. `unittest.mock.patch` is reserved for edges we do not own or cannot change yet, and stays rare and visible. The real dependency is used in a small, named set of tests where fidelity is exactly what is being checked. The underestimated cost is that a fake is a second implementation: it needs an owner, it must raise the same exception types the real one raises, and it must be exercised by the same test body as the real thing or it drifts into fiction. In a 17-service dependency graph I would fake two or three facade collaborators, not seventeen clients.

code

python · 46 lines
python
import unittest


class MemoryArchive:
    def __init__(self):
        self._saved = {}

    def save(self, invoice_id, pdf):
        self._saved[invoice_id] = pdf

    def load(self, invoice_id):
        return self._saved[invoice_id]


class FileArchive:
    def __init__(self, root):
        self.root = root

    def save(self, invoice_id, pdf):
        (self.root / invoice_id).write_bytes(pdf)

    def load(self, invoice_id):
        return (self.root / invoice_id).read_bytes()


class ArchiveContract:
    def make_archive(self):
        raise NotImplementedError

    def test_round_trips_a_document(self):
        archive = self.make_archive()
        archive.save("INV-17", b"%PDF-1.7")
        self.assertEqual(archive.load("INV-17"), b"%PDF-1.7")

    def test_missing_document_raises_key_error(self):
        with self.assertRaises(KeyError):
            self.make_archive().load("INV-99")


class MemoryArchiveTest(ArchiveContract, unittest.TestCase):
    def make_archive(self):
        return MemoryArchive()


if __name__ == "__main__":
    unittest.main()

go deeper

for a junior

You are not expected to set this policy, but know the three options exist and that a fake is a small class you write yourself rather than something a library builds for you.

for a middle

Be able to justify one choice for one test: why a fake here, why the real collaborator there, and why patching is the last resort. Notice when two test files have built the same double twice.

for a senior

Show that you have paid the maintenance cost: fake drift caught late, an error path never exercised because the fake could not fail, doubles duplicated across teams. Explain the shared-test-body technique that keeps a fake honest.

for a principal

Own the convention and its exceptions: where fakes live, who maintains them, what justifies a residual patch, how the integration tier is sized so people still run it, and how you would measure the trend without turning the metric into a target.

At one test the choice between a fake, a patch and the real thing is a matter of taste. At the scale of a codebase it is a policy question with a maintenance bill, an ownership model and a failure mode of its own. ### The decision axes **Who owns the interface?** If the team owns the collaborator, it can offer a seam, and a hand-written fake is affordable and worth it. If it belongs to a third party or to an old subsystem nobody may touch, there is often no seam, and patching is the honest answer rather than a workaround. **What is the test actually for?** A test of *our* branching logic wants a fast, deterministic stand-in. A test of the wiring - that the query is really formed correctly, that the serialisation really round-trips - proves nothing against a double, and must run against the real thing (in a container, a temporary instance, or a sandbox) or it is theatre. **What does substitution cost?** A collaborator whose real form is slow, non-deterministic, remote, or destructive earns a seam. One that is pure, fast and in-process should just run. **How many implementations are there really?** A seam with one implementation forever is speculative flexibility. Two or more - real, in-memory, recording - justifies the abstraction. ### The policy that usually works 1. **Fakes at boundaries we own, shipped with the implementation.** The fake lives in the same package as the real client, is imported by consumers, and is typed by the same protocol as the real class. That is what stops five teams hand-rolling five subtly different doubles for one service, each one wrong in a different way. 2. **Patch at the edges we do not own, and keep it rare.** A patch is a licence, not a habit. Rare use also keeps the targets few enough that a rename does not break dozens of unrelated tests. 3. **The real dependency in a small named tier.** A short, deliberate list of tests that exercise the genuine collaborator, run on merge rather than on every save, sized so that people still run them. 4. **One seam per boundary, not one per test.** The proliferation to watch for is not fakes, it is *ad hoc* doubles: a slightly different fake per test file, each drifting on its own trajectory. ### The bill nobody budgets for A fake is a second implementation of a contract. It has three characteristic failure modes. **Drift.** The real collaborator changes; the fake does not; the suite stays green and the deployment breaks. The two defences are typing the seam so a static checker flags the mismatch, and running one shared test body against both the fake and the real implementation - a common `unittest.TestCase` base class holding the behavioural assertions, with two subclasses supplying a fake and a real instance. If the fake cannot pass what the real one passes, it is not a fake, it is a fiction. **Optimism.** Hand-written fakes almost never fail. Real collaborators time out, reject, run out of quota and raise. A fake that cannot be told to raise the same exception types means every error path in the system is untested, which is usually where the outages come from. **Logic creep.** A fake that starts validating, computing and enforcing rules becomes a parallel product with its own bugs, and it can agree with your tests while disagreeing with reality. Fakes should record and return, and nothing else. ### Sizing it against the dependency graph The instinct in a service sitting in front of a 17-service dependency graph is to fake all seventeen. That is the wrong unit. Collapse the graph behind two or three collaborators expressed in the language of the consumer - an invoice source, an archive, a notifier - and fake those. The number of fakes should track the number of *boundaries* the code has, not the number of systems behind them. If it cannot be collapsed, the design, not the testing strategy, is the finding. ### Making the policy stick Write it down as a short convention: which boundaries have fakes, where fakes live, who owns them, what a residual patch needs to justify itself. Make the trend measurable - the count of patch targets per package is a cheap proxy for how many seams the design is missing. And be prepared to say when the policy does not apply: a one-off script, a spike, or a legacy module being retired next quarter does not deserve a protocol and a maintained fake. The point is that the choice is deliberate and consistent, not that one technique wins everywhere.

  • How do you stop a shared fake from drifting away from the real implementation?
    Two mechanisms, and both are cheap. Type the seam with a narrow protocol so a static checker fails the moment the real interface moves and the fake has not; and factor the behavioural assertions into one test body that is run twice, once against the fake and once against the real implementation. The second is the only thing that catches semantic drift - the fake that has the right methods and the wrong behaviour.
  • Where does the fake belong in the repository?
    Next to the real implementation, shipped as part of the same package, so every consumer imports the same one. Putting it in a test directory guarantees each consuming team writes its own, and five doubles for one collaborator drift five different ways. Shipping it also puts it under the same ownership and review as the real client, which is what keeps it current.
  • What would make you accept a patch-heavy module and move on?
    When the code is not ours to change, when it is scheduled for deletion, or when the cost of the seam clearly exceeds the remaining life of the module. A spike, a one-off migration script, or a subsystem being retired next quarter does not earn a protocol and a maintained fake. I would record the exception rather than let it become the ambient style.
  • Is there a metric you would report on this?
    Patch targets per package as a trend, plus the count of distinct doubles for the same collaborator. Neither is a target - people optimise targets by hiding patches in helpers - but both point at where the design is missing seams and where duplicated fakes are about to disagree with each other. The useful report is which modules moved, not the absolute number.

saying these in an interview costs you the question

  • Declares one technique correct everywhere with no tradeoffs
  • Treats fakes as free and never budgets their upkeep
  • Lets each test file hand-roll its own double
  • Writes a fake that never raises the real error types
  • Fakes every system behind a large dependency graph individually
  • Allows business rules to accumulate inside a fake

context