skip to content

Unit Testing

Testing one unit of behavior in isolation: fast, deterministic and fine-grained, with a deliberate choice about which collaborators stay real. Interviewers probe the failure modes — over-mocking, and tests that pin implementation details so refactoring breaks them.

on this pageshow

questions

4

What is a unit test, and what does the term unit actually refer to?

level: juniorimportance: must knowfreq 88%

answer

  1. The smallest check you run constantly
  2. Two properties: fast and deterministic
  3. The boundary is chosen, not physical
  4. Milliseconds; no clock, disk or network
  5. A red test names a small region

basics

~20 s

A unit test exercises one small piece of behaviour without touching slow or shared parts of the system, so it runs in milliseconds and fails precisely. The unit is a boundary the team chooses, not a fixed size.

solid answer

~50 s

A unit test exercises a single unit of behaviour and asserts the outcome that behaviour promises. Two properties define it: it is deterministic — the same verdict on every run, on any machine — and it is fast, typically single-digit milliseconds, because the whole value of the level is that engineers run thousands of these on every edit. The unit itself is a **chosen** boundary, not a physical one: some teams draw it around a single class with every collaborator substituted, others around a cohesive behaviour with small in-process collaborators left real. What both styles rule out is anything that makes the test slow or nondeterministic — a network hop, a real database, the wall clock, the filesystem, a background thread you have to wait on. The payoff is diagnostic precision: a red unit test names the region the defect lives in.

code

pseudocode · 5 lines
pseudocode
test renewal_with_expired_card_is_marked_past_due:
    subscription = Subscription(status = ACTIVE, card = expiredCard)
    result = renewalPolicy.evaluate(subscription, at = INSTANT_2026_04_11T02_14Z)
    assert result.status == PAST_DUE
    assert result.reason == CARD_EXPIRED

go deeper

for a junior

Be ready to define the level in one breath: one behaviour, no network, database, clock or disk, milliseconds, same verdict every run. Have a concrete example of something you tested at this level and something you deliberately did not.

for a middle

An interviewer expects you to explain that the unit is a chosen boundary and to describe both the narrow and the behaviour-shaped style, with the tradeoff each one buys. Be able to say which dependencies force a test out of the level and why.

for a senior

Show the judgement: where you draw the boundary in real code, how you keep the suite deterministic when the behaviour depends on time or files, and how you read awkward setup as a signal about the design rather than a testing chore.

for a principal

Own the economics. Be prepared to argue what the level buys per second of build time, why the suite's wall-clock budget is an engineering constraint you defend, and how you keep a definition of 'unit' consistent enough across teams that the level stays cheap.

## The level, stated plainly A unit test is the narrowest automated check a codebase runs: it puts one piece of behaviour into a known starting state, invokes it, and asserts on the result. It is the only test level that is cheap enough to run continuously — after a keystroke, before a commit, on every push — and that economics, not any formal definition, is what actually shapes the rules around it. Three properties do the defining work. **Deterministic.** The same code plus the same test must produce the same verdict every time, on every machine, in any order, forever. A check that passes nine times in ten teaches the team to re-run rather than to read, and a level that is re-run rather than read has stopped being evidence. **Fast.** The working number is milliseconds per test, so that a suite of a few thousand finishes in seconds. Speed is not a nice-to-have here — it is the constraint that makes the feedback loop tight enough to change how people work. Once a suite crosses the threshold where an engineer alt-tabs away while it runs, it has quietly become a pipeline artefact instead of a design tool. **Localising.** When it goes red it should tell you roughly where to look without a debugger. A test that spans half the system may be perfectly correct and still cost an hour per failure, which is precisely why the narrow level exists alongside the wider ones. ## What is the unit? This is the part interviews actually probe, because the honest answer is *it is a decision, not a discovery*. There is no property of production code that marks where a unit ends. Two long-standing styles draw the line differently: - **Solitary** — the boundary is one class or one function, and every collaborator it reaches for is substituted with something inert. A failure can only come from inside that one type. - **Sociable** — the boundary is one cohesive *behaviour*, and collaborators that are in-process, deterministic and cheap are left real: a value object, a small formatter, an in-memory collection, a pure calculation. Only the collaborators that reach outside the process are substituted. Neither is wrong. They trade the same two goods against each other: how sharply a failure points at a line, versus how much the test survives a refactor that reshuffles the classes without changing the behaviour. A newcomer who insists 'a unit is exactly one class' is repeating a convention as if it were a law; the useful formulation is that a unit is *the smallest boundary whose behaviour a reader of the test can name in one sentence*. ## The lines that are not negotiable Whatever boundary you choose, some dependencies push a test out of this level regardless of how few classes it touches: - the **network** and anything behind it, - a real **database** or any shared store other runs also write to, - the **filesystem**, when the test depends on real paths, real permissions or files another run can see, - the **wall clock**, when behaviour changes with the date, the hour or a time zone, - **concurrency you have to wait for** — a background worker plus a sleep in the test body. Each of these buys either time or nondeterminism, and usually both. The usual remedy is not to add tolerance to the test but to move the dependency to a parameter: pass the instant in, pass the content in, hand the behaviour an abstraction it can be given a fixed answer for. That is why the level is so often described as design feedback — code that is awkward to test at this level is usually code that reaches for ambient state instead of receiving what it needs. ## What a good one looks like ```pseudocode test renewal_with_expired_card_is_marked_past_due: subscription = Subscription(status = ACTIVE, card = expiredCard) result = renewalPolicy.evaluate(subscription, at = INSTANT_2026_04_11T02_14Z) assert result.status == PAST_DUE assert result.reason == CARD_EXPIRED ``` One arranged state, one invocation, assertions on the promise the behaviour makes rather than on the steps it took to get there — and no ambient input at all: even the instant arrives as an argument. ## Where it sits The narrow level answers 'is this rule right?' It does not answer 'is the wiring right?', 'does the query really return that?', or 'does the assembled product work?' Those are genuinely different questions with their own levels and their own costs. What the unit level uniquely offers is volume: enough checks, cheap enough, to cover the combinatorial edges of a rule — the boundary values, the empty case, the negative amount, the duplicate — which no wider level could afford to enumerate.

  • Does a test stop being a unit test the moment it touches a second production class?
    No. The boundary is a choice, and a widely used style deliberately leaves small, in-process, deterministic collaborators real so the test survives internal reshuffling. The operative criteria are speed, determinism and how well a failure localises — not a count of types instantiated. What does push a test out of the level is a dependency on the clock, the disk, a shared store or the network.
  • How fast should one unit test be, and why does the number matter?
    Single-digit milliseconds is the working target, so a few thousand of them finish in seconds. The number matters because it decides where the suite gets run: a suite that returns before attention wanders is run on every edit and shapes the design; one that takes minutes is run only by the pipeline, and by then the change has left the author's head.
  • A new unit test needs thirty lines of setup before it can call anything. What does that tell you?
    Usually that the boundary is wrong or the behaviour has too many dependencies to be one unit. Treat the setup cost as feedback on the design rather than as a chore: split the behaviour, push the ambient inputs to parameters, or move the wide part of the check to a level that is meant to be wide. Adding a setup helper hides the signal without removing the cause.

It is a bench check on one part before assembly: you power the part on its own, on a rig you control, so a fault is the part's and not the whole machine's.

saying these in an interview costs you the question

  • Any test a developer writes counts as a unit test
  • A unit is always exactly one class, by definition
  • It is still a unit test, just a slow one, if it hits a real database
  • Speed does not matter because the pipeline runs the suite
  • A test that usually passes is good enough; re-run it
  • Unit tests prove the assembled system works

context

open as a page

How should you name a unit test so its failure explains itself in a report?

level: middleimportance: should knowfreq 62%

basics

~20 s

Name a unit test for the reader of a red report, who sees only the name: state the unit or behaviour under test, the condition it is in, and the outcome expected. A name needing the word and usually means the test checks two things.

open as a page

A unit test for a subscription renewal job passes by day but fails on the nightly run. How do you diagnose it?

level: seniorimportance: should knowfreq 57%

basics

~20 s

Suspect a hidden ambient input. A test that passes at one time of day and fails at another is reading the wall clock or relying on an unspecified ordering. The fix is to pass the instant in and make the expected order explicit, not to loosen the assertion.

open as a page

How do you decide whether a unit test may exercise real collaborators, and what does that choice cost?

level: principalimportance: should knowfreq 44%

basics

~20 s

Decide by what the collaborator costs: keep it real when it is in-process, deterministic and fast; substitute it when it reaches outside the process. A narrow boundary localises failures sharply but breaks on refactoring; a wider one survives change but points at a region, not a line.

open as a page