skip to content

How do you decide how much of the system an integration test should include?

level: middleimportance: should knowfreq 61%

answer

  1. What should a red run tell you
  2. Count the suspects behind a failure
  3. One real boundary, everything else deterministic
  4. Widen only for interaction defects
  5. Three dependencies, three narrow tests

basics

~20 s

Keep one real boundary in scope and make everything else deterministic or absent. Each extra live component multiplies the suspects behind a failure, so a wide test reports that something is broken rather than what is broken.

solid answer

~50 s

Decide from what the test must prove and what a red run should tell you. If the claim is that persistence mapping is right, the scope is the component plus the real store — nothing else live. Drive the component directly instead of routing through the layers above it, replace every dependency that is not the boundary under test with something deterministic, and observe the effect at the closest honest point: read the value back through a fresh query rather than through a distant side effect. A component with three real dependencies usually deserves three narrow tests, not one wide one, because each failure then names its own boundary. Widen only for a defect class that genuinely lives between two components — a transaction spanning both, a message a real consumer must be able to act on — and put that reason in the test name so a failure still points somewhere.

go deeper

for a junior

Be ready to say that an integration test should keep one real dependency and stand in for the rest, and to explain in plain terms why a failure is easier to chase when fewer live parts are involved.

for a middle

An interviewer at this level expects the mechanics: drive the component directly rather than through upper layers, replace non-target dependencies deterministically, and observe the effect close to the boundary. Be able to turn one component with three dependencies into three tests.

for a senior

Demonstrate the trade in production terms — localisation cost against realism — and name the specific defect classes that justify widening, such as a transaction spanning two operations. Be ready to diagnose a suite whose failures cost more to triage than to fix.

for a principal

Own the policy: where the boundary between this level and the layer above sits, how many wide tests the organisation can afford to keep diagnosable, and how you stop new tests defaulting into the widest, slowest scope because it is the easiest to write.

## The question a scope decision is really answering Every integration test makes two promises: what it proves when it is green, and what it tells you when it is red. Scope is chosen for the second promise as much as the first. A test that starts six live components and fails tells you the system is broken. A test that starts one component plus the one real dependency it talks to and fails tells you the mapping between them is broken. The information content of the failure is the thing you are buying, and it falls as scope grows. ## One unknown per test The workable rule is: **exactly one thing in the test is real and under suspicion; everything else is deterministic or absent.** If the claim is that persistence mapping is correct, the real part is the component plus the real store. The pricing service it also calls is replaced by something canned. The downstream consumer that reacts to the saved bid is not started at all, because it is not needed to observe the effect being claimed. Each additional live component multiplies suspects, and the multiplication is not linear in the effort it costs you. With two live parts, a failure has three plausible homes: part A, part B, or the interaction. With four, you are reading logs from four processes to decide where to put a breakpoint, and the test's name — which usually still says something narrow — stops predicting where the fault is. ## What gets replaced, and on which side Two different substitutions are in play and they are worth naming separately. **Upstream** of the component under test, you drive it directly rather than through the layers above; there is no reason to route through a user-facing entry point to test a mapping. **Downstream and sideways**, every dependency that is not the boundary under test is replaced with something deterministic, so its latency, availability and data cannot influence the verdict. The corollary is that a component with three real dependencies usually needs three narrow tests, not one wide one. Each keeps a different boundary real, and each failure names its own boundary. ## Observation: the other half of scope Scope is not only what you start; it is where you look. A test can be narrow in what it runs and still be unreadable if the only way it observes an effect is a global side effect several layers away. Prefer the closest honest observation point: read the row back through a fresh query, consume the message from the broker, inspect the response the boundary returned. Reaching further to observe re-widens the scope through the back door. ## A worked scope decision A marketplace bidding engine has a component that accepts a bid, writes it, publishes a `BidAccepted` message and asks a pricing service for the current reserve. Three candidate tests, not one: 1. Component plus real store, pricing canned, no consumer running. Claim: the bid is written and reads back identically, including the amount format. This is the test that would have caught a locale-dependent decimal separator turning 1847.50 into `1.847,50` on the way through. 2. Component plus real broker, store replaced. Claim: an accepted bid produces exactly one message whose payload a real consumer can decode. 3. Component plus the real pricing boundary. Claim: the request is formed and its response parsed as expected. The tempting fourth test starts all of it at once and asserts an end state. It is not wrong to have one such test, but it is a different level with a different budget, and it should not be the test you reach for when you want to know whether the mapping is right. ## When widening is genuinely correct Widen when the defect class you are hunting lives **between** two components rather than inside either: a transaction that must span two operations and roll back as a unit, a message that is only correct if a real consumer can act on it, an ordering guarantee that only exists once both sides run. That is a positive reason, and when you widen for it, say so in the test name — the name should predict the neighbourhood of the fault, so that a red run points somewhere before anyone opens a log. ## Signs your scope is wrong Watch three symptoms. Triage time: if working out *where* a red run comes from routinely takes longer than fixing it, the test is too wide. Irrelevant breakage: if a test named for persistence keeps failing for reasons that have nothing to do with persistence, the extra live parts are noise. And the reverse symptom — a boundary test that never fails for a boundary reason is often not touching the boundary at all, usually because something in the chain quietly substituted a stand-in. ## What an interviewer is listening for The diagnosability argument, stated as a trade rather than a preference: wider scope costs you localisation, and you pay it only for a defect class that genuinely lives in the interaction. A candidate who answers "as much as possible, to be realistic" has not run a suite that failed at 2am; a candidate who answers "always the smallest possible" has not met a transaction that spans two components.

  • Give a case where widening the scope beyond one boundary is the right call.
    When the defect lives in the interaction itself: a unit of work that must commit across two operations and roll back as one, or a published message that is only correct if a real consumer can decode and act on it. In both, no single-boundary test can express the claim. Widen deliberately, and name the test after the interaction so a failure still points at it.
  • How do you tell from the suite that a test's scope is too wide?
    Two symptoms. Triage time: if working out where a red run came from routinely costs more than the fix, there are too many live parts. And irrelevant breakage: a test named for one concern that keeps failing for unrelated reasons is being broken by the extras rather than by its subject.
  • Why does the observation point matter as much as what you start?
    Because reaching far away to observe an effect re-widens the scope through the back door: the assertion now depends on every layer between the boundary and the place you looked. Observe as close to the boundary as is honest — read the value back through a fresh query, consume the message from the broker — while still going across the boundary rather than inspecting in-memory state.

Scope is like the size of the net you cast: a wider net catches more, but every extra thing in it is one more thing to untangle before you know what you actually caught.

saying these in an interview costs you the question

  • Says wider is always more realistic and therefore better
  • Starts every dependency because it is easier to configure
  • Ignores how long a red run takes to localise
  • Names a test narrowly while running half the system
  • Observes the effect far downstream of the boundary
  • Never widens even for a transaction spanning two components

context