skip to content

Integration Testing

Testing that components work together across a real boundary — a database, broker, HTTP service or filesystem — exercising wiring, serialization and queries rather than logic alone. Asked because this is where test containers, test data and shared state stop being theoretical.

on this pageshow

questions

4

What defects does an integration test against a real dependency catch that a fully mocked test cannot?

level: juniorimportance: must knowfreq 74%

answer

  1. The stand-in agrees with you
  2. Defects that live between two parts
  3. Wiring, serialization, mapping, query semantics
  4. Precision, absent values, ordering, constraints

basics

~20 s

Integration tests catch defects that live in the boundary itself: wiring and configuration, serialization, schema and type mapping, and real query behaviour. A programmed double returns what you assumed, so it can never contradict that assumption.

solid answer

~50 s

A test whose collaborators are all programmed doubles proves your logic given the world you imagined; it cannot fail when that imagined world is wrong, because the double was written from the same assumption as the code. Keeping one dependency real exercises what only exists at the boundary: the wiring and configuration that connects to it, the serialization and deserialization of values crossing it, the mapping between in-memory types and the dependency's schema, and the semantics of the queries or commands actually issued — ordering, absent values, precision, constraint violations, transaction visibility. Those defects survive every unit test by construction. The price is speed and diagnosability: the run is slower and a failure implicates a boundary rather than one function, so you keep the real part small and assert on what comes back across it, never on the object you handed in.

code

pseudocode · 11 lines
pseudocode
# mocked: the stand-in returns exactly what the test handed it
store = double()
store.on_load(bid_id) -> Bid(amount = 1847.50)
save_bid(store, Bid(amount = 1847.50))
assert load_bid(store, bid_id).amount == 1847.50   # always true

# integration: the value really crosses the boundary and comes back
store = real_store()
save_bid(store, Bid(amount = 1847.50))
reloaded = load_bid(fresh_reader(store), bid_id)   # new path, no cache
assert reloaded.amount == 1847.50                  # fails: 1.847,50

go deeper

for a junior

Be ready to name the split in one breath: unit tests check logic with stand-ins, integration tests check a real boundary. Have one concrete boundary defect ready — a value that changed shape crossing the boundary is the easiest to tell.

for a middle

An interviewer at this level expects mechanics: which four families of defect only appear at a real boundary, why a programmed double structurally cannot expose them, and why the assertion must be on the value read back rather than the object handed in.

for a senior

Show the judgement of what to keep real. Be able to say what a lightweight work-alike still proves and what it silently stops proving, and how you decide that a failure is a boundary defect rather than an unstable environment.

for a principal

Own the argument about evidence. Explain what each level of test actually buys, why boundary defect classes should be covered once rather than at every level, and how you keep the real-dependency layer from becoming the place every new test lands by default.

## The world you imagined versus the world you have A test that replaces every collaborator with a programmed double proves something narrow: that your logic is correct **given the world you assumed**. The double was written by the same person, in the same sitting, from the same mental model as the code it stands in for. When that model is wrong about the dependency, the code and the double are wrong together and the test stays green. That is the structural reason a separate level exists: an integration test puts one piece of the real world back in, so the test is finally able to disagree with you. ## The four defect families that only live at a real boundary **Wiring and configuration.** Addresses, credentials, pool and timeout settings, a converter registered on one code path but not the other, a component constructed without the dependency it is supposed to hold, a setting whose default differs between how you run the test and how you run the service. None of this executes at all while the dependency is a double, because there is nothing to connect to. **Serialization and deserialization.** A value leaves the process as bytes or text and comes back. Decimal precision quietly rounded, timestamps losing an offset or a fractional second, a null that returns as empty text, an enumerated value the reader does not recognise, a field omitted by the writer because it happened to be empty. A double hands back the very object it was given, so the round trip that would expose all of this never happens. **Schema and type mapping.** The dependency has its own type system, and the mapping layer between the two is code that unit tests cannot convincingly exercise: a column narrower than the value written to it, an optional field mapped to a type that forbids absence, a name mismatch that only surfaces when a real statement is prepared, a relationship loaded lazily whose session is gone by the time something reads it. **Query and command semantics.** Result ordering that is not guaranteed unless you ask for it, comparison rules on padded or case-differing text, aggregation over an empty result, uniqueness and referential constraints firing, and what a concurrent transaction can observe part-way through. These semantics belong to the real engine; a hand-written stand-in has different ones or none. ## A worked failure A marketplace bidding engine records each bid amount and re-reads it when the auction closes. The unit tests are thorough: the double returns the amount 1847.50 that the test handed it, the bid-comparison logic is proved across dozens of cases, every branch is covered. The first run against the real dependency fails on a value that never changed: the amount comes back as 1847.5 on one path and as the text `1.847,50` on another, because the value crosses the boundary through a formatter whose decimal separator and grouping follow the runtime locale setting of the process, and the two processes were configured with different ones. Comparison then reads the grouped form as a much smaller number and the wrong bid wins. No amount of logic testing could have found it: the defect lives entirely in a conversion at the boundary, and the double performed no conversion at all. ## What "real" has to mean here The evidence is only as good as the sameness of the dependency. Substituting a lightweight work-alike that speaks a similar dialect keeps the wiring and mapping parts of the test but discards most of the query and type semantics, which is precisely the part that differs between the work-alike and the thing you run in production. A work-alike is a legitimate speed compromise; it is a weaker claim, and you should be able to say which of the four families it still covers. When the dependency is one you genuinely cannot run, that is a different problem with different techniques and other leaves own it. ## Where the assertion goes An integration test that asserts on the object it just handed to the dependency has learned nothing: that object never crossed the boundary. Write, then read back along a fresh path — a new query, a new instance of the reading component, and no in-process cache in between — and assert on what came back. The round trip is the whole point. The same discipline applies to a message pushed onto a broker: assert that a consumer received and decoded it, not that a send call returned normally. ## Cost, and what it does not replace An integration test is slower by orders of magnitude, needs a real dependency to be available, and when it fails it implicates a boundary rather than a single function, so localisation takes longer. That cost is why the real part is kept small and deliberate rather than sprawling. It does not replace unit tests for branch and edge-case logic, which remain far cheaper per case, and it does not prove that a provider you do not run has kept its side of an agreement — a different level of testing answers that. ## What an interviewer is listening for Two things. First, the self-fulfilling-double argument stated plainly: a programmed stand-in cannot contradict the assumption it was built from. Second, at least two concrete defect families with an example each — mapping and serialization are the ones candidates actually hit in real work — rather than the generic answer that integration tests "check that components work together".

  • Why can a suite that is green against doubles fail the very first time it meets the real dependency?
    Because the doubles encode the same beliefs as the code: the assumed field names, the assumed format, the assumed nullability. Nothing in the suite ever tested those beliefs against an independent party. The first real run is the first time an outside implementation gets a vote, so all the accumulated boundary assumptions come due at once.
  • If a value goes into the dependency and comes back changed, where do you put the assertion?
    On the value read back through a fresh path — a new query and a new reading component, with no in-process cache in between — never on the object handed in, which never crossed the boundary. For a message boundary, assert that a consumer received and decoded the message, not that the send call returned.
  • Does swapping the real store for a lightweight work-alike buy the same evidence?
    Partly. You keep most of the wiring and mapping coverage, but the work-alike has its own query semantics, type coercions and constraint behaviour, so ordering, comparison, precision and constraint defects can still hide. It is a defensible speed compromise as long as you can say which defect families it stopped covering.

A programmed double is a rehearsal with an actor reading the lines you wrote for the other side; the real dependency is the first performance with the other side speaking for itself.

saying these in an interview costs you the question

  • Calls an integration test just a slower unit test
  • Claims programmed doubles catch mapping and serialization defects
  • Asserts on the object handed in, not the value read back
  • Thinks integration coverage removes the need for unit tests
  • Treats every boundary defect as an environment problem
  • Assumes a lightweight work-alike proves the same thing

context

open as a page

How do you decide how much of the system an integration test should include?

level: middleimportance: should knowfreq 61%

basics

~20 s

Keep one real boundary in scope and make everything else deterministic or absent. Each extra live component multiplies the suspects behind a failure, so a wide test reports that something is broken rather than what is broken.

open as a page

Your integration tests pass in isolation but fail when the run order changes. How do you find the coupling?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Replay the failing order from its recorded seed, then bisect: run the failing case last behind ever-smaller subsets until one predecessor reproduces it. That pair names the test leaving residue in the shared dependency and the test depending on it.

open as a page

How do big-bang, top-down, bottom-up and sandwich integration strategies differ in ordering and cost?

level: juniorimportance: nice to knowfreq 24%

basics

~20 s

They differ in the order components are joined. Big-bang assembles everything at once; top-down integrates downward using stubs for lower components; bottom-up integrates upward using drivers above; sandwich works from both ends toward the middle.

open as a page