skip to content

A team writes unit tests against an in-memory implementation of a Repository interface, then runs integration tests against the real database-backed implementation. What can go wrong when the two implementations' behavior diverges, and how would you catch it?

level: middleimportance: should knowfreq 55%

answer

  1. contract tests run against both implementations
  2. case sensitivity / ordering / null semantics gaps
  3. unique constraints not enforced by a naive fake
  4. Testcontainers for the real-DB half
  5. false confidence from green unit tests

basics

~20 s

The fake in-memory version might behave slightly differently than the real database, like ignoring case or not enforcing uniqueness, so tests pass on the fake but fail for real. Catch it with one shared test suite run against both.

solid answer

~50 s

In-memory Repository implementations (backed by a HashMap or list) are great for fast, isolated unit tests, but they can silently diverge from the real database's behavior: default sort order, case sensitivity in string matching, whether a unique constraint is enforced, null-handling in comparisons, and transaction/rollback semantics are all places a naive in-memory fake gets 'close enough' rather than exactly right. This creates false confidence — tests pass green using the fake, but the same code breaks against the real database in production or in slower integration tests. The standard mitigation is a contract test: a single suite of behavioral tests written against the Repository interface, run twice — once against the in-memory implementation, once against the real database-backed one (often via Testcontainers) — so any divergence in observable behavior is caught immediately rather than discovered in production.

go deeper

for a junior

Should recognize that a fake and a real database might not behave exactly the same, even with matching method signatures.

for a middle

Should name at least two concrete divergence examples (e.g., ordering, uniqueness, case sensitivity) and know that a shared contract test suite is the standard fix.

for a senior

Should be able to design a contract test suite in practice, know tools like Testcontainers, and judge when a domain is complex enough that an in-memory fake isn't worth the risk at all.

for a principal

Should be able to set team-wide policy on when in-memory fakes are appropriate versus mandating real-database integration tests, balancing CI cost/speed against the historical cost of divergence bugs reaching production.

## The two implementations An **in-memory Repository implementation** typically wraps a `HashMap<Id, Entity>` or similar structure and implements the interface's methods (`save`, `findById`, `findByX`, `delete`) using plain Java/Kotlin collection operations — filtering a list, checking map containment, and so on. It's attractive for unit tests because: - it starts instantly; - it requires no database, no schema migration, no network round-trip; - it can be reset between tests with a single line. A **database-backed implementation**, by contrast, delegates the same interface methods to a real SQL engine (via JDBC, an ORM, or a query builder), which brings real transactional semantics, real indexes, and real constraint enforcement. Both implementations are written to satisfy the identical Repository interface's method signatures, which is precisely what lets a service class be wired to either one interchangeably at test time versus runtime — that interchangeability is the entire point of writing the fake in the first place. ## Where the behavior diverges The two implementations satisfy the same interface, but 'satisfies the interface' only guarantees method signatures match — it says nothing about whether the *behavior* matches, and that's exactly where divergence creeps in. Common gaps: - **(1) Ordering** — a `findAll()` over a HashMap-backed store has no guaranteed order, while a SQL table without an explicit `ORDER BY` also technically has no guaranteed order but in practice often returns rows in insertion or index order, so tests that accidentally depend on ordering can pass against one implementation and fail against the other, or vice versa. - **(2) String matching** — an in-memory `equals()`-based lookup is case-sensitive and exact, while many databases perform case-insensitive comparisons depending on collation settings, so a `findByEmail("[email protected]")` might find a stored `[email protected]` in production but not against a naive fake. - **(3) Constraint enforcement** — a real database rejects a duplicate insert against a unique index with a constraint-violation exception; an in-memory fake that just does `map.put(id, entity)` will happily overwrite or duplicate unless someone explicitly coded that check in. - **(4) Null and empty-collection semantics** — SQL's three-valued logic around `NULL` (where `NULL = NULL` is not true) rarely gets faithfully reproduced in a hand-rolled fake. - **(5) Transaction/rollback behavior** — an in-memory fake typically has no notion of a transaction being rolled back on exception, so a test that verifies 'nothing was saved when an exception was thrown mid-operation' can pass trivially against the fake for the wrong reason. - **(6) Concurrency** — a real database serializes conflicting writes according to its isolation level and may raise a deadlock or optimistic-locking failure under contention, behavior an in-memory `HashMap` guarded by nothing (or by a single coarse lock) will never reproduce, so concurrency bugs the production database would surface stay invisible in the fake-backed test suite entirely. ## What the trade-off buys, and what it costs The value this trade-off buys is real: **fast feedback loops** matter enormously for developer productivity, and a large suite of business-logic unit tests that run in milliseconds against an in-memory fake is far more valuable, in aggregate, than the same suite taking minutes against a real database. The cost is the **false-confidence risk** described above — and the failure mode in production is exactly what you'd expect: a bug that every unit test passed on ships, and it's only caught (if at all) by an integration test, a staging environment, or a customer report, at which point it costs far more to diagnose because the discrepancy is between two abstraction layers that were both individually 'correct' by their own interface contract. ## The mitigation The standard mitigation is a *contract test* (also called a shared behavioral test suite): a single set of tests written once against the Repository *interface*, parameterized or run twice, once wired to the in-memory implementation and once wired to the real database-backed implementation (commonly via Testcontainers spinning up a real, ephemeral Postgres/MySQL instance in CI). If both implementations pass the identical suite, you have real evidence that the two are behaviorally interchangeable for everything the contract test suite exercises, not just for the specific scenarios each individual test author happened to think of when writing unit vs integration tests separately. This doesn't eliminate the gap entirely — it only covers what the contract test suite checks — but it converts an invisible, discovered-late risk into a build-time, explicit one. ## A real-world instance A well-known real-world instance of this exact pattern is Spring Data's repository test infrastructure combined with Testcontainers: teams commonly define an abstract base test class with methods like `shouldEnforceUniqueEmail()` or `shouldReturnEmptyWhenNotFound()`, then subclass it twice — once wiring an in-memory fake, once wiring a `@DataJpaTest` backed by a real containerized Postgres — so CI runs the identical assertions against both and immediately flags any behavioral drift as the schema or fake evolves independently over time.

  • Besides contract tests, what's another way teams reduce this risk?
    Some teams skip the in-memory fake entirely for anything beyond the simplest lookups and instead run all Repository-touching tests against a real, fast, ephemeral database (e.g., an in-memory-mode real database engine, or a lightweight Testcontainers instance reused across a test run) so there's only one implementation to diverge from in the first place, trading some test speed for eliminating the divergence risk altogether.
  • Would you recommend an in-memory Repository fake for every project?
    Not universally — it's most valuable when business logic has many pure unit tests that only need a Repository to stand in for persistence at all, and less valuable when the domain logic depends heavily on database-specific query behavior (complex joins, window functions, database-enforced constraints), where a real-but-fast database in tests gives more trustworthy coverage for a modest speed cost.

It's like rehearsing a play with a stand-in actor who reads lines slightly differently than the real lead — rehearsals go smoothly, but opening night reveals timing and delivery gaps nobody caught because the understudy was never checked against the same script cues as the star.

saying these in an interview costs you the question

  • Assumes an in-memory fake and the real database implementation are automatically behaviorally identical because they share an interface
  • Has no answer for how to verify the two implementations agree
  • Doesn't mention any concrete example of behavioral divergence (ordering, case sensitivity, constraints, nulls)
  • Suggests just deleting the in-memory fake and always using the real database everywhere with no discussion of the speed trade-off

context