Your repository tests run on H2 — what bugs slip through that a PostgreSQLContainer would catch?
answer
- Compatibility mode copies syntax, not the engine
- The riskiest SQL you own runs once
- Casing and quoting rules diverge
- Error codes drive your domain exceptions
- Locking is where real incidents come from
basics
~20 sEverything that depends on the real engine: PostgreSQL-only SQL and types, migration scripts that only ever ran against H2, identifier casing and quoting rules, constraint and error semantics, and concurrency behaviour such as locking and isolation. H2 tests prove your Java, not your database.
solid answer
~40 sH2's PostgreSQL compatibility mode emulates a *subset of syntax*; it does not share PostgreSQL's engine. So four classes of defect pass H2 and fail production. **SQL and types**: `jsonb` operators, arrays, `ON CONFLICT` upserts, `LATERAL`, full-text search and extension-provided functions. **Migrations**: your Flyway or Liquibase scripts are the most production-critical SQL you own, and running them against a different engine tests almost nothing. **Semantics**: identifier casing and quoting differ, as do the error codes your code maps to domain exceptions, so unique-violation handling can be untested. **Concurrency**: MVCC behaviour, `SELECT … FOR UPDATE`, lock waits, deadlock detection and isolation levels are engine-specific, and this is where the worst production bugs live. A real container costs startup time; it buys you a test that means something.
code
sql · 9 lines-- routine in PostgreSQL, not portable to an in-memory substitute
INSERT INTO subscriptions (customer_id, plan, metadata)
VALUES (42, 'pro', '{"source":"import"}'::jsonb)
ON CONFLICT (customer_id) DO UPDATE
SET plan = EXCLUDED.plan;
SELECT id FROM subscriptions
WHERE metadata @> '{"source":"import"}'::jsonb
FOR UPDATE SKIP LOCKED;go deeper
Understand that an in-memory database is a different product, so SQL that works in tests can still fail in production, especially anything beyond plain queries.
Enumerate the divergences concretely — dialect features, identifier casing, error codes, migrations — and explain why compatibility mode cannot close a behaviour gap.
Lead with the risk classes and put migrations and concurrency at the top, because those are where undetected defects become incidents rather than annoyances.
Own the trade explicitly: making Docker a hard prerequisite for the build in exchange for meaningful persistence tests, and where the remaining in-memory tests are still legitimate.
## The premise to reject "H2 in PostgreSQL mode is close enough" is the assumption under test here. H2 is an independent database that *accepts* a portion of PostgreSQL's syntax. It does not embed PostgreSQL's parser, planner, type system, MVCC implementation or error codes. Compatibility mode narrows the syntax gap; it cannot narrow the behaviour gap, because there is no shared code. Everything below follows from that one fact. ## Class 1: SQL and type coverage The features teams reach for in PostgreSQL are exactly the ones least likely to be emulated: `jsonb` and its containment and path operators, native arrays, `ON CONFLICT … DO UPDATE` upserts, `LATERAL` joins, range types, full-text search with `tsvector`, and anything provided by an extension. Type semantics differ too — timestamp-with-time-zone handling, numeric precision behaviour, text versus varchar. The practical consequence is a chilling effect: developers avoid the database features that would make the code simpler and faster, because "the tests won't run". Testing against the real engine is as much about what it *permits* you to write as about what it catches. ## Class 2: migrations Migration scripts are among the highest-risk SQL in any codebase — they run once, in production, usually under a deploy window. If your tests apply them to H2, then either the scripts are written to a lowest common denominator (and are not the scripts production runs), or they are PostgreSQL scripts that H2 partially accepts (and the test proves nothing about them). Running migrations against a real container at the start of the suite converts the riskiest artefact you own into the most-exercised one. Many teams find their first container-backed run fails immediately on a migration — that failure is the value being delivered, not a setback. ## Class 3: semantics and error handling Unquoted identifiers are folded differently: PostgreSQL lower-cases, H2 historically upper-cases, and a quoted identifier is case-sensitive in both. Mixed-case names and quoting therefore behave differently between the two. Error codes matter more than people expect. Code that catches a constraint violation and maps it to a domain error — "email already registered" — depends on the SQLState and the constraint name the engine reports. If that mapping is only ever exercised against H2, the branch you most want to trust is the one you never truly ran. Also in this class: default values and generated keys, sequence versus identity behaviour, `NULL` ordering, collation-dependent sorting, and the ORM dialect chosen at runtime — you are testing a different dialect implementation from the one production uses. ## Class 4: concurrency The most expensive production incidents live here. Row-level locking, `SELECT … FOR UPDATE` and its `SKIP LOCKED` variants, deadlock detection and which transaction gets aborted, lock-wait timeouts, and the precise behaviour of isolation levels are all engine-specific. A queue-worker pattern or an inventory decrement that works under one engine's concurrency model can corrupt data under another. No amount of H2 testing tells you anything here. ## What H2 is still good for Speed, and only speed. If a test genuinely exercises Java logic — mapping, validation, a service orchestrating collaborators — and merely needs *a* database to exist, an in-memory one is legitimate. The mistake is letting that convenience decide the fidelity of your *repository* and *migration* tests, which exist precisely to test the database interaction. ## Paying the cost honestly A container adds startup time to the suite. The mitigation is architectural — share one database container across the suite rather than starting one per test class — and it works well enough that container-backed persistence testing is now the default in most JVM shops. Docker availability on developer machines and build agents becomes a hard requirement, which is a real operational commitment worth stating. ## Interview framing Structure the answer by class of defect rather than reciting a feature list, name migrations explicitly (it is the point most candidates miss), and acknowledge the cost with a concrete mitigation. Ending with "and it lets us actually use the database we chose" shows you understand the upside as well as the risk.
- Where is an in-memory database still the right choice?Where the test is about Java, not SQL: service orchestration, mapping, validation, or anything that merely needs a database to exist. Keep repository tests, migration verification and anything touching locking or dialect-specific SQL on the real engine, and be deliberate about which bucket each test belongs to.
- Teams migrating from H2 to a container often see immediate failures. Is that a problem?It is the return on the change. The usual first failures are migration scripts the real engine rejects and exception-mapping code that assumed different error codes — genuine defects that were previously invisible. Budget time for that first pass rather than treating it as the new setup being broken.
- What is the honest cost of container-backed database tests?Startup time and a hard Docker dependency for every developer machine and build agent. The time cost is largely mitigated by sharing one database container across the suite rather than per class; the Docker prerequisite is an operational commitment that has to be made explicitly, not discovered by the first person whose build fails.
H2 in compatibility mode is a flight simulator with the same cockpit layout: excellent for practising your procedures, useless for discovering how this particular airframe behaves in a stall.
saying these in an interview costs you the question
- Says H2 in PostgreSQL mode is close enough
- Never runs real migrations in tests
- Assumes constraint-violation mapping is portable
- Thinks only exotic SQL differs between engines
- Ignores locking and isolation differences entirely