skip to content

How does testing a CQRS system differ from testing a service with a single shared model, and what specific strategy would you use to test that a read-model projection ends up correct after a write, given the projection updates asynchronously?

level: seniorimportance: should knowfreq 45%

answer

  1. three test levels: write, projector, sync
  2. poll-with-timeout not sleep
  3. idempotency: replay same event twice
  4. contract tests on event schema
  5. sync projector in tests for speed

basics

~20 s

You test the write side and read side separately, then test the connection between them by writing data, waiting for the read side to catch up (or polling), and checking it matches — instead of one instant check like you would with a single database.

solid answer

~40 s

Testing a CQRS system means testing three things instead of one: the write model's business logic/invariants in isolation (fast, synchronous, standard unit/integration tests), each read projection's transformation logic in isolation (feed it a canned event, assert the resulting read-model row), and the end-to-end sync behavior across the async boundary, which needs a polling-with-timeout assertion ('eventually the read model shows X') rather than an immediate assertion, since a straight synchronous check will flake against real event-processing latency. Contract tests on the event schema between write and read side are also valuable to catch the version-skew failures common in CQRS before they hit production.

go deeper

for a junior

Should recognize that read model checks need some kind of wait/retry rather than an instant assertion, at a conceptual level.

for a middle

Should be able to describe testing write and read sides in isolation plus an end-to-end poll-based check, and know why sleep() is a flaky anti-pattern here.

for a senior

Should design the full layered strategy (write, projector, sync, contract tests) and specifically call out idempotency testing for at-least-once delivery.

for a principal

Should weigh test-suite cost/speed trade-offs across teams (in-process projector doubles for speed vs. real pipeline for fidelity) and set org-wide contract-testing practice so independently-deployed write/read teams don't silently break each other.

## Why it is harder than testing one model Testing a CQRS system is harder than testing a single-model service specifically because correctness now spans an **asynchronous boundary**, and naive test strategies either: - give **false confidence** (testing each side in isolation and assuming the wiring between them works); - or become **flaky** (testing end-to-end with a fixed sleep that sometimes isn't long enough). A sound strategy tests at three separate levels, each catching a different class of bug. ## Level one — the write model in isolation The first level is the write model in isolation, and this is the easy, familiar part: standard unit and integration tests against the command-handling logic — given this command and this current aggregate state, does it produce the right resulting state and the right events, and does it correctly reject invalid commands (business rule violations, invariant breaches)? Because the write side is typically synchronous and transactional, these tests look exactly like tests for a single-model system: no special handling of timing or async behavior is needed. This layer's job is to prove the write model's business logic is correct and that it emits events with the right shape and content — it says nothing yet about whether the read side interprets those events correctly. ## Level two — each projector in isolation The second level is each read-model projector in isolation, and the key testing insight here is that a projector is really just a **pure(ish) transformation function**: given an event (or a sequence of events), it produces or updates a read-model row. You don't need a running event bus or a live write model to test this — feed the projector a canned event fixture directly and assert on the resulting read-model state, including edge cases like: - **out-of-order event delivery**; - **duplicate delivery** (most message systems offer at-least-once, not exactly-once, delivery, so the projector must be idempotent); - **unknown/newer event versions** it should degrade gracefully on. Testing idempotency specifically — replay the same event twice, assert the read model isn't double-counted or duplicated — catches a whole class of bugs that only show up in production under retry conditions, which is exactly when you don't want to discover them. ## Level three — the end-to-end sync path The third level is the end-to-end sync path, and this is where CQRS testing genuinely differs from single-model testing: you issue a real command through the write model, and then assert that the read model eventually reflects it — but 'eventually' is the operative word. A synchronous assertion immediately after the write ('assert read model shows X') will pass in a fast CI environment and then flake in a slower one, or vice versa, because it's racing the actual projection latency. The correct pattern is **poll-with-timeout**: retry the read-model assertion on a short interval until it passes or a generous timeout elapses, then fail clearly if the timeout is hit — this makes the test robust to real (variable) latency while still catching genuine sync failures within a bounded, reasonable wait. Some teams instead make the test harness deterministic by running the projector synchronously/in-process during tests (processing the event inline rather than through a real async queue), trading a bit of production fidelity for a fast, non-flaky test; this is a reasonable trade-off as long as the team also runs a smaller number of true end-to-end async tests against the real pipeline to catch anything the synchronous test double would miss. ## The fourth practice — contract testing the event schema A fourth, often-skipped but high-value practice is contract testing on the event schema itself: a test suite that asserts the event a write model actually emits still matches what each downstream projector expects to consume. This is what catches the schema-version-skew failure mode (a write-side change silently breaking a projector) before it reaches production rather than after, and is especially valuable when the write model and the projectors are owned by different teams or deployed independently, since there's no compiler forcing them to agree. ## The four techniques on one order A concrete example: an order service's write model handles `PlaceOrder` commands and emits an `OrderPlaced` event. - **Tests at level one** confirm `PlaceOrder` correctly rejects an order with no line items and correctly emits `OrderPlaced` with the right total when valid. - **Tests at level two** feed a canned `OrderPlaced` event (and a duplicate copy of it) directly into the 'customer order history' projector and assert exactly one row appears with the right fields — proving idempotency without needing a real queue. - **A test at level three** places a real order through the full stack and polls the order-history read API for up to five seconds asserting the new order appears, catching any break in the actual wiring between write and read sides. - **A contract test** asserts the `OrderPlaced` event's schema, as actually emitted, still satisfies what the projector's parser expects — catching, for instance, a renamed field before it ships. Together these four techniques give confidence at every layer without either false confidence from over-isolated tests or flaky, slow end-to-end-only tests.

  • Why is testing projector idempotency specifically important, beyond just testing that it produces the right output once?
    Most message delivery mechanisms guarantee at-least-once, not exactly-once, delivery, so in production a projector will eventually receive duplicate events due to retries after transient failures. If it isn't idempotent, that duplicate delivery silently corrupts the read model — e.g. double-counting a total — and this bug class specifically only manifests under retry/failure conditions, so it won't show up in a naive happy-path test.
  • What's the risk of running the projector synchronously/in-process for all your tests instead of against the real async pipeline?
    It gives fast, deterministic tests but doesn't exercise the real queue/broker, serialization format, or actual latency characteristics, so it can miss bugs that only occur in the real async transport — like serialization mismatches or ordering issues under real concurrent load. That's why it should complement, not fully replace, a smaller set of true end-to-end async tests.

Like testing a relay race: you time each runner alone (write model, projector), then you also have to watch the actual handoffs (the sync path) because a runner can be individually fast and still drop the baton — and you can't just assume the race finished in a fixed number of seconds, you have to watch until the baton is actually received.

saying these in an interview costs you the question

  • Tests the read model with an immediate assertion right after a write with no wait/poll
  • Never mentions idempotency or duplicate-event handling in projector tests
  • Treats CQRS testing as identical to single-model testing
  • No strategy for catching event-schema mismatches between write and read side before production
  • Uses fixed sleep() calls to 'wait for consistency' instead of polling with a condition

context