skip to content

How do you structure a data-access test so that its assertions observe state only after the unit-of-work boundary has closed?

level: seniorimportance: must knowfreq 57%

answer

  1. three phases, three boundaries
  2. never assert on the returned object
  3. the test must not wrap the exercise
  4. re-read through a second observer
  5. plain SELECT shows the real columns

basics

~20 s

Give each phase its own boundary: seed and close, run the code under test in a context it enters itself, then assert in a fresh context or with a plain SELECT. The assertion must never read through the tracked set that produced the state.

solid answer

~40 s

Three phases, three boundaries. **Arrange** by seeding the rows and closing that unit of work, so nothing from setup stays tracked. **Act** by invoking the code under test the way production invokes it, letting it open and close its own boundary — the test must not wrap it in an outer transaction, or the boundary can never close. **Assert** from a second observer: a fresh unit of work, or a plain `SELECT` that bypasses the mapper entirely and shows what the columns actually hold. Where a full boundary is impractical, clearing the tracked set at least defeats the identity map, though the transaction and its uncommitted view remain. The habit worth keeping is that the assertion is a *re-read*, never the object the exercise handed back.

go deeper

for a junior

Learn the rule of thumb: assert on data you read back, not on the object you just saved. Reading it back after the write is finished is what makes the assertion about the database.

for a middle

Describe the three-phase shape and why each boundary matters, and explain the difference between clearing the tracked set and actually closing the unit of work.

for a senior

Show the design decisions: which paths get a committed, boundary-crossing test, when a plain SELECT beats a mapped re-read, how the test avoids wrapping the exercise, and how you keep the runtime acceptable.

for a principal

Decide the tiering for the codebase and make it a convention rather than a per-test choice, so the expensive shape lands on the paths that earn it and the cheap shape is not quietly asserting on memory everywhere else.

The single structural change that makes data-access tests tell the truth is to stop letting one open unit of work span the whole test. What follows is the shape that works and the tradeoffs it carries. ## The three-boundary shape 1. **Arrange.** Seed the rows the scenario needs, then end that unit of work. Two things happen: the writes are actually emitted and accepted, and none of the seeded objects remain tracked, so the exercise phase must load them for real. 2. **Act.** Call the code under test the way production calls it — through the entry point that opens its own boundary per request, message or job. Critically, the test must not have an open transaction the production code silently joins; if it does, the boundary the code relies on never closes and the exercise is not the production shape. 3. **Assert.** Open a fresh unit of work, or read the columns directly with a plain `SELECT`, and check the state there. This is the only read that reflects what the database holds rather than what the test remembers. ## Why the assertion must be a re-read Asserting on the object the exercise returned is the most common way a test quietly stops proving anything. That object is the tracked instance the code just manipulated: its fields are whatever the code set, regardless of whether a statement was emitted, whether the engine rewrote a value, or whether the transaction committed. A re-read after the boundary crosses the gap between "the code believes it did this" and "the database contains this". ## Choosing the second observer | Observer | Shows | Costs / caveats | |---|---|---| | A fresh unit of work | What the mapper reads back, including its own conversions | Still goes through the mapping, so a symmetric mapping bug can cancel out | | A plain `SELECT` for the columns | Exactly what landed in each column | Verbose; couples the test to the schema | | Clearing the tracked set only | A real query instead of an identity-map hit | Same transaction, uncommitted view; no detachment | A symmetric bug — writing and reading a value through the same wrong conversion — is the case that argues for the plain `SELECT` on at least the fields that matter: identity, the linking column of an association, a status value another system reads, a monetary or temporal value where precision is easy to lose. ## What else becomes assertable once the boundary closes - **Usability of the returned graph.** If the production caller receives an object after the boundary and walks it, let the test do the same walk outside the boundary. A read of something never loaded then fails in the test instead of in production. - **What the operation actually emitted.** With the phases separated, the statements attributable to the exercise are exactly the ones the code under test caused, which makes counting them meaningful rather than polluted by setup. - **Ordering and constraint timing.** Checks that fire at commit rather than at statement time only fire when the boundary actually ends. ## The costs, honestly Committed test data has to be cleaned up, which is a real design decision for the suite and slower than never committing anything. Setup that goes through the boundary twice costs time. And some code cannot be invoked without its surrounding runtime, so "act the way production acts" is sometimes an approximation. The pragmatic answer is a tiered suite: - most tests stay cheap and check mapping and behaviour, with the tracked set cleared before the assertion so at least the identity map cannot answer; - a smaller tier commits and re-reads across true boundaries, covering the paths where the hand-off matters — anything returning a graph to a caller, anything that writes in a second context what it read in the first, anything with engine-supplied values; - a very small tier reads with plain `SELECT` on the fields where a symmetric mapping error would be expensive. ## The check to apply to any existing test Ask: *could this test still pass if no statement had ever been sent?* If yes, it is asserting on memory. Then ask: *does anything in this test observe the state through a different reader than the one that wrote it?* If nothing does, the test has no second opinion, and a second opinion is the entire point.

  • Why must the test avoid opening a transaction that the code under test joins?
    Because then the boundary the production code depends on never closes inside the test. Objects stay tracked, deferred reads keep working, commit-time checks never run, and the assertion reads uncommitted state through the same context. The exercise is no longer the production shape.
  • When is a plain SELECT worth the verbosity over a fresh mapped read?
    When a mapping error would be symmetric — the same conversion applied on write and read cancels out and the mapped re-read agrees with the wrong value. Reading the raw columns is worth it for identity, association linking columns, values other systems read, and anything where precision or truncation matters.
  • What does clearing the tracked set give you, and what does it not?
    It defeats the identity map, so the next read is a real query and stale or wrong mapping surfaces. It does not detach the way a closed boundary does, does not commit, and does not run commit-time checks, so code that fails because its context has ended still cannot fail there.
  • How do you keep this shape from making the suite slow?
    Tier it. Keep the bulk of tests cheap with a cleared tracked set, reserve committed boundary-crossing tests for hand-off paths and engine-supplied values, and seed bulk data with direct statements rather than through the mapper's write path.

saying these in an interview costs you the question

  • Asserts on the object the exercise returned rather than a re-read
  • Wraps the code under test in the test's own transaction
  • Thinks clearing the tracked set is the same as closing the boundary
  • Believes a mapped re-read catches every mapping error
  • Assumes objects seeded in setup are still tracked when the exercise runs