skip to content

In a data-access test, why can reading a just-saved object back inside the same unit of work pass without the database storing it?

level: juniorimportance: must knowfreq 64%

answer

  1. the object never left memory
  2. one id, one instance, in memory
  3. identity map answers the read
  4. no statement emitted before a flush
  5. assert after the write is forced out

basics

~20 s

A mapper holds saved objects in an in-memory identity map and defers the write until a flush. A read by the same id hands back that instance, so the assertion passes whether or not a statement ever reached the database.

solid answer

~40 s

Most mappers work through a **tracked set**: a save call registers the object and its intended change, and the statements are emitted later, at a flush. The same layer keeps an **identity map** so that one id maps to one instance inside the unit of work. When the test reads that id back, the mapper answers from the identity map rather than issuing a `SELECT`, so the assertion compares the fixture object with itself. Nothing about the mapping, the column widths, the not-null constraints, the engine-supplied defaults or durability was exercised. To make the test mean something, force the accumulated changes out and then observe the state from outside the unit of work that produced it.

go deeper

for a junior

Remember two things: a save call registers a change, and the statement goes out at a flush. Until then the object lives in memory, and reading it back by id can be answered from memory.

for a middle

Be able to explain the identity map and why it makes an in-unit read-back tautological, and name what only a flush surfaces (constraints, widths, types) versus what only a fresh read surfaces (defaults, coercion, generated values).

for a senior

Show how you restructure the test: write, close the boundary, re-read in a new one or with a plain SELECT, and assert on the columns the database owns. Say what the old shape let through in a system you worked on.

for a principal

Frame the tradeoff: boundary-crossing tests cost setup and runtime, so decide which paths earn one and what you accept from cheaper in-memory checks. A suite that never crosses a boundary gives false confidence to everyone reading its green result.

A data-access test that saves an object, reads it back by id and asserts on the result feels like it verifies persistence. In a layer with a tracked set it usually verifies almost nothing, and the reason is worth understanding precisely, because the same mechanism hides several other classes of defect. ## What a save call actually does In a mapper built around a **unit of work**, a save call is a registration, not a write: - the object is added to the **tracked set** — the collection of objects the layer is responsible for at this moment; - an intended change (insert, or an update discovered later by dirty checking) is recorded against it; - the layer also puts the object in its **identity map**, a per-unit-of-work index from identity (usually the primary key) to the single instance representing that row; - no statement is necessarily emitted. Statements go out at a **flush** — triggered explicitly, before a query whose results the pending changes could affect, or at commit. Layers differ here: some write eagerly on the call and keep no tracked set, some buffer everything until commit, and some flush automatically before queries. A test written against a buffering layer therefore has a window in which the "saved" object exists only in memory. ## Why the read-back returns without a query The identity map exists so that two loads of the same row inside one unit of work yield the same object — that is what makes dirty checking and reference equality coherent. Its side effect in a test is that a read by id is answered from memory when the instance is already held. The assertion then compares the fixture to itself. ## What each assertion point actually proves | Where the assertion is taken | What it proves | What it still cannot catch | |---|---|---| | After the save call, same unit of work | The object is in the tracked set | Anything the database would say: constraints, widths, types, defaults | | After a flush, same unit of work | Statements were accepted by the engine | Engine-supplied values, since the read is answered from memory | | After the boundary closes, in a fresh unit of work | The rows exist and read back with the stored values | Behaviour at realistic row volume; concurrent access | | By a plain `SELECT` outside the mapper | Exactly what landed in the columns | Whether the mapper can read its own output back | ## The failures this blind spot conceals 1. **Mapping violations.** A value too long for its column, a null in a not-null column, or a wrong type is rejected when the statement is emitted. If no statement is emitted, the test is green. 2. **Engine-supplied values.** Column defaults, generated keys, truncation or coercion, and values written by triggers are only visible once the row is read back from the database. An identity-map read shows the in-memory value the test itself set. 3. **Ordering and cascade effects.** The order in which the layer emits inserts and updates, and what it does with children of a modified collection, is decided at flush time. 4. **Durability.** A statement that has been emitted is not yet a durable row; the transaction still has to commit. ## Layers without a tracked set A query builder or an explicit-write layer emits the statement on the call, so a mapping violation does fail immediately. The blind spot narrows but does not vanish: an assertion taken inside the same open transaction still sees uncommitted state, deferred constraint checks have not run, and a read-back may be served from the transaction's own view rather than proving anything about what other readers will see. ## What to assert instead - **Force the write out before asserting**, so the engine gets a chance to reject it. A test that never triggers a flush cannot fail for a mapping reason. - **Observe state after the boundary closes** — re-read in a second unit of work, or read the columns with a plain `SELECT`. That is the only read that reflects what the database actually holds. - **Assert on the stored values, not on the object you passed in.** Check the fields the database is responsible for: generated identity, defaults, rounded or truncated values, the linking column on the other side of an association. - **Watch the statements the operation emits** when the point of the test is how the data is fetched rather than what it contains. The rule of thumb: a persistence test that never crosses a unit-of-work boundary is a test of your own in-memory bookkeeping. Make the data leave the process and come back before you believe it.

  • The test forces a flush and then re-reads by id in the same unit of work. What is still unverified?
    The statements have been accepted, so mapping violations would now fail. But the read is still answered from the identity map, so anything the engine supplies — a default, a generated key filled in only on re-read, a coerced or truncated value — is not observed. Durability is also unverified until the transaction commits.
  • How does this differ in a layer with no tracked set?
    The write goes out on the call, so a constraint or type violation fails at once and the read-back really is a query. What remains is that the assertion sits inside the same open transaction: uncommitted state is visible to it and to nothing else, and checks deferred to commit have not run.
  • Why is asserting that the object now has a non-null identifier a weak persistence check?
    It shows the layer obtained a key, which some layers do without writing the row at all — from a pre-allocated block of keys, or from a value the application assigned. A key in memory is not evidence that a row exists.

Filing a form in your out-tray and then checking the out-tray to confirm it was received.

saying these in an interview costs you the question

  • Assumes a save call in a buffering layer has already sent a statement
  • Believes a read by id always issues a query
  • Treats a green read-back as proof the column mapping is right
  • Uses flush and commit as if they were the same event
  • Asserts on the object handed to save rather than on stored values