skip to content

Read-Only Units & Routing

Marking a unit of work read-only and sending it to a replica or a tenant's datasource, plus the stale reads and mis-pointed pooled connections that follow. A scaling question with a correctness trap.

on this pageshow

questions

6

What does marking a unit of work read-only change about how a data-access layer handles the objects it loads?

level: juniorimportance: must knowfreq 60%

answer

  1. what the unit stops remembering
  2. no snapshot, so no dirty check
  3. flush is skipped at the end
  4. modifications inside are simply dropped

basics

~20 s

A read-only unit stops keeping a value snapshot of each loaded object, so no dirty check and no flush run at the end. Reads get a lighter, cheaper path, and anything modified inside the unit is simply not written.

solid answer

~50 s

A normal unit of work keeps every object it loaded plus a snapshot of the values it read, compares object to snapshot at the end (dirty checking), and flushes the differences as `UPDATE` statements. Marking the unit read-only tells the layer none of that is needed: no snapshot is retained, no comparison pass runs, and there is nothing to flush. The saving is real on big reads - roughly one copy of the data instead of two, and no per-object scan at the end. Many layers also push the read-only attribute down to the transaction or the connection, which lets the engine refuse writes or take a cheaper path, and it is usually the same marking a routing layer reads when deciding to send the unit to a replica. It is still a transaction and it still borrows a connection.

go deeper

for a junior

Remember the two things that stop happening: the layer keeps no before-copy of loaded objects, and no update statements are produced at the end. Say plainly that a change made inside such a unit does not reach the database.

for a middle

Explain the mechanics: snapshot, dirty check, flush, and which of them the marking removes. Note that it is still a transaction on a borrowed connection, and that some layers also push the attribute down so the engine knows the unit will not write.

for a senior

Show where it pays off - large list, export and report paths where per-object bookkeeping doubles the memory of a big result. Mention that the marking is a promise, not an enforced guard, so a write in a supposedly read-only path can disappear silently.

for a principal

Frame it as a default: read-only for query paths, read-write only for command paths. That default buys memory and gives the routing layer a signal to work with, but it makes correctness depend on the marking being honest across every team touching the code.

## What a unit of work normally does for loaded objects A data-access layer that maps rows onto objects almost always keeps the objects it loaded for the lifetime of the **unit of work**: an identity map, so the same row loaded twice yields the same instance, plus - in layers that track changes automatically - a **snapshot** of the column values as they were read. At the end of the unit, or before a query in layers that flush eagerly, the layer walks the tracked set and compares each object with its snapshot. That comparison is **dirty checking**, and the statements it produces are the **flush**. It is what makes `load, set a field, commit` persist without an explicit save call. That convenience is not free: - a second copy of every loaded row's values, held until the unit ends; - a comparison pass over the whole tracked set at every flush; - objects pinned in memory for the unit's whole life, even if the caller only formatted them into a response; - in layers that flush before queries, a large tracked set slows down every later query in the same unit. ## What the read-only marking removes Declaring the unit read-only says up front that nothing in it will be written, so the layer can skip the machinery: 1. **No snapshot** is retained for loaded objects. 2. **No dirty check** runs, so no `UPDATE` is ever generated from a tracked object. 3. **Nothing is flushed** at the end; commit closes a transaction that emitted only reads. 4. In many layers the attribute is **propagated to the transaction or connection**, so the engine itself knows the unit will not write. 5. In a routing layer, the same marking is usually the **signal that the unit may be served by a replica**. | | read-write unit | read-only unit | |---|---|---| | snapshot of loaded values | kept per object | not kept | | dirty check at the end | runs over the tracked set | does not run | | flush | emits pending statements | nothing to emit | | a field you modify | written on commit | discarded | | memory per loaded row | object plus snapshot | object only | | eligible for replica routing | no | yes, where the layer routes | ## What it does not change This is where candidates over-claim. A read-only unit is still: - **a transaction** - it opens and closes one, and it can still be held open too long; - **a connection borrower** - it takes a connection from the pool for its lifetime and returns it at the end; - **subject to the same isolation and visibility rules** - read-only is not a lower isolation level and does not mean dirty reads; - **able to load lazily** - resolving a deferred association is a read, so it generally still works inside the unit. And it does not make the database's own work disappear. The query costs what it costs; what shrinks is the layer's per-object bookkeeping on top of it. On a query returning five rows the difference is noise. On one returning fifty thousand it is often the difference between a comfortable endpoint and a heap problem. ## Two different levels of read-only Data-access layers differ here, and it is worth saying so out loud rather than asserting one behaviour. Some treat read-only purely as a **local optimisation hint**: skip tracking, skip flush, but do nothing to the transaction, so an explicitly issued write statement still reaches the database and commits. Others **declare the transaction itself read-only**, and then a write is rejected by the engine. A candidate who assumes the strict behaviour everywhere will eventually be surprised by a modification that vanished with no error at all. The practical rule that follows: treat the marking as a **promise you make**, not a guard the platform enforces for you. If a code path must write, it belongs in a unit that was never marked read-only in the first place. ## Where to use it Good candidates are the paths that structurally cannot write: list and search endpoints, exports and reports, batch scans that project rows into a transfer model, and anything feeding a rendering layer. The benefit scales with the number of objects loaded, so the endpoints worth marking are exactly the ones that load a lot. Some teams go further and make read-only the default for query paths, leaving read-write for the handful of command paths - which is safe only if the marking is honest, because the write that slips into a read-only path is the failure mode that follows from this whole design.

  • Does a read-only unit of work still need a transaction at all?
    Usually yes. A transaction gives the reads a consistent point of view, so several statements in the same unit agree with each other, and it is the boundary the connection is bound to. Skipping it entirely means each statement runs on its own and a multi-statement report can see the effects of concurrent writes landing between its queries.
  • Where does the memory saving actually come from?
    Mainly from not retaining a snapshot of the values of every loaded object; that is roughly a second copy of the data. Secondarily, layers that skip registration for change purposes can let objects be collected sooner instead of holding them until the unit ends. Both scale with row count, which is why the marking matters on large reads and not on small ones.
  • Is deferred loading still available inside a read-only unit?
    Generally yes - resolving a deferred association issues a read, and reads are exactly what the unit is for. What differs between layers is whether the objects arriving that way are also untracked. It is still worth planning fetches deliberately, because a read-only marking does nothing about a load-per-row pattern.

It is the difference between reading a document and reviewing it with track changes on: the review keeps a before-copy of every line so it can report edits, while plain reading keeps nothing and has nothing to submit at the end.

saying these in an interview costs you the question

  • Thinks read-only means no transaction is opened at all
  • Assumes every layer raises an error on a write inside a read-only unit
  • Claims the speedup comes from the database rather than from skipping change tracking
  • Confuses read-only with a lower isolation level or with taking no locks
  • Believes deferred loading stops working inside a read-only unit
open as a page

When each tenant has its own datasource or schema, what must the data-access layer do before a unit of work starts and after it ends?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Resolve the tenant and bind its datasource or schema before the first statement, since the connection is fixed for the unit's life. On release, reset any state the switch set, or the pool lends the connection out still pointed there.

open as a page

A request writes a row, then a read-only unit routed to a replica reads it back stale - what causes this and what fixes it?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Replication is asynchronous, and each unit is routed on its own declared attributes with no memory that an earlier unit in the request wrote. The fix is a request-scoped pin sending later units to the primary.

open as a page

How does a routing data-access layer decide to send a unit of work to a replica, and when is that decision fixed?

level: middleimportance: should knowfreq 48%

basics

~20 s

The read-only marking declared on the unit is the usual routing signal, and it is read before the unit takes a connection. That connection then points at one server for the unit's whole life, so nothing can re-route it later.

open as a page

What happens to an object that code modifies inside a unit of work that was marked read-only?

level: middleimportance: should knowfreq 52%

basics

~10 s

Usually it is discarded: with no snapshot and no dirty check, the layer emits no update and the unit commits having written nothing. Some layers instead declare the transaction read-only and reject the write.

open as a page

Would you route every read-only unit of work to a replica by default or make replica routing opt-in per unit, and what does each choice cost?

level: principalimportance: should knowfreq 42%

basics

~20 s

Routing every read-only unit by default maximises offload but makes correctness depend on each marking also being stale-tolerant, which it never meant. Opt-in is safe and under-used; most systems vet a few paths and pin after writes.

open as a page