When a data-access layer serves a read, in what order does it consult its cache tiers?
answer
- the read decides the path, not the layer
- identifier lookups and queries take different paths
- keyed by identifier, not by statement
- fall-through down is also fill-down back up
basics
~20 sA lookup by identifier checks the unit of work's map, then the shared identifier-keyed tier, then the database, filling each on the way back. A query skips both, reaching a statement-result tier or the database; a raw statement consults nothing.
solid answer
~40 sThe order depends on the read, because the tiers are keyed differently. A **lookup by identifier** can build every key: the unit of work's identity map first, then the tier shared across units of work if the type is admitted to it, then the database — and each tier it passed is filled on the way back up. A **query** cannot be keyed by identifier at all, so it goes to a statement-result tier if the layer has one, otherwise straight to the database; the rows that come back are then resolved against the identity map, and a row whose object is already tracked comes back as that instance rather than as fresh values. A statement sent through the layer's raw escape hatch consults nothing and fills nothing.
go deeper
Remember that a read by identifier can be answered without touching the database, while a query normally cannot, because the caches nearest the code are indexed by identifier.
Walk both paths out loud, including the fill on the way back and the resolution of returned rows against the identity map. That resolution step is what most candidates leave out.
Use the order diagnostically: which tier served a read is deducible from whether the saving disappears at the end of a request, at a write, or not at all.
Judge the stack by what each hit actually saves against what each costs to keep correct, and be explicit about which tiers your layer really implements rather than assuming a canonical order.
## Two reads, two different paths "In what order does the layer check its caches?" has no single answer, because the tiers are keyed differently and a read can only consult a tier whose key it is able to build. A lookup by identifier can build every key in the stack. A query cannot: nothing in an identifier-keyed tier knows which rows satisfy `status = 'OPEN'`. So there are two paths down the stack, and knowing which one a given read takes is most of the skill. ## Path one — a lookup by identifier 1. **The unit of work's identity map.** If the object is already tracked, it is returned as it stands, usually with no statement at all. This is why the same lookup repeated inside one request costs nothing after the first. 2. **The tier shared across units of work**, when one is enabled and the type is admitted to it. A hit yields dehydrated state, from which the layer builds an object and puts *that object* into the unit of work's map. 3. **The database.** A `SELECT` by primary key. On the way back up the row becomes an object, the object goes into the map, and — if the type is admitted — its state goes into the shared tier. The fall-through is therefore also a fill-down: the tiers a read passed through are populated on the way back. ## Path two — a query A query is identified by its text and its parameter values, so the identifier-keyed tiers cannot answer it at all. 1. **The statement-result tier**, if the layer offers one and the statement is marked as cacheable. Layers that offer it typically hold a result's identifiers rather than its rows, so even a hit has to turn each identifier back into an object. 2. **The database**, otherwise — always. 3. **The identity map, on the way back.** This is the step candidates miss. Every returned row is resolved against the map: when the row's object is already tracked, most mappers hand back *that instance* and do not overwrite its fields with the freshly-read column values. That last rule is why a query can return an object whose fields do not match the row the database just sent. It is not a cache hit — the statement really ran — it is identity being preserved. | Read | Consults | Fills | |---|---|---| | lookup by identifier | map, then shared tier, then database | map, shared tier | | query | statement-result tier, then database, then map resolution | map, statement-result tier | | raw statement | nothing | nothing | | set-based write | nothing | nothing, and invalidates nothing by itself | ## The escape hatch consults nothing Every layer offers some way to send a statement as written. That read goes straight to the database: no tier is asked on the way down, no tier is filled on the way back, and the values arrive untracked. It is the one read you can trust to reflect the row as stored, and the one read that can never be reused by anything. ## Why the order is a debugging tool - A read that is fast the second time **within** a request but slow the second time **across** requests was answered by the private map, not by any shared tier. - A read that stays fast across requests and goes slow after any write to the table was answered by a statement-result entry. - A read that returns your own not-yet-committed change was answered by identity resolution, not by a cache at all. - A read that returns a value no tier inside the mapper should still hold points at the store the application drives, which no write through the mapper will ever invalidate. ## What a hit actually saves A hit in the map saves the whole round trip and the object build. A hit in the shared tier saves the round trip but still pays for building an object per unit of work. A hit in the statement-result tier saves running the query but may still pay one lookup per identifier it returned. The tiers are not interchangeable units of "cache" — they save different work, in decreasing amounts as you go down. ## The honest caveats Layers differ, and overstating the order is an easy way to be wrong. Some consult a shared tier only for types explicitly admitted to it; some skip it when the lookup is part of a larger fetch plan; some offer no statement-result tier at all; and a query builder with no tracked set has no map to resolve against, so every read hands back a fresh set of values. Describe the shape of the order, then say which parts the layer in front of you actually implements.
- Why can a query return an object whose field values differ from the row just read?Because the identity map wins the resolution step. If a returned row belongs to an object the unit of work already tracks, most mappers hand back the tracked instance and discard the freshly-read values for it, so an in-memory change you have not flushed is still visible. That is identity being preserved, not a cache hit.
- What happens on the way back up when nothing in the stack held the object?The statement runs, the row becomes an object, the object is put into the unit of work's map, and — if the type is admitted to it — its dehydrated state is put into the shared tier. The fall-through downwards is matched by a fill on the way back, which is why the second read of the same identifier is cheap.
- Where does a statement sent through the layer's raw escape hatch fit in this order?Outside it entirely. The statement is sent as written: no tier is consulted going down, none is filled coming back, and the values arrive untracked. That is precisely why such a read cannot serve you something stale, and equally why nothing can ever be reused from it.
saying these in an interview costs you the question
- Thinks every read consults every tier in the same order
- Believes a query is answered from the unit of work's identity map
- Assumes a hit anywhere means no statement was sent
- Thinks tiers are only read during a read, never written
- Expects a raw statement to consult or refresh the layer's caches