skip to content

In a data-access layer with several caches, which copies are private to one unit of work and which are shared?

level: juniorimportance: should knowfreq 58%

answer

  1. not one cache but a stack
  2. scope: who can see a stale entry
  3. private map, shared copy, statement results, app store
  4. lifetimes: the unit of work, then longer

basics

~20 s

A data-access layer keeps copies at up to four scopes: the unit of work's own identity map, private and short-lived; a tier shared across units of work; results held against a statement; and a store the application drives itself.

solid answer

~40 s

Think in scopes, not in one cache. The **identity map** a unit of work keeps is private to that unit of work and is discarded with it, so nobody else ever sees those objects. A tier **shared across units of work** holds dehydrated row state keyed by identifier, lives as long as the process or cluster that owns it, and is read by every request. A **statement-result** tier holds what a query returned, keyed by the statement and its parameter values, and is voided by writes to the tables that statement read. A store the **application** fills above the mapper holds whatever shape it chose, and nothing inside the mapper knows it exists. Scope decides everything that matters afterwards: who can be served a stale entry, and how long the mistake survives.

go deeper

for a junior

Recall the four scopes by name: a map private to one unit of work, a copy shared beyond it, results held against a statement, and a store the application fills itself.

for a middle

Explain what each tier stores, live objects versus dehydrated state versus a result, and what ends an entry's life in each: closing the unit of work, an eviction, an expiry, a write.

for a senior

Show you can answer which tier served a read from evidence rather than belief: statement counts, per-tier hit counters, and what changes when the read is repeated in a fresh unit of work.

for a principal

Frame the stack as blast radius. A wrong entry in a private map ends in milliseconds; the same mistake in a shared tier is served to every request until something evicts it.

## "The cache" is the wrong noun A data-access layer almost never keeps a single cache. It keeps a small stack of them, and the useful way to tell them apart is not by what they store but by **scope** — who can see an entry — and by **lifetime** — how long the entry survives. When someone asks whether a read was served from cache, the correct first move is to ask back: *from which one?* Two of these tiers usually come with a full mapper, one is optional and off by default in most layers, and one the application builds itself. A plain query builder with no tracked set offers none of the first three, which is why the same application code caches very differently depending on the layer underneath it. ## The four scopes 1. **The unit of work's identity map.** While a unit of work is open it keeps every object it has loaded, indexed by type and identifier. Its purpose is identity — one row, one object, for the duration — and caching is a side effect of that purpose. It holds **live, tracked objects** whose changes are being watched. It is private: no other unit of work, thread or request can reach it. It dies when the unit of work closes, so a wrong entry in it lasts milliseconds. 2. **The tier shared across units of work.** An optional tier, keyed by type and identifier, that outlives the unit of work that filled it. It cannot hold live objects — an object belongs to the unit of work tracking it — so it holds **dehydrated state**: the column values, with identifiers standing in for associations. Each unit of work rebuilds its own instance from that state. Its scope is a process, or a whole cluster when the store behind it is shared; its lifetime runs until eviction, expiry, or a write that voids the entry. 3. **The statement-result tier.** Keyed by the text of a statement together with its parameter values, it remembers what a query returned. It is the only tier that can shortcut a *query* rather than a lookup, and it is the most fragile of the four: any write to a table the statement read must void it, so it earns its keep only over data read far more often than written. 4. **The application-driven store.** Anything the application fills and reads above the mapper — typically transfer models, already assembled and already shaped for a response. The mapper does not know this store exists, so nothing in the mapper will ever invalidate it. | Tier | Holds | Visible to | Ends when | Voided by | |---|---|---|---|---| | Unit-of-work map | live tracked objects | one unit of work | that unit of work closes | discarding the object | | Shared tier | dehydrated row state | every request in the process or cluster | eviction or expiry | a write to that row | | Statement-result tier | a result for one statement plus parameters | every request repeating it | eviction or expiry | a write to any table it read | | Application store | whatever shape was chosen | whoever that store is shared with | the application says so | only the application | ## Scope decides the blast radius The rows above differ in what they store, but the column that matters operationally is *visible to*: - A wrong entry in the **unit-of-work map** is visible to one request and is gone when the request ends. It is the cheapest mistake in the stack. - A wrong entry in the **shared tier** is visible to every request on that node — or on every node — until something evicts it. The same bug now has an audience. - A wrong entry in the **statement-result tier** is worse in one specific way: it can misreport not just a value but the *membership* of a result, so a row that should now appear is simply missing. - A wrong entry in the **application store** is the longest-lived, because the layer that could have noticed the write has no idea the entry exists. ## What follows from this - "Is it cached?" is not an answerable question. "Which tier answered this read, and how long would a wrong answer survive there?" is. - A repeated read that costs nothing inside one request may cost a statement in the next one, because the private tier went away with the unit of work. - Two requests reading the same row hold two objects, even when one shared entry served both. Sharing state is not sharing objects. - The higher a tier sits — further from the database, closer to the response — the more it saves per hit and the less able the layer is to tell it that something changed. The mechanics of each tier are subjects of their own: how an identity map guarantees one object per row, how a shared tier decides what to admit and what to evict, where an application-owned cache boundary belongs, and how any of them go stale. What belongs here is the comparison itself — and the habit of naming the tier before arguing about a hit rate.

  • Does a query inside a unit of work get answered by that unit's identity map?
    No. An identity map is keyed by identifier, not by statement, so the query still goes to the database or to a statement-result tier. What the map does is resolve the rows that come back: any row whose object is already tracked is handed back as that instance rather than as a second copy.
  • Two units of work read the same row through a shared tier — do they end up sharing the object?
    No. A shared tier stores dehydrated state rather than live objects, precisely because an object belongs to the unit of work tracking it. Each unit of work rebuilds its own instance from the shared entry, so the read is cheap while the two objects stay independent and separately mutable.

A note on your own desk, a note on the team whiteboard, and a note posted in the lobby all say the same thing. What differs is who reads it and how long a wrong one stays up.

saying these in an interview costs you the question

  • Says the cache as if a layer kept only one of them
  • Thinks the unit of work's map answers queries, not just identifier lookups
  • Assumes a shared tier hands the same live object to every request
  • Treats a process-local tier as visible to every other instance
  • Believes a hit anywhere means the value matched the database