In a data-access layer with an identity map, what does loading the same row twice inside one tracked set give you?
answer
- one row, one object
- keyed by type plus identifier
- by-key hit needs no statement
- scope is the tracked set, not the process
basics
~20 sYou get the same object both times. A tracked set's identity map, keyed by type plus identifier, makes one row exactly one object for the life of that set; a by-key lookup it already holds needs no statement.
solid answer
~50 sA unit of work keeps an **identity map**: a lookup, keyed by type plus identifier, of every object it has already materialised. When a second load produces the same key, the layer hands back the object it is already holding instead of building a second one, so one row is exactly one object for the life of that set. Two consequences follow. First, a lookup **by key** for an object already in the map can be answered straight from memory, with no statement sent — a query on non-key columns still travels, and its returned rows are reconciled afterwards. Second, an edit you made after the first load is visible to whoever loads that row again, because there is only one instance to edit. The guarantee is scoped to one tracked set: a second set has its own map and its own object for the same row.
go deeper
Recall the one-liner: one row becomes one object inside one tracked set, so the second load hands back the same instance. Knowing that much already explains most of what a mapper does that plain row reading does not.
Explain the two separate effects — instance reuse on every path, statement avoidance only on a by-key hit — and why a query on other columns must still travel to the database before its rows are reconciled.
Show where the scope boundary bites in running systems: a held instance is as old as its load, and closing or clearing the set throws the map away, so freshness has to be arranged deliberately.
Frame it as a contract you are choosing: instance reuse buys coherent edits inside one operation and costs you freshness and cross-set equality. Decide where that trade is worth paying and where reads should bypass tracking entirely.
## The rule the map enforces A **tracked set** — the unit of work a data-access layer holds between load and commit — maintains an **identity map**: an in-memory lookup whose key is *object type plus identifier* and whose value is the object already materialised for that row. The rule it enforces is short: **one row becomes exactly one object for the life of that set.** Without it, every read would build a fresh object from the columns just returned. Load the same customer through two different code paths in one operation and you would hold two objects, edit one, and silently lose the other edit at write time. The identity map is what makes "the customer" a single thing inside one operation instead of a per-read snapshot. ## What goes into the map, and when - an object materialised by a **lookup by key**; - every object materialised from a **query's** returned rows; - an object that becomes tracked when it is scheduled for insert, so that later reads inside the same set see that same instance; - in most layers, objects reached through a **lazily loaded association** as they are resolved. Rows that never become tracked objects — a projection into a transfer shape, an aggregate, a raw column read — do not enter the map, and are not reconciled against it. That is a common surprise: a projection can report values that differ from the tracked object you are holding. ## The two paths, and which one skips the database | Path | Statement sent? | What comes back | |---|---|---| | Lookup by key, object already in the map | Usually none | The held instance, as it currently stands | | Lookup by key, not yet held | One read | A new object, then placed in the map | | Query on other columns | Always — the predicate runs in the database | Rows reconciled against the map; already-held keys yield the held instance | | Projection into a transfer shape | Always | Plain values; nothing enters or consults the map | The "no statement" case is only the by-key path, and only because the layer can answer the question — *which object is row 7?* — from what it already holds. It cannot answer a predicate over columns from memory, so a query always travels. ## What the guarantee is not - **Not a cache that outlives the set.** When the set is closed or cleared, its map goes with it; the next set starts empty and re-reads. Tiers that survive across sets are a different mechanism. - **Not engine-level read stability.** The two reads agree because they returned the *same object*, not because the database promised a stable view. Whether another committed writer can change the row underneath you is decided by the isolation level, not by the map. - **Not a freshness promise.** A held object reflects the columns as of the moment it was materialised, plus your own edits. It can be stale, and the map is exactly what keeps handing you the stale instance. - **Not shared across sets or threads.** Two sets open at once hold two objects for one row. A tracked set is normally used by one thread at a time. ## Why interviewers ask it first Because almost every later surprise in a mapper is this rule showing through: a query that returns your edited values rather than the database's, reference equality that works in a test and fails in production once a second set is open, a "refresh" that appears to do nothing because the layer served the held instance again. Candidates who can state the rule and its scope — *one row, one object, for the life of one set* — can usually derive the rest on the spot. ## Saying it well A strong answer names the map, gives its key (type plus identifier, not just the identifier — two different types can share the number 7), states the scope (one tracked set), and separates the two effects: **instance reuse** on every path, and **statement avoidance** only on the by-key path. Then it adds the honest caveat that data-access layers differ in how aggressively they populate the map and in whether a re-read can be forced to overwrite what is held — the caller normally has an explicit way to ask for a re-read or to discard the object so the next read is real.
- Why is the map keyed by type plus identifier rather than by the identifier alone?Identifiers are only unique within their own table, so the number 7 can name an order and a customer at the same time. Keying by type as well keeps those two rows in separate slots. It also lets the layer answer a by-key lookup without knowing anything about the row beyond the class the caller asked for.
- If the map avoids a statement on a by-key hit, why does a query on a non-key column still hit the database?The predicate has to be evaluated against every row in the table, and the set only holds the handful of objects it has already read. Evaluating in memory would silently miss rows nobody has loaded. So the statement runs, and the map is applied to the rows that come back rather than to the question going out.
- What happens to the map when the tracked set is closed or cleared?It is discarded along with the set. Objects that were in it become detached: still usable as plain objects, but no longer reconciled with anything, and no longer reused by the next set. The next set starts with an empty map and re-reads the rows it needs.
A cloakroom ticket. Hand in the same number twice and you get the very same coat back, not a copy — and if the coat is already on the counter, the attendant does not walk to the rack.
saying these in an interview costs you the question
- Says two loads of one row give two equal but distinct objects
- Thinks the map survives closing the set, like a cache
- Believes any query can be answered from the map without a statement
- Assumes the map guarantees the object still matches the row in the database
- Treats the map as process-wide rather than per tracked set