When a list endpoint needs only a few fields from a widely reused graph, why might a projection or a cache beat any fetch-plan change?
answer
- round trips are not the only cost
- fetch plan keeps whole objects
- shape problem versus repetition problem
- project read paths, load write paths
- reuse, low churn, staleness tolerance
basics
~20 sA fetch-plan change still loads whole objects into the tracked set for every row. A projection asks only for the columns the response shows, and a cache removes the statement entirely for data read far more often than it changes.
solid answer
~40 sChanging how a link is fetched fixes the round trips but leaves everything else: full rows for every parent and child, objects registered in the tracked set, dirty checking over data nobody edits, and a payload sized by the mapping rather than the response. When the response shows a handful of fields, a **projection** to a flat transfer model removes all of that in one statement and leaves no deferred link to trip over later. When the same values are read on nearly every request and change rarely, a **cache** removes the statement for most requests, which no fetch plan can do. Both cost something: projections do not participate in change tracking, and caches add staleness and invalidation. Reach for them when the read's shape or its reuse, not its round trips, dominates.
go deeper
Learn the distinction first: a fetch plan changes how many statements run, a projection changes how much each statement returns, a cache can remove the statement entirely.
Explain what a projection gives up — no change tracking, a shape to maintain alongside the mapping — and why that is acceptable on a read-only path.
Diagnose which cost dominates before choosing, and be ready to justify a cache with reuse, churn and staleness numbers rather than a feeling.
Own the split: which paths are read models, which stay object graphs, and how much cached state the system is prepared to keep correct.
## What a fetch-plan change does and does not remove Switching a link from deferred loading to a join or a bulk second statement fixes exactly one thing: the number of round trips. Everything else about the read is unchanged. - Every selected column of every parent and every child still crosses the wire. - Every returned row is still turned into a managed object and registered in the unit of work's identity map. - Those objects are still candidates for dirty checking at flush time, even in a read that writes nothing. - The payload is still shaped by the mapping — every mapped column — not by the response. For a read whose real cost is round trips, that is the right fix and the story ends. For a read whose cost is **volume** or **repetition**, it is a partial fix that leaves the dominant term untouched. ## When the shape is the problem: project If the response for each row is a few scalars, ask the database for a few scalars. A projection selects named columns — including columns reached through the link — into a flat transfer model. | | Objects with a fetch plan | Projection | |---|---|---| | Statements | 1-2 | 1 | | Columns transferred | every mapped column | only what the response shows | | Objects tracked | all rows, both sides | none | | Deferred links left | yes, elsewhere in the graph | none | | Usable for writes | yes | no | | Maintenance | mapping only | mapping plus a per-read shape | The second column is the honest one. A projection is a **separate read model**: it does not participate in change tracking, it cannot be edited and saved back, and when the mapping changes nobody's compiler reminds them that a projection also names those columns. Which is why the rule of thumb is to project *read paths* and keep object loading for paths that actually modify state. The design of narrow read models is a subject of its own; here it is one remedy among several. ## When the repetition is the problem: cache Some links are read constantly and change almost never — a currency table, category names, feature descriptors, a small permission map. No fetch plan can beat *not issuing the statement*. A cache earns its place when three things hold together: 1. **High reuse.** The same values are read across requests and across users, not once per report. 2. **Low churn.** Writes are rare enough that invalidation is a small, well-understood set of paths. 3. **Tolerance for staleness.** The read stays correct if the value is a few seconds or minutes old, or the cache is invalidated synchronously on write. Miss any of those and a cache is a liability: the hit rate is low, the invalidation logic spreads through the write paths, and a stale value surfaces as a bug that reproduces on one machine only. Cache scopes and where their boundaries belong are their own topic; the decision here is narrower — *is repetition the dominant cost?* ## Choosing between them, and against them Ask three questions in order. 1. **How much of what we fetch does the response use?** Small fraction, and the read never writes: project. 2. **How often is the same data read again?** Constantly, and it barely changes: cache. 3. **Otherwise?** It is a round-trip problem, and a fetch plan is the right tool. The three are not exclusive. A widely reused graph often ends up with a projection for the list read and a cache in front of the small reference lookups that the projection joins to; the object graph then serves only the write paths, which is where change tracking earns its cost. ## The failure mode to avoid Reaching for a cache first, because it is the remedy that shows the largest improvement on a graph and asks the fewest questions about the read. It also hides the shape problem permanently: the first request still pays full cost, memory grows with the key space, and the day the invalidation is wrong the bug looks like a data-integrity failure rather than a performance choice. Fix the shape first, cache the remainder if it is still worth caching, and keep the statement count in front of you while you do it.
- Why is a projection often faster even when the fetch plan already made it one statement?Because the statement returns far less. Narrow columns move fewer bytes, the database may satisfy the query from an index alone, and the layer skips materialising and registering managed objects for every row. On a wide parent with a large page, that dominates the round trip that was already removed.
- What breaks when a projection is used on a path that also writes?Nothing is tracked, so there is no change detection and no version check on save; code ends up re-loading the object or writing fields directly, which reintroduces the lost-update risk the tracked path was handling. Keep writes on the object path and let the projection serve the read.
- How do you tell a cache is the dominant remedy rather than a plaster?Look at reuse across requests, not within one. If the same small key set is requested by most requests and changes rarely, cached hits eliminate real work. If each request has its own keys, the hit rate is near zero and the cost was volume or round trips, which caching does not touch.
saying these in an interview costs you the question
- Assumes fewer statements means the read is now efficient
- Caches a graph before checking whether the read needed it
- Uses projections for paths that must also write changes
- Thinks a fetch plan reduces the columns transferred
- Adds a cache with no story for invalidation on write