A tracking layer loads hundreds of thousands of rows for a report and slows as it runs; what does change detection cost here, and how do you keep those reads out of the tracked set?
answer
- two copies per row, nothing collectable
- each flush rescans the whole set
- no snapshot, no compare
- untracked objects swallow writes silently
basics
~20 sTracking holds each row twice, object plus snapshot, and every flush compares the whole accumulated set, so cost grows as the read proceeds. Read untracked, or project the needed columns into an unmapped shape, so no original values are kept.
solid answer
~50 sTwo costs compound. Each loaded row is held roughly twice, as the object and as its snapshot of original values, and the tracked set pins both alive for as long as the unit of work lives, so memory climbs and the collector cannot reclaim anything. Then each flush compares every object in the set, not just recent ones, so interleaving loads with flushes makes the scan repeat over an ever-larger set and the run gets progressively slower. For a path that will never write, both costs buy nothing. The fixes are to read **untracked** so no original values are kept, or to select only the columns needed into a plain transfer shape the layer does not track at all. The caveat is that untracked objects look completely ordinary, so an assignment on one is silently discarded — mark such paths clearly as read-only.
go deeper
Remember that a layer with change detection keeps extra state per loaded row, so loading very many rows through it costs far more than the rows themselves. Reads that never write should say so.
Explain the two costs separately: the second copy plus retention for memory, and the whole-set compare per flush for time. Know that a read-only or untracked read removes both.
Diagnose before fixing: check whether memory falls when the unit of work ends, count flushes and statements on the path, and then choose between an untracked read and a projection with the trade-off stated.
Own the guardrails. Decide where in the architecture read paths are allowed to materialise mapped objects at all, and make the untracked path impossible to mistake for a writable one.
## What a large tracked read actually spends A read that materialises mapped objects through a layer with change detection spends on three things, and only the first is the data you asked for. 1. **The objects themselves** — unavoidable, whatever the layer. 2. **The snapshot** — a second record of every loaded row's values, so the compare has something to diff against. For a wide row this roughly doubles the memory of the read. 3. **Retention** — the tracked set holds a reference to every object it tracks. Nothing loaded can be reclaimed while the unit of work is alive, even if the application dropped its own reference on the previous loop iteration. On a report that walks hundreds of thousands of rows, the third point is what converts a memory cost into a failure: the set is a live root for the whole graph, so the usual assumption that processed objects become garbage does not hold. ## Why it gets slower as it runs The compare is over the **whole tracked set**, because until the layer has compared an object it does not know whether it differs. If the run only loads, that scan happens once. If it interleaves loads with anything that triggers a flush, the scan repeats — and each repeat covers everything loaded so far, not just the newest rows. The work follows the sum of the set sizes rather than the row count, which is why the first thousand rows feel instant and the hundredth thousand crawls. | symptom | mechanism behind it | |---|---| | memory climbs and never falls | the tracked set retains every loaded object and its snapshot | | throughput decays as the run proceeds | each flush rescans an ever-larger set | | collector time rises sharply | a large live set with nothing collectable | | statements appear on a read-only path | the compare finding fields that never round-trip equal | ## Getting the reads out of the tracked set - **Read untracked.** Most layers expose a read that materialises objects but records no original values — often surfaced as a read-only read. There is nothing to compare, so the second copy and the scan both disappear. This is the smallest change to the code, since the objects still have the shape the report expects. - **Select the columns you need into an unmapped shape.** A projection into a plain transfer object is not a mapped object at all, so tracking never applies, and it also stops the read from dragging in columns and associations the report never uses. This usually beats untracked mapped objects on both memory and query cost. - **Bound what is in the set at once.** Process in ranges and let the unit of work for each range end, so retention resets rather than accumulating for the whole run. - **Do not mix a write path into the report's unit of work.** Every flush the write path triggers rescans everything the report has loaded. - **Prefer streaming the result to materialising it whole**, so the row cursor rather than a list bounds the memory — this is orthogonal to tracking but compounds with it. ## Confirming the diagnosis rather than assuming it Before changing anything, establish which cost you are paying. Measure the memory held after the read completes and see whether it falls when the unit of work ends — if it does, retention is the cause. Count how many times the layer flushes during the run; a read-only report should flush zero times, and a non-zero count points at a write mixed into the path. Count the statements the report issues: UPDATEs on a read-only path mean the compare is finding phantom differences, which is a separate bug that the untracked read also happens to hide. ## The trap in the fix An untracked object is indistinguishable from a tracked one at the call site. It has the same class, the same fields, the same behaviour — and an assignment to it is quietly discarded, because there is no recorded state to compare and no membership in any set that will be flushed. That is fine while the code is the report you just wrote, and it becomes a bug the first time someone reuses the loader on a path that means to write. So make the boundary visible: keep untracked reads behind a separately named entry point, return a read-only transfer shape from it where you can, and say in the name that the result is not writable. The performance is worth having, but only if the next reader of the code can tell which kind of object they are holding.
- Why does memory not drop as the report processes and discards each row?Because the tracked set still references every object it loaded, so the application dropping its own reference changes nothing. The objects and their snapshots stay reachable for as long as the unit of work lives, which makes the set a live root for the entire read rather than a cache that can be trimmed.
- When is projecting into an unmapped shape better than an untracked read of mapped objects?Almost whenever the report needs a few columns. The projection reads less from the database, materialises smaller objects, cannot trigger the loading of associations, and is unambiguously not writable. An untracked read of mapped objects is the better choice only when the report genuinely needs the full mapped shape or existing logic that operates on it.
- How do you stop an untracked loader from being reused on a write path by mistake?Make the difference visible in the type or the name rather than in a comment: a separately named read-only entry point, and where possible a return type that is not the mapped class at all. If the result is a transfer shape with no setters, the mistake cannot compile rather than failing silently at runtime.
saying these in an interview costs you the question
- Blames the database when the cost is in-memory tracking and retention
- Thinks dropping the application's reference lets the loaded objects be collected
- Assumes flush cost depends only on how many objects were edited
- Believes an untracked read runs outside the transaction or returns stale data
- Adds memory instead of taking the read out of the tracked set
- Reuses an untracked loader on a write path and assumes edits still persist