Spring Data JDBC has no identity map or first-level cache. What does that mean for the mapping model and how does it shape aggregate design?
answer
- No identity map / no first-level cache
- Two loads = two distinct objects
- No dirty tracking, no auto-flush, no lazy
- Save = whole aggregate atomically
- Cross-aggregate = AggregateReference (id only)
basics
~20 sSpring Data JDBC doesn't track loaded entities. Loading the same row twice returns two distinct objects; there's no dirty checking, no lazy loading, no identity map. You load a whole aggregate and save the whole aggregate explicitly, and cross-aggregate links use AggregateReference (id only).
solid answer
~50 sUnlike a JPA persistence context, Spring Data JDBC keeps **no identity map / first-level cache** and does **no dirty tracking or lazy loading**. Each query builds fresh, plain objects; loading the same row twice yields two separate instances that are not `==`. There is no automatic flush — changes persist only when you explicitly call `save`, which rewrites the entire aggregate (root plus its `@MappedCollection` children and embeddeds) in one unit of work. Because nothing is cached or proxied, the mapping model is simple and predictable: what you loaded is a snapshot, not a live managed entity. This is why relationships within an aggregate are loaded eagerly and as a whole, while links to **other** aggregates are modeled by `AggregateReference` (id only, never auto-loaded). The design consequence: draw small aggregates, treat them as atomic load/save units, and fetch across boundaries explicitly.
code
java · 16 lines// No identity map: same row, two different instances
Order a = orderRepository.findById(1L).orElseThrow();
Order b = orderRepository.findById(1L).orElseThrow();
assert a != b; // NOT the same reference (no identity map)
// No dirty tracking: mutation alone does not persist
a.setStatus("SHIPPED");
// ... nothing written to DB yet ...
orderRepository.save(a); // explicit save rewrites the whole aggregate
// Optimistic locking must be explicit because there is no managed session:
class Order {
@org.springframework.data.annotation.Id Long id;
@org.springframework.data.annotation.Version Long version; // opt-lock
String status;
}go deeper
Know you must call save; changing a loaded object doesn't auto-persist.
Explain no identity map/cache: two loads are two objects, no lazy loading.
Connect it to aggregate-as-unit-of-work and explicit @Version for concurrency.
Drive aggregate boundary design, load/save atomicity, and cross-aggregate access patterns from the no-identity-map model.
## What 'no identity map' means An **identity map** (first-level cache / persistence context in JPA/Hibernate) guarantees that within one session, a given database row maps to exactly one in-memory object, and it tracks changes to that object. **Spring Data JDBC deliberately has none of this.** Consequences: 1. **Two loads = two objects.** `repo.findById(1)` called twice returns two distinct instances; they are `equals()` only if you implemented equals on id, never reference-`==`. 2. **No dirty checking / auto-flush.** Mutating a loaded object does nothing to the database until you explicitly `save(...)`. There is no transactional write-behind that flushes changes for you. 3. **No lazy loading / proxies.** Every relationship is either loaded eagerly (within-aggregate) or not at all (cross-aggregate via `AggregateReference`). There are no lazy collections that throw `LazyInitializationException`. 4. **Objects are plain snapshots.** A loaded entity is just data; the framework does not manage its lifecycle after the query returns. ## Why the designers chose this Spring Data JDBC intentionally trades JPA's convenience for **simplicity and predictability**, organized around DDD aggregates: - **Aggregate = unit of load and save.** Calling `save` on the root persists the whole aggregate (root + nested entities via `@MappedCollection` + `@Embedded` value objects) atomically; typically it deletes-and-reinserts child rows to reconcile the collection. - **Explicit boundaries.** Because there is no identity map to silently stitch a giant object graph, you keep aggregates small and reference other aggregates by id. ## How the mapping features fit together - `@Embedded` and `@MappedCollection` are **within** an aggregate → loaded/saved as one unit. - `AggregateReference<T, ID>` is the **across-aggregate** link → only the id is stored/read, nothing is loaded, matching the no-identity-map philosophy (there's no session to hang a managed reference off of). - `NamingStrategy` governs how all of this maps to columns; it is stateless and cache-free like the rest of the model. ## Consequences / gotchas for senior design - **Concurrency:** with no dirty tracking, use `@Version` optimistic locking explicitly if you need lost-update protection. - **Reference equality:** never rely on `==` between two loads; implement `equals/hashCode` on the id if you put entities in sets. - **Collections replaced wholesale:** saving an aggregate typically rewrites child rows; huge child collections make save expensive — a reason to keep aggregates small. - **No caching means repeated reads hit the DB;** add an explicit cache (e.g. Spring Cache) if needed — the mapping layer won't do it. - **Fetching across aggregates is your job:** resolve `AggregateReference.getId()` with a second repository call; there is no auto-join, which avoids hidden N+1 but requires deliberate loading.
- Without dirty checking, how do you guard against lost updates under concurrency?Add a @Version field for optimistic locking; Spring Data JDBC checks and bumps it on save and throws OptimisticLockingFailureException on a stale write. Or use explicit DB-level locking if needed.
- Why is 'no identity map' a natural fit with AggregateReference?There is no session/persistence context to hold a managed reference to another aggregate, so cross-aggregate links store only the id and load nothing — consistent with the snapshot, load-the-whole-aggregate model.
saying these in an interview costs you the question
- Expecting mutations to auto-persist without calling save (assuming JPA dirty checking)
- Assuming two findById calls return the same cached instance
- Expecting lazy loading of relationships