skip to content

Spring Data JDBC has no identity map or first-level cache. What does that mean for the mapping model and how does it shape aggregate design?

level: principalimportance: should knowfreq 30%

answer

  1. No identity map / no first-level cache
  2. Two loads = two distinct objects
  3. No dirty tracking, no auto-flush, no lazy
  4. Save = whole aggregate atomically
  5. Cross-aggregate = AggregateReference (id only)

basics

~20 s

Spring Data JDBC doesn't track loaded entities. Loading the same row twice returns two distinct objects; there's no dirty checking, no lazy loading, no identity map. You load a whole aggregate and save the whole aggregate explicitly, and cross-aggregate links use AggregateReference (id only).

solid answer

~50 s

Unlike a JPA persistence context, Spring Data JDBC keeps **no identity map / first-level cache** and does **no dirty tracking or lazy loading**. Each query builds fresh, plain objects; loading the same row twice yields two separate instances that are not `==`. There is no automatic flush — changes persist only when you explicitly call `save`, which rewrites the entire aggregate (root plus its `@MappedCollection` children and embeddeds) in one unit of work. Because nothing is cached or proxied, the mapping model is simple and predictable: what you loaded is a snapshot, not a live managed entity. This is why relationships within an aggregate are loaded eagerly and as a whole, while links to **other** aggregates are modeled by `AggregateReference` (id only, never auto-loaded). The design consequence: draw small aggregates, treat them as atomic load/save units, and fetch across boundaries explicitly.

code

java · 16 lines
java
// No identity map: same row, two different instances
Order a = orderRepository.findById(1L).orElseThrow();
Order b = orderRepository.findById(1L).orElseThrow();
assert a != b;                  // NOT the same reference (no identity map)

// No dirty tracking: mutation alone does not persist
a.setStatus("SHIPPED");
// ... nothing written to DB yet ...
orderRepository.save(a);        // explicit save rewrites the whole aggregate

// Optimistic locking must be explicit because there is no managed session:
class Order {
    @org.springframework.data.annotation.Id Long id;
    @org.springframework.data.annotation.Version Long version; // opt-lock
    String status;
}

go deeper

for a junior

Know you must call save; changing a loaded object doesn't auto-persist.

for a middle

Explain no identity map/cache: two loads are two objects, no lazy loading.

for a senior

Connect it to aggregate-as-unit-of-work and explicit @Version for concurrency.

for a principal

Drive aggregate boundary design, load/save atomicity, and cross-aggregate access patterns from the no-identity-map model.

## What 'no identity map' means An **identity map** (first-level cache / persistence context in JPA/Hibernate) guarantees that within one session, a given database row maps to exactly one in-memory object, and it tracks changes to that object. **Spring Data JDBC deliberately has none of this.** Consequences: 1. **Two loads = two objects.** `repo.findById(1)` called twice returns two distinct instances; they are `equals()` only if you implemented equals on id, never reference-`==`. 2. **No dirty checking / auto-flush.** Mutating a loaded object does nothing to the database until you explicitly `save(...)`. There is no transactional write-behind that flushes changes for you. 3. **No lazy loading / proxies.** Every relationship is either loaded eagerly (within-aggregate) or not at all (cross-aggregate via `AggregateReference`). There are no lazy collections that throw `LazyInitializationException`. 4. **Objects are plain snapshots.** A loaded entity is just data; the framework does not manage its lifecycle after the query returns. ## Why the designers chose this Spring Data JDBC intentionally trades JPA's convenience for **simplicity and predictability**, organized around DDD aggregates: - **Aggregate = unit of load and save.** Calling `save` on the root persists the whole aggregate (root + nested entities via `@MappedCollection` + `@Embedded` value objects) atomically; typically it deletes-and-reinserts child rows to reconcile the collection. - **Explicit boundaries.** Because there is no identity map to silently stitch a giant object graph, you keep aggregates small and reference other aggregates by id. ## How the mapping features fit together - `@Embedded` and `@MappedCollection` are **within** an aggregate → loaded/saved as one unit. - `AggregateReference<T, ID>` is the **across-aggregate** link → only the id is stored/read, nothing is loaded, matching the no-identity-map philosophy (there's no session to hang a managed reference off of). - `NamingStrategy` governs how all of this maps to columns; it is stateless and cache-free like the rest of the model. ## Consequences / gotchas for senior design - **Concurrency:** with no dirty tracking, use `@Version` optimistic locking explicitly if you need lost-update protection. - **Reference equality:** never rely on `==` between two loads; implement `equals/hashCode` on the id if you put entities in sets. - **Collections replaced wholesale:** saving an aggregate typically rewrites child rows; huge child collections make save expensive — a reason to keep aggregates small. - **No caching means repeated reads hit the DB;** add an explicit cache (e.g. Spring Cache) if needed — the mapping layer won't do it. - **Fetching across aggregates is your job:** resolve `AggregateReference.getId()` with a second repository call; there is no auto-join, which avoids hidden N+1 but requires deliberate loading.

  • Without dirty checking, how do you guard against lost updates under concurrency?
    Add a @Version field for optimistic locking; Spring Data JDBC checks and bumps it on save and throws OptimisticLockingFailureException on a stale write. Or use explicit DB-level locking if needed.
  • Why is 'no identity map' a natural fit with AggregateReference?
    There is no session/persistence context to hold a managed reference to another aggregate, so cross-aggregate links store only the id and load nothing — consistent with the snapshot, load-the-whole-aggregate model.

saying these in an interview costs you the question

  • Expecting mutations to auto-persist without calling save (assuming JPA dirty checking)
  • Assuming two findById calls return the same cached instance
  • Expecting lazy loading of relationships

context