Should a Repository interface return database-framework types like an ORM's managed entity proxy or a framework's Page<T> pagination wrapper, or should it return plain domain objects? What breaks if persistence-framework types leak through the interface?
answer
- managed entity vs plain domain object
- lazy-initialization exception after session closes
- dirty-checking causes surprise write-back
- Page<T> leakage ties callers to the framework
- read-side CQRS exception is often tolerated
basics
~10 sA Repository should return plain domain objects, not framework-specific wrappers. If it leaks framework types, callers end up depending on that framework everywhere, which defeats the whole point of hiding persistence details.
solid answer
~40 sIdeally a Repository returns plain domain objects, detached from any persistence-framework machinery, so business code has zero compile-time dependency on the ORM or driver. In practice many teams accept some leakage — returning JPA-managed entities directly, or a Spring Data Page<T> — because a mapping layer for every entity is extra work. The cost is real: managed entities can throw lazy-initialization exceptions once the session closes, couple the domain layer to the ORM, and make swapping persistence technology harder since the leaked type's contract is baked into calling code. The pragmatic middle ground: map to domain objects at the boundary for anything the domain touches, but tolerate framework types in a read-only reporting layer decoupled from the write-side model.
go deeper
Should recognize that a Repository ideally returns plain objects rather than framework-specific ones, even if unfamiliar with the specific failure modes.
Should be able to name at least one concrete risk of returning a managed entity directly, such as a lazy-loading exception.
Should articulate both major risks (lazy-init failures and surprise dirty-checking write-back), know a mapping tool like MapStruct, and be able to judge when leakage is acceptable versus not.
Should be able to set an architectural policy distinguishing write-side domain Repositories from read-side query layers, and justify that split in terms of concrete production incidents or maintainability costs, not just theoretical purity.
## The purist form The purist form of the Repository pattern says the interface's return types should be **plain domain objects** — ordinary classes/data classes the domain model defines, with no annotations, no lazy proxies, no framework-specific base classes. In that world, a `CustomerRepository.findById(id): Customer?` returns a `Customer` that is indistinguishable, from the caller's perspective, from a `Customer` constructed by hand in a unit test. The implementation is responsible for mapping whatever the underlying storage technology hands back (a JPA entity, a JDBC `ResultSet` row, a JSON document from a NoSQL driver) into that plain domain object before returning it, and for mapping the domain object back into storage-appropriate form on `save()`. ## The shortcut most codebases take In practice, this purity has a cost, and many production codebases don't fully pay it. - The most common shortcut is letting an **ORM's managed entity class double as the domain object** — e.g., a JPA `@Entity`-annotated `Customer` class is both what Hibernate persists and what the `CustomerRepository` returns directly. This saves writing and maintaining a separate mapping layer (domain object ↔ persistence entity) for every type in the system, which for large systems with dozens or hundreds of entities is genuinely significant effort. - Similarly, exposing a **framework's pagination wrapper** directly (Spring Data's `Page<T>`, which bundles the result list with total-count and page-metadata) from a Repository method saves re-implementing pagination metadata by hand. ## What leakage costs The trade-off shows up as leakage, and leakage has concrete, observable costs. 1. **First, compile-time coupling**: once `Page<T>` or a JPA-managed entity type appears in a Repository interface's signature, every caller — service classes, controllers, test code — now has a compile dependency on that framework, even code that has nothing to do with persistence and shouldn't need to know it exists. Swapping ORMs, or introducing a second Repository implementation backed by a different technology, now requires touching every call site that used the leaked type, not just the Repository implementation, which defeats a large part of the pattern's original promise. 2. **Second, runtime fragility**: a managed entity returned directly from a Repository is often still attached to (or was attached to) a persistence context / session. If a service method accesses a lazily-loaded field after that session has closed — for instance, in a different thread, or after the request's transactional boundary ended — it triggers a `LazyInitializationException` or equivalent, a bug class that's notoriously easy to introduce accidentally and confusing to debug for anyone unfamiliar with the ORM's session lifecycle. 3. **Third, unintended write-back**: a managed entity is often subject to the ORM's dirty-checking, meaning if calling code mutates a field on what it thinks is 'just a returned object,' the ORM may silently persist that mutation on the next flush, even though the caller never called `save()` — a surprising, hard-to-trace form of implicit write path that a plain, detached domain object would never exhibit. ## The pragmatic middle ground The pragmatic middle ground most senior engineers land on is **contextual rather than absolute**: | Which part of the system | What the engineers land on | |---|---| | the domain/write side of the system — anywhere business rules are enforced and objects are meant to be mutated, validated, and saved back | mapping to plain domain objects at the Repository boundary is worth the mapping-layer cost, because the failure modes above (implicit write-back, session-lifecycle bugs, framework coupling in business logic) are genuinely expensive in a part of the system meant to be trustworthy and stable | | a read-only reporting or query-side layer — where the goal is efficiently shaping data for a screen or export, and there's no risk of an accidental write-back because nothing is ever saved | many teams explicitly accept a separate, framework-aware query layer (sometimes following CQRS-style separation) that's allowed to return `Page<T>` or projection-specific DTOs directly, precisely because that layer was never meant to be persistence-ignorant in the first place | ## A real-world instance A concrete, widely encountered real-world instance: Spring Data JPA repositories that extend `JpaRepository<Customer, Long>` where `Customer` is the `@Entity` class itself are the default, out-of-the-box pattern in most Spring tutorials and a large fraction of production Spring codebases — convenient to start with, but exactly the leaky version described above. Teams that outgrow it typically introduce a mapping layer (MapStruct-generated mappers, or hand-written `toDomain()`/`toEntity()` functions) so the `@Entity` class becomes purely an infrastructure-layer JPA mapping artifact, and a separate, framework-free domain class is what the Repository interface actually exposes to the rest of the application — a refactor commonly triggered by exactly one of the bugs described above (a `LazyInitializationException` reaching production, or an unexpected silent update) rather than being planned from day one.
- What tool do teams commonly use to reduce the boilerplate cost of mapping between a persistence entity and a plain domain object?MapStruct is a widely used compile-time code generator for exactly this: given interface method signatures describing the mapping (entity to domain object and back), it generates the mapping implementation at build time, avoiding both hand-written boilerplate and the runtime reflection cost of libraries like ModelMapper.
- Is it ever acceptable to expose a framework's Page<T> type from a Repository interface?Many teams accept it specifically on read-only query paths that are explicitly outside the core write-side domain model — for example, a reporting or search endpoint's dedicated query Repository — since there's no risk of the dirty-checking or session-lifecycle issues that make leakage dangerous on the write side; the key is that this is a deliberate, scoped exception rather than the default everywhere.
Handing out a managed ORM entity is like giving someone a live TV remote still paired to your set-top box instead of a written note of the channel number — they can accidentally change your channel (write-back) or the remote stops working the moment they leave the room (lazy-init failure after the session closes).
saying these in an interview costs you the question
- Says returning the ORM entity directly is always fine with no mention of lazy-initialization or dirty-checking risk
- Can't explain what a LazyInitializationException is or why it happens
- Treats mapping between entity and domain object as pure unnecessary busywork with no acknowledgment of the coupling/write-back trade-off
- Has no concept of separating read-side and write-side leakage tolerance