How does the Unit of Work pattern relate to the Repository pattern, and what's the difference in what each one is responsible for?
answer
- Repository = retrieval facade per aggregate type
- UoW = cross-type change tracking + single commit
- repositories register loaded entities with the shared UoW
- Hibernate Session/JPA EntityManager = both roles combined
- save() on a managed entity can be a no-op until commit
basics
~20 sA Repository is where you go to find and fetch objects, like a filing cabinet drawer for one type of thing. The Unit of Work is the assistant that remembers everything you changed across all the drawers and files it all away together at the end.
solid answer
~50 sRepository and Unit of Work solve different, complementary problems. A Repository provides a collection-like interface for finding and retrieving aggregates of one type - it answers 'how do I get the Customer with id 42' - and abstracts away the query mechanics. Unit of Work tracks what changed across potentially many different repositories/entity types during one business operation and coordinates writing all of it back in a single transaction - it answers 'given everything that changed across Customers, Orders, and Payments in this operation, commit it all atomically.' In practice they're commonly paired: repositories load entities into the same shared Unit of Work so it can track them, business logic mutates the entities it got from various repositories, and a single unitOfWork.commit() call at the end persists everything. Many modern ORMs blur the line by making the Session/EntityManager itself act as both a query façade and the Unit of Work, which is why the two patterns are frequently confused as being the same thing.
go deeper
Should have a rough sense that 'finding stuff' and 'saving stuff' are different concerns, even without precise pattern names.
Should be able to state each pattern's responsibility clearly and explain, at a high level, how they're typically wired together in a request.
Should be able to explain why popular ORMs merge the two roles into one object, and correctly reason about when a save() call does or doesn't trigger immediate SQL.
Should be able to discuss the testability and architectural-boundary benefits of keeping the two separate even when the underlying framework doesn't, and make a call on when that separation is worth preserving explicitly in a codebase's own abstractions.
Repository and Unit of Work are two of the most frequently confused patterns from *Patterns of Enterprise Application Architecture*, largely because popular ORMs implement both roles inside a single object, but they exist to solve genuinely different problems and are worth separating conceptually even when a specific framework merges them. ## What a Repository is responsible for A Repository's job is **retrieval**: it presents a collection-like interface — think `findById`, `findAll`, `findOverdueOrders` — for one aggregate type, hiding the actual query mechanics (SQL, an HTTP call to another service, an in-memory list in a test double) behind a domain-shaped API. From the calling code's perspective, a Repository looks like an in-memory collection you can query, even though it's backed by a database. Its scope is deliberately narrow: one Repository typically handles one aggregate root type (`CustomerRepository`, `OrderRepository`), and its methods are about reading (and, in some formulations, also adding/removing) rather than about deciding when or how changes get persisted. ## What a Unit of Work is responsible for Unit of Work's job is different: it doesn't know how to query anything, and it doesn't care what type of object it's tracking. Its entire responsibility is **bookkeeping and transaction coordination** across a business operation — remembering which objects, of any type, were created, mutated, or marked for deletion since the operation began, and, at commit time, writing all of those changes to the database as one atomic transaction in dependency-respecting order. It is, by design, agnostic to how those objects were originally found; it just needs to be told about them. | Pattern | Its job | Its scope | |---|---|---| | **Repository** | retrieval | one aggregate type | | **Unit of Work** | bookkeeping and transaction coordination | a business operation | ## How the two compose The two patterns compose naturally: 1. repositories, when asked to load an entity, register that entity with the shared Unit of Work so future mutations to it get tracked; 2. business logic then calls into several repositories over the course of one operation, mutates the objects however the business logic requires; 3. and finally calls a single `unitOfWork.commit()` rather than separately calling `save()` on each repository. This division of labor is genuinely useful even outside frameworks: it lets you unit-test business logic against fake repositories without touching a real Unit of Work, and it lets you swap out how objects are fetched without touching how changes are committed. ## The trade-off, and what ORMs actually do The trade-off with keeping them separate is the extra indirection — two abstractions and two interfaces to learn and wire together, versus one combined object. That's exactly the trade most ORMs make in the other direction: - Hibernate's `Session` and JPA's `EntityManager` act simultaneously as a query façade (you can run `session.get(...)`, `session.createQuery(...)`) and as the Unit of Work (dirty checking, identity map, flush/commit) in one object; - Spring Data's repository interfaces are typically a thin, per-aggregate-type wrapper layered on top of that same underlying `EntityManager`. In a typical Spring/Hibernate stack, then, the 'Repository' and the 'Unit of Work' end up sharing the exact same underlying persistence context, and calling `save()` through a Spring Data repository often does nothing more than register the entity with the EntityManager's Unit of Work-style tracking, rather than triggering an immediate SQL write. ## The confusion this causes in production code The common confusion this causes in production code is calling `repository.save(entity)` and assuming an immediate, isolated database write happened, when in a typical JPA-backed Spring Data repository, `save()` on an already-tracked (managed) entity is frequently a no-op from the ORM's point of view — the entity was already dirty-tracked by the Unit of Work the moment it was mutated, and the actual SQL doesn't execute until the surrounding transaction (often demarcated by a `@Transactional` method boundary, not the repository call itself) commits and the Unit of Work flushes. Developers who don't understand that `repository.save()` and the Unit of Work's commit are two different moments in time are frequently surprised by exactly when their write actually reaches the database. ## A concrete example A concrete example is a typical Spring service method: an `@Transactional` `placeOrder` method that loads a `Customer` via `customerRepository.findById(id)`, calls a `debitBalance` method on it directly, and separately calls `orderRepository.save(new Order(...))`: - the debit on the Customer is picked up purely by Unit of Work dirty checking (no explicit save call needed at all); - the new Order is explicitly registered via `save()`; - both are written to the database together only when the transactional method returns and Spring's transaction manager triggers the underlying EntityManager's commit — which is the Unit of Work's commit, wearing the Repository pattern's clothing.
- If Repository and Unit of Work are conceptually separate, why do so many developers think they're the same thing?Because the most widely used ORMs - Hibernate's Session, JPA's EntityManager - implement both roles in a single object: it's a query façade you can call get()/createQuery() on, and it's also the Unit of Work doing identity mapping and dirty checking. Spring Data repositories are typically a thin per-type wrapper over that same shared EntityManager, so in practice the two patterns share one underlying mechanism, which blurs the conceptual line most developers encounter day to day.
- In a Spring Data JPA app inside a @Transactional method, does calling repository.save(entity) on an entity that was already loaded (and is therefore already managed) trigger an immediate database write?Not necessarily - if the entity is already managed by the persistence context, it's already being dirty-tracked, so save() is frequently a no-op from the ORM's perspective; the actual SQL write happens later, when the surrounding transaction commits and the Unit of Work flushes. This surprises developers who expect save() itself to be the moment of the database write.
- What's a benefit of keeping Repository and Unit of Work conceptually separate, even in a codebase where the ORM merges them internally?It makes business logic easier to unit-test, because you can substitute a fake/in-memory Repository for tests without needing a real Unit of Work or database at all - the business logic only depends on the Repository's collection-like interface, not on transaction-commit mechanics.
Repository is the librarian who helps you find and check out books from one specific section; Unit of Work is the front-desk ledger that quietly notes every book you've checked out, renewed, or returned across the whole library visit and settles your account once, when you leave.
saying these in an interview costs you the question
- Thinks Repository and Unit of Work are just two names for the same thing with no distinction
- Believes repository.save() always triggers an immediate database write
- Can't explain what a Repository is responsible for versus what a Unit of Work is responsible for
- Assumes every framework must implement them as separate objects
- Doesn't realize dirty-tracked entities can be persisted without any explicit save call