skip to content

What is the Repository pattern, and what problem does it solve for code that needs to save and load domain objects?

level: juniorimportance: must knowfreq 70%

answer

  1. collection-like interface
  2. domain vs infrastructure seam
  3. Fowler PoEAA
  4. swap DB for in-memory in tests
  5. persistence ignorance

basics

~20 s

A Repository is an object that lets your business code save and fetch domain objects using simple method calls, like a collection, without knowing whether the data lives in a SQL database, a file, or memory.

solid answer

~40 s

The Repository pattern (from Fowler's Patterns of Enterprise Application Architecture) is an abstraction sitting between domain logic and the data source, exposing a collection-like interface: add(), remove(), findById(), findByX(). Calling code works with domain objects as if held in an in-memory collection, unaware of SQL, an ORM, HTTP, or file I/O underneath. The domain layer depends only on the Repository interface; a concrete implementation (JPA-backed, JDBC-backed, in-memory for tests) plugs in behind it. This decouples business rules from persistence technology, so you can swap databases, or unit-test logic against a fake in-memory repository, without touching domain code. It's one of the most common ways to keep the core of an application free of infrastructure concerns.

go deeper

for a junior

Should be able to say a Repository lets you save/load objects without writing SQL directly in business code, and give a one-line reason why that's useful (easier to test, easier to change database).

for a middle

Should articulate the interface/implementation split and name at least one concrete benefit (swappable persistence, or in-memory fakes for tests) with a real framework example like Spring Data.

for a senior

Should discuss where the interface should live, why persistence-ignorance matters for the domain layer, and can name at least one leaky-abstraction failure mode from real experience.

for a principal

Should be able to place the pattern in the broader architecture (hexagonal/ports-and-adapters), discuss when it's over-engineering, and speak to organizational costs of adopting it broadly across many services.

## What a Repository is A Repository is a class or interface that **mediates between the domain model and the underlying data store**, presenting a collection-like interface for accessing domain objects. Concretely: instead of a service method running a SQL query or calling an ORM session directly inline, it calls something like `orderRepository.findByCustomerId(customerId)`, which returns a list of `Order` domain objects. The service never touches a `ResultSet`, a raw SQL string, a JDBC `Connection`, or ideally even an ORM-specific type. From the caller's point of view, a Repository behaves like an in-memory `Set` or `List`: you `add()` an object, `remove()` an object, or query for objects, and it just works. The fact that `add()` triggers an `INSERT` statement, or that `findById()` runs a `SELECT ... WHERE id = ?`, is an implementation detail hidden entirely behind the interface. ## The pieces it splits into Structurally, the pattern splits into two pieces: 1. **An interface** (or abstract type) that lists the operations the domain needs — this is what business logic is allowed to depend on. 2. **One or more concrete implementations** that fulfill it against a specific storage technology. A test suite can supply **a third implementation**, backed by a plain in-memory collection, satisfying the exact same interface with none of the real machinery underneath. ## The problem it solves The problem this solves is **coupling**. Without a Repository, persistence code (SQL strings, ORM annotations, connection handling, transaction boundaries) tends to spread directly into business logic — a controller or service method builds a query, executes it, and maps rows to objects all in one place. That makes the business logic: | The symptom | What is behind it | |---|---| | **hard to read** | persistence noise drowns out the actual rule being enforced | | **hard to test** | you need a real or embedded database to exercise any code path that touches data | | **hard to change** | switching from Hibernate to a hand-written JDBC layer, or from Postgres to DynamoDB, means hunting down persistence code scattered across dozens of files | The Repository pattern draws a hard seam: domain code depends on a Repository *interface* it owns or is defined near, and a separate infrastructure module supplies the *implementation*. This is the **persistence-ignorance** goal at the heart of layered and hexagonal architectures — the domain doesn't know persistence exists. ## The trade-off The trade-off is **indirection and a small amount of extra ceremony**. Every entity that needs custom querying gets its own interface plus at least one implementation; simple CRUD screens that just need `save`/`findById`/`delete` may feel like the Repository is a needless wrapper around what an ORM's session object already provides for free. Teams that add Repositories reflexively, for every entity, in every project, sometimes end up with a thin pass-through layer that adds files without adding real decoupling — especially if the 'implementation' is just Spring Data JPA, itself already an abstraction over the database. The pattern earns its cost: - when the domain logic is meaningfully complex; - when you need to unit test business rules without spinning up a database; - when the persistence technology is genuinely likely to change or vary (e.g., you need to support both a database and an in-memory adapter for embedded/offline mode). ## Failure modes Failure modes in production tend to be about **leaky abstractions** rather than the pattern itself failing outright. - A Repository that returns lazily-loaded ORM entities (e.g., Hibernate proxies) instead of fully-hydrated domain objects can throw a `LazyInitializationException` the moment the service layer touches a field outside the original transaction — the abstraction promised 'just a domain object' but actually handed back something still wired to a live database session. - A Repository whose query methods leak pagination or filtering semantics specific to one database engine (say, cursor-based pagination that only Postgres supports efficiently) makes it impossible to swap implementations later, defeating the point. - A Repository interface that grows dozens of ad hoc `findByThisAndThatOrThatOther()` methods becomes its own maintenance burden, effectively re-implementing a query language one method at a time. - A subtler failure is **silent behavioral drift between implementations**: if the in-memory test double and the real database-backed implementation aren't held to an identical, explicitly verified contract, a test can pass against the fake for reasons that have nothing to do with the real database's actual behavior, and a bug ships that no unit test ever had a chance to catch. ## A concrete example A concrete, widely recognized example: Spring Data JPA's `CrudRepository`/`JpaRepository` interfaces are a direct, generic implementation of this pattern — you declare an interface extending `JpaRepository<Order, Long>`, optionally add derived-query methods like `findByCustomerIdAndStatus(...)`, and Spring generates the implementation at runtime via a dynamic proxy. Domain/service code depends only on the interface; the concrete backing (Hibernate + a specific database) is wired in by configuration. In testing, that same interface can be backed by an in-memory `Map`-based fake instead, letting business-rule unit tests run in milliseconds with zero database dependency — which is usually the single biggest practical payoff engineers cite when asked why they reach for this pattern.

  • Where should the Repository interface live — in the domain module or the infrastructure module?
    The interface should live with (or be owned by) the domain layer, since it's part of the contract the domain depends on, while the implementation lives in infrastructure. This follows the Dependency Inversion Principle: infrastructure depends on the domain-defined interface, not the other way around, which keeps the domain module free of any import on JPA, JDBC, or a specific database driver.
  • Is a Repository the same thing as the DAO (Data Access Object) pattern?
    They're closely related but not identical in intent: a DAO is typically table- or record-oriented and can expose CRUD methods tied closely to the underlying storage shape, while a Repository is domain-object-oriented and models a collection of aggregates with business-meaningful query methods. In practice many codebases use the terms interchangeably, but a purist Repository hides persistence details more completely than a typical DAO does.

A Repository is like a librarian's front desk: you ask for 'the book about whales' and get a book handed to you — you never see the stacks, the card catalog system, or whether the library uses the Dewey Decimal system or barcodes internally.

saying these in an interview costs you the question

  • Says a Repository is just 'a DAO with save and delete'
  • Can't explain why the interface would live separately from its implementation
  • Thinks Repositories are only useful for testing and has no other rationale
  • Confuses Repository with Active Record (where the object saves itself)

context