skip to content

In Domain-Driven Design, what is a Repository, and why is it often described as giving client code the illusion of an in-memory collection of objects?

level: juniorimportance: must knowfreq 82%

answer

  1. add/remove/findById shape
  2. domain owns the interface, infra implements it
  3. collection metaphor, not literal collection
  4. hides mechanism, not cost

basics

~20 s

A repository is a piece of code that lets the rest of the app save and fetch domain objects (like 'add this order' or 'find order #5') without knowing or caring whether they're stored in a database, a file, or memory - it just feels like working with a list.

solid answer

~40 s

A Repository is an interface, owned by the domain layer, that mediates access to aggregates: methods like add(), remove(), and findById() that behave the way a collection would. Client code calls repository.findById(id) or repository.add(aggregate) without knowing whether the implementation runs SQL, calls a document store, or just holds objects in memory. The 'collection illusion' is Eric Evans' framing: it lets domain and application code stay expressed in domain terms - aggregates going in and out of a conceptual set - while an infrastructure-layer implementation does the real work of mapping, querying, and transactions. It's an abstraction boundary, not a literal in-memory collection.

go deeper

for a junior

Should describe the basic shape (add/remove/find) and say it hides the database from the rest of the app; doesn't need to explain layering or the dependency-inversion angle yet.

for a middle

Should articulate that the interface lives in the domain layer and the implementation in infrastructure, and should know the illusion doesn't imply cheap operations.

for a senior

Should be able to design a repository interface for a real aggregate, justify which finder methods belong on it versus a separate read path, and explain how the illusion supports testability.

for a principal

Should be able to set org-wide conventions for repository interface shape, call out when the collection metaphor is actively harmful for a subsystem, and connect the pattern to broader architectural boundaries (hexagonal/ports-and-adapters).

## What a Repository actually is A **Repository** in Domain-Driven Design is an abstraction that sits between the domain model and whatever technology actually stores data. Concretely, it's an interface — typically declared in the domain (or application) layer — with methods shaped like collection operations: - `add(aggregate)`, `remove(aggregate)`, `findById(id)` - and maybe a small number of purpose-named finder methods such as `findOverdueOrders()` The interface is implemented in the **infrastructure layer**, where the real persistence mechanics live: SQL statements, an ORM session, HTTP calls to a document store, or (in tests) a plain in-memory map. Client code — application services, domain services, other aggregates orchestrating a use case — depends only on the interface, never on the implementation, so it can say `orderRepository.findById(orderId)` and get back a fully-formed `Order` aggregate without knowing or caring how that aggregate was assembled. ## Why it is called a collection illusion The reason this is called a 'collection illusion' traces back to Eric Evans' original description in Domain-Driven Design: a repository should feel, from the caller's perspective, like a conceptual set of all aggregates of that type, already in memory, that you can add to, remove from, and query. That framing matters because it dictates the shape of the interface: - you don't call `orderRepository.executeQuery(sql)`, you call `orderRepository.findById(id)`; - you don't call `save(dto)` with an anemic data bag, you call `add(order)` with a real aggregate carrying its invariants. The illusion pushes the vocabulary of the interface toward the domain and away from the storage technology, which is the entire point — it lets someone read application-service code and understand what is happening (an order is fetched, mutated, and persisted) without being distracted by how persistence works underneath. ## The three problems it solves Why does this exist as a distinct pattern rather than just calling the ORM or database client directly? Three problems it solves: 1. **First**, it decouples the domain model from a specific persistence technology, so swapping Postgres for a document store, or introducing a cache, touches only the repository implementation, not domain or application logic. 2. **Second**, it makes the domain model testable — a domain or application service test can be handed an in-memory fake repository (a `HashMap` behind the same interface) and run with zero database, zero mocking-framework ceremony, and zero I/O latency. 3. **Third**, and most DDD-specific, it enforces the aggregate as the unit of consistency and access: if repositories only ever hand out and accept whole aggregate roots, nothing outside the aggregate can reach in and mutate an internal entity or value object directly, bypassing the invariants the aggregate root is meant to protect. ## The trade-off The trade-off is that the illusion is **never fully free**, and pretending it is causes real damage. A method that looks as cheap as `List<Order> findAll()` may, in a naive implementation, pull an entire table into memory — the collection metaphor **hides cost the same way it hides mechanism**. Query-shaped needs (pagination, filtering, sorting, aggregation, joins across aggregates) don't map cleanly onto `add/remove/findById`, so teams face a real choice: - keep the repository interface narrow and pure (only identity-based access and a few named, intention-revealing finder methods) and push reporting/search needs to a separate read-side, - or widen the repository with more finder methods and risk it slowly becoming a thin wrapper around a query builder — at which point the domain-facing illusion has effectively collapsed back into a data-access object. ## Failure modes Failure modes show up in predictable ways. 1. The most common is leaking the ORM's own entity or session objects out through repository methods instead of returning the domain's own aggregate type — callers end up depending on lazy-loading proxies, and calling a getter outside a transaction throws a lazy-initialization exception because the 'illusion' quietly depended on an open session. 2. Another is exposing something like `IQueryable<Order>` or a raw query-builder from the repository, which hands infrastructure query semantics straight to domain/application code and defeats the whole purpose of the abstraction. 3. A third is having more than one repository type per aggregate (one for `OrderHeader`, another for `OrderLine`), which breaks the aggregate-as-consistency-boundary idea, since callers can now mutate line items without going through the order root. 4. A fourth is repository methods that silently do far more I/O than their name suggests, e.g. a `findById` that triggers N+1 queries to hydrate every child collection eagerly. ## A worked example A concrete worked example: an e-commerce system's `OrderRepository` exposes `findById(OrderId): Order?`, `add(Order): void`, and one named finder, `findUnfulfilledOrdersOlderThan(Instant): List<Order>`, used by a nightly job. The checkout application service calls `findById`, invokes domain methods like `order.applyDiscount(code)` on the returned aggregate, and calls `add` (or `save`) to persist it — it never sees a SQL string, a table name, or an ORM annotation. The infrastructure-layer `JpaOrderRepository` (or a hand-rolled SQL mapper) is the only place that translates the `Order` aggregate to and from rows, and it can be swapped or heavily optimized without a single line in checkout logic changing.

  • If a repository is supposed to feel like an in-memory collection, why shouldn't its findAll() method just return every row from a table?
    Because 'collection-like' describes the interface's vocabulary, not a promise about cost - returning every row can be catastrophic for a large table. In practice teams either avoid unbounded finders entirely, add pagination/criteria parameters, or push bulk/reporting reads to a separate read-model outside the repository. The illusion covers how you talk about aggregates, not how much data moves.
  • Where should the Repository interface itself live: the domain layer or the infrastructure layer?
    The interface belongs in the domain (or application) layer, expressed in domain language and returning domain aggregates; only the implementation belongs in infrastructure. This is what lets the domain layer depend on nothing except its own abstractions, with infrastructure depending inward on the domain - the Dependency Inversion Principle applied at the persistence boundary.
  • Can a repository return a partially-loaded or projected object instead of the full aggregate?
    Not if it's meant to preserve the aggregate-collection illusion - callers need a full, invariant-consistent aggregate to safely invoke behavior on it, and returning a partial object risks that behavior operating on stale or incomplete state. Projections for display or reporting are a legitimate need, but they belong to a separate read path, not the aggregate repository.

A repository is like a hotel concierge desk for a specific kind of request (say, restaurant reservations): you ask in plain language ('book me a table for two at 8pm'), and the concierge deals with whichever restaurant's booking system, phone line, or paper diary is actually behind the scenes - you never touch that machinery directly.

saying these in an interview costs you the question

  • Repository methods return the ORM's mapped entity/proxy type instead of the domain's own aggregate type
  • Repository exposes a generic query-builder or IQueryable escape hatch to callers
  • Treats findAll() as always safe because 'it's just a collection'
  • Can't explain why the interface lives in the domain layer while the implementation lives in infrastructure
  • Confuses the repository with a generic CRUD/DAO wrapper around a single table

context