skip to content

As the tech lead redesigning persistence for a system with a deep object graph - Order pointing to LineItem, LineItem to Product, Product to Supplier, Supplier to Address - some screens need only order totals while others need the full graph. How do you decide, association by association, whether to default to lazy or eager loading, and what alternative to loading full domain objects would you consider for read-heavy endpoints?

level: principalimportance: should knowfreq 45%

answer

  1. fetch strategy is per-query, not per-association
  2. keep mapping default lazy (fail-safe, not fail-expensive)
  3. widen fetch explicitly per use case that needs it
  4. read-heavy endpoints: skip full entities, use projections/DTOs
  5. generalizes to CQRS-style read models at scale

basics

~20 s

There's no single right default: pick lazy or eager per association based on how often each screen actually needs it, and for read-heavy endpoints, skip loading full objects altogether and query a flat projection with just the columns that screen displays.

solid answer

~50 s

Treat the mapping-level lazy/eager default as a fallback for rarely-exercised paths, not the real answer - the real fetch strategy should be decided per query, using fetch joins, batch fetching, or entity graphs to eagerly pull exactly what a given use case needs and nothing more, while leaving the default conservative (lazy) so accidentally-unscoped queries don't silently over-fetch the whole graph. Deep associations like Product to Supplier to Address are prime candidates for staying lazy by default, since most order-processing code never touches supplier or address data. For read-heavy, high-traffic endpoints - an order-summary list, a dashboard - the strongest lever isn't tuning lazy/eager at all; it's bypassing the full domain-object graph entirely and running a purpose-built projection/DTO query (or a dedicated read model) that selects only the columns that screen renders, avoiding both N+1 and over-fetching, and decoupling the read path's performance from however the write-side domain model happens to be shaped.

go deeper

for a junior

Not generally expected to answer this; a junior can be credited for recognizing that not every screen needs all the data.

for a middle

Should recognize that eager-by-default over-fetches on paths that don't need the association, even if they can't yet name entity-graph or projection mechanisms.

for a senior

Should describe per-query fetch overrides (fetch joins/entity graphs) and propose a DTO/projection for a read-heavy endpoint as an alternative to loading full entities.

for a principal

Should articulate the general policy (conservative default, per-query widening, projections/read-models for hot read paths), name the CQRS connection, and weigh the added maintenance surface against the traffic that justifies it.

## Why one global default is the wrong frame The instinct to pick one global lazy/eager default per association and apply it everywhere is the wrong frame at this scale, because the 'right' fetch behavior is a property of a specific query and its use case, not a property of the association itself. Order-to-LineItem might be eager on an order-detail page but should stay lazy on an order-list page; Product-to-Supplier might be lazy almost everywhere except a supplier-performance report. Baking one choice into the mapping forces every consumer of that association to accept the same trade-off, which is exactly how systems end up either: - **riddled with N+1 problems** (everything defaulted to lazy, no query bothers to fetch-join what it needs), - or **chronically over-fetching** (everything defaulted to eager, so a lightweight order-list endpoint silently drags along line items, products, suppliers, and addresses it never renders). ## The default policy that scales The practical default policy that scales is: keep the mapping-level default conservative - lazy - across the board, including deep, rarely-needed hops like Product-to-Supplier and Supplier-to-Address, precisely because a conservative default fails safe (an unscoped query is merely slow to lazily fetch on demand, or throws a clear LazyInitializationException-style error outside a session, which is debuggable) rather than failing expensive silently (an eager default quietly joins in four extra tables on every single query touching Order, and nobody notices until a profiler run). Then, for every specific query/use case that does need part of the graph, explicitly widen the fetch for that one query: - a `JOIN FETCH` or entity-graph hint for order-detail pulling Order+LineItem+Product in one shot, - a batch-fetch configuration for a report that touches Supplier across many products, and so on. This turns fetch strategy into a **per-query decision** made by whoever owns that query, informed by what that specific screen or endpoint actually renders, rather than a single global guess baked into the domain model that every future consumer inherits whether it fits or not. ## Skipping the domain graph on read paths The deeper architectural move for read-heavy endpoints, though, is recognizing that loading the full domain object graph at all is often the wrong tool for a read path, regardless of how carefully lazy/eager is tuned. The domain model (Order, LineItem, Product, Supplier, Address as mapped entities with identity, associations, and behavior) exists to support writes and business logic that need that behavior - applying a discount, validating a state transition, enforcing an invariant across LineItems. A dashboard or list endpoint that only ever displays five columns has no use for that behavior; loading full entities for it means paying identity-map bookkeeping, proxy creation, and change-tracking overhead for data that will be serialized to JSON and discarded within milliseconds. A dedicated projection or DTO query - a `SELECT` that names only the columns a specific screen needs, mapped straight into a flat, non-entity read type - sidesteps N+1 concerns and eager/lazy tuning entirely, because there's no lazy association left to accidentally trigger and no unused columns being carried along. At a larger scale, this generalizes into CQRS-style separation: a distinct read model (sometimes a separate, denormalized table or materialized view, sometimes just a set of projection queries against the same schema) optimized purely for the shapes read-heavy endpoints need, decoupled from however the write-side domain model is structured for correctness. ## What the approach costs The trade-off of this approach is added surface area: instead of one canonical way to load an Order, there are now several: 1. the full entity graph for write paths and detail views, 2. several fetch-join variants for specific detail screens, 3. and one or more DTO projections for list/dashboard endpoints. That means more query code to write and keep correct as the schema evolves, and a discipline requirement (someone has to actually choose the right one per endpoint rather than defaulting to 'just load the entity, it's already there'). The cost is worth paying specifically where traffic and data volume are high enough that N+1 or over-fetching would otherwise show up as real latency or database load; for a low-traffic admin screen, loading the full lazy graph and eating an occasional extra query is often the pragmatic, lower-maintenance choice. ## What this looks like in practice A concrete, recognizable pattern for this is exactly what large e-commerce and SaaS systems do: an Order aggregate with lazy associations everywhere by default, a handful of fetch-join or entity-graph variants for the checkout/detail flows that genuinely need the nested graph, and a separate, flat order-summary projection (sometimes backed by its own denormalized table populated by domain events) powering the order-list and dashboard views that get hit far more often than the detail view ever does - the classic CQRS read-model split, motivated directly by exactly this loading-strategy trade-off rather than adopted as an abstract architectural preference.

  • Why is a conservative (lazy) mapping-level default safer than an eager one, even though lazy loading is what causes N+1 in the first place?
    A lazy default fails in a way that's visible and local - a specific unscoped query is a bit slower, or throws a clear error if accessed outside a session - and it's fixable by widening that one query's fetch. An eager default fails silently and globally: every query touching that association pays the extra join or fetch cost forever, including paths that never needed it, and nothing points a developer at the problem until someone profiles a slow endpoint.
  • How do entity graphs or fetch-join hints let you override the default per query without changing the mapping itself?
    Mechanisms like JPA's @EntityGraph or an explicit JOIN FETCH clause in a query let a specific query request a wider (or narrower) fetch shape than the entity's default mapping specifies, scoped only to that one query's execution - so the mapping can stay lazy by default globally while an order-detail query, say, asks specifically for Order plus LineItem plus Product to be fetched together in that one call.
  • When would you NOT bother building a separate projection/read model for a read-heavy endpoint, and just accept the full entity graph?
    When traffic or data volume is low enough that the extra query cost is negligible in absolute terms, or when the endpoint's data needs are unstable and still changing rapidly (so a hand-tuned projection would need constant rework), or when the team lacks the capacity to maintain two parallel read paths - in those cases, the added complexity of a dedicated projection isn't worth it, and loading the entity graph with a reasonably scoped fetch join is the pragmatic choice.

Like a restaurant kitchen that keeps every ingredient in the pantry by default (lazy) and only pulls out exactly what a specific order's recipe calls for, rather than dragging the entire pantry to every table (eager-everywhere) or, for the express takeout counter, skipping the kitchen entirely and grabbing a pre-made combo off a shelf (a projection/read model).

saying these in an interview costs you the question

  • Proposes one global lazy-or-eager setting per association as the complete solution
  • Doesn't distinguish between fixing an association's default versus fixing one query's fetch strategy
  • Jumps straight to CQRS/read-models for every scenario regardless of actual traffic or complexity cost
  • Can't explain why an eager default is riskier to leave unnoticed than a lazy one
  • Assumes full domain entities are always the right thing to return from a read endpoint

context