skip to content

A service loads 100 orders, then for each order calls a lazily-loaded getLineItems() to compute a total. The endpoint that used to take 50ms now takes 4 seconds under real data. Walk through why lazy loading causes this, and how you'd fix it.

level: seniorimportance: must knowfreq 80%

answer

  1. N parents + 1 initial query
  2. loop triggers per-row extra query
  3. invisible in code review
  4. fix: fetch join / batch size / DTO projection

basics

~20 s

Lazy loading fetches each order's line items only when asked. Looping over 100 orders and asking each one for its items triggers 100 separate database round-trips (plus the original 1 to get the orders) instead of one combined query — that's the classic N+1 problem.

solid answer

~50 s

This is the N+1 query problem: one query loads the N parent rows (orders), and then, because line items are lazily loaded, each iteration of the loop triggers a separate query to fetch that one order's items — N additional round trips, each paying full network/connection overhead for often-tiny result sets. Fixes include: replacing the lazy per-parent fetch with a single batched query (an explicit fetch join for this specific use case, or a WHERE order_id IN (...) query for all needed items at once), using batch-fetch-size settings so the ORM groups deferred loads into chunks instead of one-by-one, or restructuring the read path around a purpose-built projection/DTO query that never touches lazy associations in the first place. The general principle: keep the eager/lazy decision but make bulk-read code paths use explicit batch fetching instead of leaning on default per-object lazy triggers.

go deeper

for a junior

Should recognize that looping and calling a getter can trigger many database calls, even if they can't name 'N+1' precisely.

for a middle

Should name the N+1 pattern explicitly and describe at least one fix (batching or a joined/bulk fetch).

for a senior

Should be able to choose between fetch-join, batch-size, and DTO-projection fixes based on the specific access pattern and its trade-offs.

for a principal

Should discuss how to catch N+1 systemically (query-count assertions in CI, APM alerting) and how to set team defaults/conventions that prevent it recurring across a codebase.

## Tracing the queries The N+1 query problem is the single most common production symptom of lazy loading, and it's worth tracing exactly why it happens rather than just naming it. 1. Start with the query that loads the 100 orders — that's one round trip to the data store, query #1. 2. Each returned `Order` object, though, has its line items represented by a lazy placeholder rather than actual data (per whichever Lazy Load variant is in use — a virtual proxy or a lazily-initialized collection are the common cases in ORMs). 3. The code then loops over the 100 orders and calls `getLineItems()` on each one to sum a total. 4. Because each order's line items were never fetched by the original query, that first access on **each** order independently triggers its own load — a fresh round trip to fetch just that order's items. 5. Loop iteration 1 fires query #2, iteration 2 fires query #3, and so on, up through iteration 100 firing query #101. Hence **N+1**: the 1 initial query for the parents, plus N further queries, one per parent, for their children — 101 total queries where a well-written access pattern needs only 1 or 2. ## Why it scales into a cliff Why does this hurt so much in production but not in a small test dataset? Because the damage scales **linearly with N**, and each of those N extra queries pays a mostly-fixed per-round-trip cost that dwarfs the actual work of fetching a handful of line-item rows: - connection/thread handoff; - network latency; - driver overhead; - database query-planning overhead. With N=5 in a unit test, the extra overhead is a few milliseconds and invisible. With N=100 orders in production, each round trip costing even 30-40ms of network+parsing overhead, you get exactly the jump from 50ms to 4 seconds described in the scenario. ## Why nothing in the code looks wrong The failure is also structurally invisible in code review: the loop body just calls a getter that **looks** like a cheap in-memory field access, so nothing about the code's shape signals 'this issues a database query.' It typically first shows up not in development, where test fixtures are small, but under real traffic and real data volume, often flagged by: - a slow-query monitor; - an APM trace showing a suspicious burst of near-identical queries; - a user complaint about a slow page — well after the code has shipped. ## Three ways to fix it Fixing it means replacing the per-parent lazy trigger with a single bulk fetch for the whole batch, and there are a few standard techniques. 1. The most direct is an **eager fetch join** scoped to this one query — in JPQL/HQL terms, something like `SELECT o FROM Order o JOIN FETCH o.lineItems WHERE ...`, which tells the ORM to pull line items in the **same** query as the orders via a SQL join, producing 1 query total instead of 101. This is deliberately scoped to the one use case that needs it, rather than making the association eager everywhere — a global eager-fetch default just moves the 'load everything unconditionally' cost to every code path, including the ones that never touch line items at all. 2. A second technique is a **manual batched fetch**: after loading the 100 orders, issue one additional query with `WHERE order_id IN (...)` across all 100 order ids at once, fetch all their line items in that single query, and stitch them onto the in-memory `Order` objects yourself, or let the ORM's batching feature do this automatically (Hibernate's `@BatchSize` annotation, for example, groups deferred single-entity loads into chunks of, say, 20 at a time, turning 101 queries into roughly 6). 3. A third, often cleanest, technique for read-heavy endpoints is to bypass the lazy-loaded domain graph entirely for this access path and write a purpose-built **projection or DTO query** that selects exactly the columns needed (order id, and a pre-aggregated sum of line item totals via a JOIN and GROUP BY) in one round trip, which sidesteps the N+1 risk altogether because there's no lazy field being touched in a loop in the first place. ## The deeper lesson The deeper lesson is that lazy loading's default behavior is fine for single-object, single-field access (load one `Customer`, read one lazy field, done — one extra query is nothing) but actively dangerous inside any loop over a collection of parents, because the **same** single-object cost that was negligible once becomes N times the cost when repeated, and lazy loading, by design, gives you no visibility at the call site that you're inside such a loop.

  • Why not just make the association eager everywhere to avoid N+1 entirely?
    Because most code paths that touch the parent object never need the child collection at all, so making it eager pays the cost of loading it on every single load of the parent, everywhere in the codebase, even in the 99% of cases that don't need it. That trades a sometimes-bad N+1 cost for an always-bad unconditional cost, which is usually worse in aggregate.
  • How would you detect an N+1 problem before it reaches production?
    Enable SQL/query logging or statement counting in integration tests and assert an expected query count for a given operation over a realistic-sized dataset, or use an APM/observability tool that flags bursts of structurally-identical queries within a single request. Testing against a dataset large enough to make the O(N) pattern visible (not just N=1 or N=2) is essential, since the bug is invisible at tiny N.

Like sending a separate courier trip to the warehouse for every single item on a 100-item shopping list instead of one truck that picks up all 100 items in a single trip.

saying these in an interview costs you the question

  • Suggests eager-loading everything as the universal fix
  • Can't explain why the extra queries happen structurally
  • Doesn't know batching/fetch-join techniques exist
  • Thinks the fix is always 'just add an index'

context