What is the N+1 problem in a GraphQL API, and how does batch loading solve it?
answer
- 1 + N round trips
- field resolvers run per-item
- defer, collect keys, dispatch once
- GraphQL can't auto-JOIN
- @BatchMapping or DataLoader
basics
~20 sWhen a query returns a list of N items and each item triggers its own lookup for a related field, you get 1 query for the list plus N extra queries. Batch loading collects those N lookups and runs them as one query.
solid answer
~40 sGraphQL resolves each field independently, so a query returning N books where each book resolves its author fires 1 query for the books plus N separate author queries — the classic N+1 problem. Because it is data-source agnostic, GraphQL cannot auto-join like SQL. Batch loading fixes this: instead of resolving each author immediately, the resolver registers the author key and returns a deferred value. Once all N books have registered their keys, the framework dispatches a single batched call (e.g. findAllById(authorIds)) and hands each field its result. In Spring for GraphQL you get this with @BatchMapping controller methods or a DataLoader registered via BatchLoaderRegistry. It collapses N+1 into 2 round trips regardless of N.
go deeper
Must be able to define N+1 and say batching turns 1+N into 2 queries.
Should connect it to Spring's @BatchMapping / DataLoader and explain deferred resolution.
Should discuss demand-driven fetching and correlating batch results back to keys.
Frames it as an architectural constraint of GraphQL's field-resolution model and the tradeoff vs eager fetching.
## The problem A GraphQL query like: ```graphql { books { title author { name } } } ``` resolves fields top-down and independently. Spring for GraphQL first runs the `books` resolver (1 query → N books). Then, for **each** book, it runs the `author` field resolver — and if that resolver naively does `authorRepository.findById(book.authorId())`, you fire **N** more queries. Total: **1 + N** = the **N+1 problem**. With 100 books that is 101 database round trips. GraphQL is **data-source agnostic** — the engine has no idea your `author` field maps to a foreign key, so unlike a hand-written SQL `JOIN` it cannot fold the lookups together automatically. You must do it. ## How batch loading fixes it The core trick is **deferred (lazy) resolution**. Instead of loading the author immediately, the field resolver: 1. Records the **key** it needs (the `authorId`) and returns a **promise/future** that is not yet completed. 2. The GraphQL engine keeps resolving sibling fields, so all N books register their author keys. 3. When the engine finishes that level and is about to dispatch, it calls your **batch function once** with the full list of keys: `[id1, id2, … idN]`. 4. Your batch function runs a single query — `findAllById(authorIds)` — and returns the results; the engine completes each book's future with the matching author. Result: **2** round trips (books + authors) instead of **1 + N**. ## In Spring for GraphQL Two built-in mechanisms: - **`@BatchMapping`** — a controller method that receives a `List<Book>` (all the parents at that level) and returns a `Map<Book, Author>` (or an ordered `List<Author>`). Spring wires it to a DataLoader for you. - **`DataLoader` + `BatchLoaderRegistry`** — register a batch function manually; the resolver calls `dataLoader.load(key)` and returns the resulting `CompletableFuture`. ## Why not just eager-fetch everything? Because the client chooses the shape at runtime. You do not know in advance whether the client will even ask for `author`, so you cannot bake a JOIN into every query. Batch loading is demand-driven: it only batches the fields the client actually requested. ## Gotchas - Batching only helps when the same field is resolved for **many parents at the same level** — a single-object query sees no benefit. - The batch result must be **correlated back** to each key (by map key or by list position); a mismatch silently returns wrong/null data.
- Why can't GraphQL just do a JOIN automatically like SQL does?GraphQL is data-source agnostic — the engine only sees field resolvers, not tables or foreign keys. It has no knowledge that `author` corresponds to a joinable column, so the developer must supply the batching strategy.
- Does batch loading help a query that returns a single book?No. Batching only collapses many lookups at the same level into one. With a single parent there is exactly one author lookup, so there is nothing to batch — the benefit scales with N.