skip to content

In graph RAG, when do you route a query to community summaries instead of entity traversal?

level: seniorimportance: should knowfreq 52%

answer

  1. does the question name an anchor?
  2. aggregate questions have no top-k
  3. clusters summarised at build time
  4. map over summaries, then reduce
  5. Leiden hierarchy sets the cost dial

basics

~20 s

Route on whether the question has an entity anchor. Questions about specific named things are answered locally, by traversing that entity's neighbourhood. Questions about the corpus as a whole have no anchor to traverse from and are answered from precomputed hierarchical community summaries.

solid answer

~60 s

Graph RAG has two retrieval modes with different cost profiles. **Local** queries name concrete entities — a region, a programme, a person — so you resolve the anchor node, expand a hop or two, and answer from that neighbourhood plus its attached source text. **Global** queries ask for corpus-wide synthesis: over a decade of an NGO's field reports, "what themes recur across regions?" has no entity to anchor on, and no top-k of individual chunks represents the whole corpus. Those are served by clustering the entity graph at build time — Leiden community detection is the common choice — summarising each community, and building a hierarchy of summaries. At query time you map the question over the relevant community summaries, then reduce the partial answers into one. Routing is usually a cheap classifier or the model itself deciding whether the question names an anchor. Getting it wrong is expensive in both directions: a global query answered locally returns one region's view as if it were the corpus, and a local query answered globally fans out over many summaries and loses the specific detail.

go deeper

for a junior

Know that graph RAG can answer two different kinds of question: ones about a specific named thing, and ones about the corpus overall. Being able to name the split is enough here.

for a middle

Explain the mechanics of each mode: anchor resolution plus hop expansion for local, and build-time clustering into communities with LLM-written summaries that get mapped over and reduced for global.

for a senior

Show you handle routing and its failure modes — especially that a global question answered locally returns a fluent, cited answer from one corner of the corpus — and that you evaluate global answers on coverage, not just fluency.

for a principal

Own the cost model. The hierarchy level, how often clustering is re-run, whether global mode is a default path or an explicit analysis surface, and what caching exists are budget decisions that determine whether global answering is affordable at your query volume.

## Two questions that look alike and are not Consider a decade of an NGO's field reports. "What did the water programme in the eastern districts report about borehole maintenance?" and "What themes recur across all regions over the decade?" are both retrieval questions over the same corpus, but they need structurally different retrieval. The first has an **anchor**: named entities that exist as nodes. The second does not. It is a question about the corpus in aggregate, and there is no chunk, and no small set of chunks, that represents an aggregate. This is the local/global split, and it is the main architectural idea a graph index buys you beyond multi-hop traversal. ## Local queries: anchor, expand, answer A local query is answered by resolving the entities in the question to nodes — usually via similarity search over node names and their generated descriptions, since users rarely type canonical names — then expanding into the neighbourhood. You collect the neighbouring entities, the relations connecting them, and the source text snippets each relation was extracted from, and hand that bundle to the model. The knobs are hop depth, how many neighbours to keep per hop, and whether to filter by relation type. Two hops is generally as far as you go before the neighbourhood grows faster than the context budget. Local retrieval is cheap: it touches a small part of the index and reads a bounded amount of text. ## Global queries: cluster, summarise, map-reduce A global query cannot be answered by reading a neighbourhood, and reading the whole corpus per query is unaffordable. The standard construction moves the work to build time. After entities and relations are extracted, the graph is partitioned into **communities** — densely interconnected groups of entities — using a modularity-based clustering algorithm, most commonly Leiden. Because these algorithms produce a hierarchy, you get communities at several granularities: fine-grained clusters at the leaves, broad thematic groupings near the root. An LLM then writes a summary of each community from its member entities, relations and source snippets, and higher-level summaries are composed from the level below. At query time, a global question is answered map-reduce style: the question is applied against each community summary at the chosen level, producing partial answers with a self-reported usefulness score; low-value partials are discarded and the rest are reduced into a final answer. The corpus is therefore covered in aggregate — every region's reports contributed to some community summary — without reading it all per query. **Choosing the hierarchy level is the cost dial.** A high level means a handful of very broad summaries: cheap, fast, but coarse. A low level means many detailed summaries: better coverage and specificity, and many more LLM calls per question. This is a per-query-class decision, not a global constant. ## Routing between the two Routing is typically done by a small classifier, a prompt asking whether the question names specific entities and expects a specific fact, or by exposing both as tools to an agent and letting the model choose. Signals that a query is global: superlatives and aggregates ("most common", "overall", "recurring", "across"), an absence of resolvable entity mentions, and a request for themes rather than facts. Get routing wrong and the failure mode is characteristic and hard to spot. A **global query answered locally** returns a confident, well-sourced answer drawn from one anchor's neighbourhood — one region, one programme — presented as if it characterised the corpus. Nothing looks broken; the answer is simply unrepresentative. That is the more dangerous direction, because the output is fluent and cited. A **local query answered globally** fans out over community summaries, costs many times more, and returns generalities where the user wanted a specific fact, because summarisation has already discarded the detail. Because misrouting is silent, evaluation should include a labelled routing set and should score global answers on coverage — did the answer draw on multiple regions or only one — not just on fluency. ## Practical notes Hybrid routing is common: run local retrieval and, if the anchor resolution is weak or the answer is unsupported, fall back to global. Global answering is expensive enough that many systems cache answers per question class, or restrict global mode to an explicit "analyse the corpus" surface rather than the default chat path. And global mode is only as good as the community summaries, which are built once — if the corpus has grown substantially since the last clustering pass, global answers quietly reflect the older corpus.

  • What does a global query answered by local retrieval look like to the user?
    Entirely plausible, which is the danger. The system resolves whichever entity it can, traverses that neighbourhood, and returns a fluent, well-sourced answer that describes one region or programme as though it characterised the whole corpus. Nothing errors. Catching it needs a coverage metric — how many distinct communities or sources contributed — rather than a quality judgement on the prose.
  • How do you pick which level of the community hierarchy to answer a global query from?
    Treat it as a cost-versus-specificity dial and tune it per query class on an eval set. High levels give a few broad summaries: fast and cheap, but generic. Low levels give many detailed summaries: better coverage and concrete detail, at many more LLM calls per question. Measure answer quality and cost at two or three levels rather than guessing.
  • Why can't you answer a global question just by retrieving a very large top-k of chunks?
    Because similarity ranking selects the chunks most like the question, not a representative sample of the corpus, so the sample is biased towards whatever phrasing the question used. Even a large k covers a vanishing fraction of a big corpus and skews to over-represented documents. Community summaries are built to cover the corpus by construction, which is exactly the property ranking lacks.

saying these in an interview costs you the question

  • Treats every graph query as neighbourhood traversal
  • Thinks a large top-k can answer a corpus-wide aggregate question
  • Assumes community summaries are computed per query
  • Ignores that misrouting a global query fails silently
  • Claims one hierarchy level suits every global question

context