skip to content

Why does flat vector RAG fail on multi-hop questions that graph RAG answers?

level: middleimportance: must knowfreq 68%

answer

  1. chunks are scored independently
  2. each document holds only half
  3. the connector is never named in the query
  4. edges are stored, not inferred
  5. seed on entity nodes, then traverse

basics

~20 s

Flat retrieval scores every chunk against the query independently, so a bridging fact that mentions neither end of the question ranks low and never enters the context. Graph RAG stores entities and relations explicitly, so traversal follows the connecting edges instead of hoping one chunk contains them.

solid answer

~50 s

A multi-hop question needs two or more facts that live in different documents and are joined by an intermediate entity the query never names. Take a pharmacovigilance corpus and the question "which studies connect this drug to arrhythmia?": the drug-to-adverse-event link sits in a safety report, and the event-to-trial link sits in a study registration. Neither document is a good embedding match for the full question, because each contains only half of it, so single-shot similarity retrieval ranks both below documents that merely mention the drug a lot. Graph RAG extracts entities and typed relations at index time, so the join is materialised as edges. At query time you seed on the entity nodes — often with vector search over node names and descriptions — then traverse one or two hops and pull the text attached to the neighbours. The connection is followed, not guessed.

go deeper

for a junior

Be able to say that ordinary RAG scores each chunk separately against the question, and that a question whose answer is split across two documents may retrieve neither. Naming the problem clearly is enough at this level.

for a middle

Explain why the bridging passage ranks low — it matches only half the query — and why the graph fixes it by materialising relations as edges at index time so retrieval can traverse rather than re-rank.

for a senior

Show you diagnose before you build: prove no single chunk holds the answer, weigh query decomposition and long-context stuffing as cheaper routes, and talk about hop depth and neighbourhood explosion as real tuning constraints.

for a principal

Own the economics. Index-time extraction cost is paid once and amortised over queries, so the case for a graph rests on what share of the real query log is genuinely multi-hop and on whether the entity structure is reused enough to repay the build.

## What "multi-hop" actually means A multi-hop question is one whose answer requires composing two or more separate facts, where at least one of the linking entities is not mentioned in the question. "Which studies connect this drug to arrhythmia?" is multi-hop: the corpus may never contain a sentence saying "drug X is linked to arrhythmia in study Y". Instead a pharmacovigilance report says "post-marketing surveillance of drug X recorded arrhythmia events", and a separate registration document says "trial NCT-whatever measured arrhythmia incidence in patients on drug X". The answer is the join, and the join exists in no single passage. Contrast this with a single-hop question — "what dose of drug X was used in trial Y?" — where one passage carries the whole answer. Single-hop is what flat retrieval is good at, and it is the majority of most query logs. ## Why single-vector retrieval structurally misses the bridge Standard RAG embeds the question once and scores each chunk by similarity to that one vector. Three things go wrong on multi-hop questions. First, **partial-match dilution**. A chunk containing only half of the question's content is, by construction, only half-similar to it. A chunk containing the drug and a great deal of unrelated pharmacokinetics can easily out-score a chunk that carries the crucial arrhythmia link, because embedding similarity rewards overall topical overlap rather than the presence of a specific bridging fact. Second, **the connector is invisible to the query**. If the two halves are joined through an intermediate entity — an adverse-event code, a study sponsor, a subsidiary — that the user never named, no amount of query-side similarity can reach for it. The query vector does not know the entity exists. Third, **independence**. Every chunk is scored on its own. Retrieval has no mechanism for "this chunk is worth including *because* that other chunk was included". Multi-hop retrieval is inherently conditional, and single-shot scoring is not. A common instinct is to raise top-k. It rarely helps: if the bridging chunk ranks 400th, moving from 5 to 50 does nothing, and the chunks you do add dilute the context and raise cost. The ranking problem is not a quantity problem. ## What the graph changes Graph RAG spends LLM effort at index time instead of query time. Each chunk is passed through an extraction pass that emits entities (drug, adverse event, trial, sponsor) and typed relations between them, each carrying the source text it came from. Extraction across the whole corpus is then merged so that the same real-world entity becomes one node no matter how many documents mention it. The result is that a relationship spanning two documents becomes a **path of length two in a single structure**. Retrieval then becomes a two-stage operation. You first locate the anchor nodes — usually by embedding-similarity search over entity names and their generated descriptions, sometimes by exact name match — and then you expand: one hop, two hops, optionally filtered by relation type. What you send to the model is the subgraph plus the source snippets attached to those edges. The bridge is now reachable by construction, because you are walking edges rather than re-ranking prose. This seed-then-traverse shape is why graph retrieval is usually described as complementary to vector retrieval rather than a replacement: vectors are excellent at fuzzy entry-point resolution, traversal is what vectors cannot do. ## Knowing you really have a multi-hop problem Before reaching for a graph, confirm the diagnosis. Take failing queries and check whether the answer exists in one retrievable chunk that simply ranked badly — that is a chunking, embedding or reranking problem, and far cheaper to fix. If instead no single chunk in the corpus contains the answer, and a human would have to read two documents and connect them, the failure is structurally multi-hop. Two cheaper routes also exist and should be considered honestly. Query-side decomposition breaks the question into sub-questions and retrieves for each, which handles many multi-hop cases without any index-time cost. And for a small corpus, putting everything in a long context window sidesteps retrieval entirely. The graph earns its keep when the corpus is too large for that, the hops are numerous, and the same entity structure is reused across many queries. ## Costs to state out loud The graph is not free: extraction is at least one LLM pass per chunk, merging entities is error-prone, and traversal that expands too greedily floods the context with weakly-related neighbours. Hop depth is a real tuning knob — two hops is typically the practical ceiling before neighbourhood size explodes and precision collapses.

  • How would you tell a multi-hop failure apart from a plain chunking or reranking failure?
    Check whether any single chunk in the corpus actually contains the answer. If one does and it simply ranked poorly, the fix is chunking, embeddings or a reranker. If a human would need to read two documents and join them through an entity the question never names, the failure is structurally multi-hop and no ranking change reaches it.
  • Does simply raising top-k solve multi-hop retrieval?
    Almost never. The bridging passage is not just outside the cut-off, it is genuinely low-scoring against the query text, so it may sit hundreds of positions down. Raising top-k adds cost, dilutes the context and increases the chance the model latches onto a topically similar but irrelevant passage, while leaving the ranking cause untouched.
  • Why is two hops usually the practical ceiling for traversal depth?
    Neighbourhood size grows roughly multiplicatively with degree, so a third hop on a well-connected entity can pull in thousands of nodes and their attached text. Beyond two hops precision collapses and the context fills with weakly-related material. Deeper needs are better served by relation-type filters, edge scoring, or a query that anchors on a more specific entity.

saying these in an interview costs you the question

  • Claims a larger top-k fixes multi-hop retrieval
  • Thinks a better embedding model alone bridges unstated relations
  • Assumes every RAG question benefits from a graph
  • Confuses multi-hop retrieval with multi-step generation by the model
  • Believes the model can infer missing links from unrelated chunks

context