When is a graph RAG index not worth its build and refresh cost?
answer
- high fixed cost, amortised over queries
- label the query log first
- refresh is the forgotten bill
- stale summaries are trusted and wrong
- graph the subdomain, not the corpus
basics
~20 sA graph pays off only when index-time cost is amortised over many queries that genuinely need structure. Skip it when the query log is mostly single-fact lookups, when the corpus churns faster than it can be rebuilt, when entity density is low, or when cheaper retrieval reaches parity.
solid answer
~50 sGraph RAG front-loads cost: at least one LLM extraction pass per chunk, then entity resolution, then clustering and a summarisation pass per community per hierarchy level. That is a large one-off spend plus a recurring refresh bill, repaid only if enough queries need relationships or corpus-wide aggregates. As of mid-2026 the honest decision procedure is empirical, not architectural. Sample the real query log and label what share is genuinely multi-hop or global; if it is a small minority, the graph is subsidising queries a vector index already answers. Then build the graph over one representative corpus slice, run it against the vector baseline on that slice's eval set, and compute cost per correctly answered query rather than quality alone. Refresh is the part teams underestimate — new documents change community structure, so summaries drift stale unless you recluster, and full rebuilds on a fast-churning corpus are unaffordable. A frequent good answer is partial adoption: graph the entity-dense subdomain, leave the rest on plain retrieval.
go deeper
Know that building a graph over a corpus costs real money — a model call per chunk and more on top — and that it only makes sense if questions actually need relationships.
Break the cost into its parts — extraction, resolution, clustering, summarisation, refresh — and explain that local traversal is cheap at query time while corpus-wide answering is the expensive premium path.
Show you measure before committing: label the real query mix, pilot on a corpus slice, compare against the vector and long-context baselines, and price refresh at your genuine ingest rate rather than assuming a one-off build.
Own the portfolio call. Decide what fraction of the corpus is worth graphing, what refresh cadence the global query class actually requires, what quality gain justifies the recurring bill, and be willing to conclude that selective adoption or no graph at all is the right answer.
## Where the money goes Graph RAG's cost is dominated by index-time LLM work, and it comes in four distinct buckets. **Extraction** runs at least one model call per chunk to pull entities and relations, sometimes several passes for higher recall. This scales linearly with corpus size and is the single largest line item on a large corpus. **Entity resolution** adds blocking, similarity scoring and adjudication of the ambiguous band, where the adjudicator is often another model call or a human. **Community detection and summarisation** clusters the entity graph and then writes a summary per community — and, because the clustering is hierarchical, per community *at each level* you intend to serve. This multiplies with how many levels you keep. **Refresh** is the recurring bill and the one most often left out of the business case. Appending new documents is cheap; keeping the structure correct is not. Against this, query-time cost for local traversal is comparable to ordinary RAG, while global map-reduce answering is substantially more expensive per query. So the shape is: high fixed cost, moderate marginal cost, and an unusually expensive premium query class. ## The refresh problem A new batch of documents introduces new entity mentions, some of which are new surface forms of existing entities and some genuinely new. Extraction on the new chunks is straightforward. What is not straightforward is that new nodes and edges change the community structure — a cluster may now split, or two may merge — and every summary built from an affected community is now describing a corpus that no longer exists. The options are all compromises. Full rebuild gives correctness at full cost and is only viable on slow-moving corpora. Incremental attachment extracts from new documents, resolves against existing nodes, and appends edges without reclustering — cheap, but community summaries drift stale and global answers silently reflect an older corpus. Periodic reclustering of affected regions splits the difference and is what most production systems settle on, with a scheduled full rebuild as a backstop. The decision rule follows from the query mix: if global queries matter, staleness is a correctness issue and refresh cadence must be tight. If almost all queries are local traversals, stale summaries barely matter and incremental attachment is fine. ## Signals the graph is not worth it **The query log is mostly single-hop.** Instrument first. If eighty percent of questions are answerable from one passage, the graph is being paid for by queries that never touch it, and better chunking, metadata filters or a reranker will move the aggregate number more per unit of spend. **The corpus churns faster than it rebuilds.** An incident-response knowledge base ingesting continuously cannot support a rebuild cadence that keeps summaries meaningful. Structure that is always stale is worse than no structure, because it is trusted. **Entity density is low.** Extraction over narrative prose with few named entities and few explicit relations produces a sparse, noisy graph. The graph's value comes from real relational structure in the source material; if the documents do not contain it, extraction will hallucinate a thin version of it. **The corpus is small enough for long context.** If everything fits comfortably in a modern context window, multi-hop and even aggregate questions can often be answered by giving the model the whole corpus. That baseline should be beaten, not assumed away. **A cheaper technique reaches parity.** Query-side decomposition handles a good share of multi-hop questions without index-time cost. Structured metadata and filters answer many aggregate questions exactly, and better than any summary. If a cheap option gets within a few points on your eval set, the graph's remaining gain has to justify its whole build and refresh bill. ## How to decide, concretely Do not decide architecturally. Run the experiment. First, label a sample of the real query log by class: single-fact, multi-hop, corpus-aggregate. That distribution is the size of the prize and takes an afternoon. Second, build the graph over a representative slice of the corpus, not all of it. Measure build cost per document and extrapolate; teams are routinely surprised by a factor of several. Third, evaluate against the vector baseline on the same slice, and report **cost per correctly answered query** broken down by query class rather than a single quality score. A graph that wins overall by three points while costing six times as much has told you to adopt it selectively, not globally. Fourth, price refresh explicitly at your real ingest rate, including whatever reclustering cadence the global query class demands. The common right answer is partial adoption. Graph the entity-dense, high-value subdomain — the contract corpus, the safety-signal corpus — and route only queries about it to the graph, leaving everything else on the vector index. That captures most of the benefit at a fraction of the build and refresh bill, and it keeps the failure blast radius small if resolution quality turns out to be worse than the pilot suggested.
- Why is refresh cost usually underestimated more than build cost?Build cost is a visible one-off number that appears on a bill. Refresh is structural: new documents change community membership, so summaries built from affected communities go stale even though ingest succeeded and nothing errored. Teams price extraction on new chunks, forget reclustering and re-summarisation, and discover months later that global answers describe a corpus that no longer exists.
- What baseline should a graph RAG proposal always be measured against?Two. First, the existing vector pipeline with cheap improvements applied — better chunking, metadata filters, a reranker — since those often close much of the gap for a fraction of the cost. Second, for small corpora, putting the whole corpus in a long context window. Report cost per correctly answered query by query class against both, not a single aggregate quality score.
- How would you scope a graph RAG pilot so a negative result is still cheap?Build over one representative corpus slice rather than everything, chosen to be entity-dense and tied to a high-value query class. Measure build cost per document so you can extrapolate honestly, evaluate only on that slice, and keep routing explicit so the graph serves one query class. If resolution quality disappoints, you have spent a slice's budget and the production path is untouched.
saying these in an interview costs you the question
- Adopts a graph because the corpus is large, not because queries are relational
- Prices extraction but ignores reclustering and re-summarisation
- Assumes the graph is built once and never refreshed
- Skips the vector and long-context baselines entirely
- Compares quality without comparing cost per answered query