When is enabling RAGFlow's RAPTOR or knowledge-graph extraction worth the indexing cost?
answer
- opt-in LLM passes at index time
- cost scales with corpus, not usage
- one helps synthesis, one helps multi-hop
- summaries are model output, and citable
- adding documents makes them stale
basics
~20 sBoth are opt-in LLM passes over your chunks at index time. Enable them when queries need information no single chunk holds — corpus-level summaries or multi-hop links across documents. For fact-lookup corpora they add cost, latency and a new source of wrong content for no gain.
solid answer
~50 sRAGFlow keeps both features as toggles because both are expensive and neither helps every corpus. **RAPTOR** clusters your chunks and has the LLM write summaries of each cluster, recursively, adding those summaries back into the same knowledge base as additional retrievable content. It pays off when users ask thematic or whole-document questions — "what are the recurring complaints" — that no individual chunk answers, because a summary node can match where leaf chunks cannot. **Knowledge-graph extraction** has the LLM pull entities and relationships out of chunks and build a graph over the knowledge base, which helps multi-hop questions that require connecting facts stated in different documents. The costs are symmetrical: LLM tokens proportional to corpus size, a much longer ingestion, a re-run whenever the corpus changes materially, and — the part people underweight — generated summaries and extracted relations are model output, retrievable and citable as if they were source text. Decide from the query distribution, not from the feature list.
go deeper
Know these are optional switches in the knowledge-base configuration that use the LLM during indexing, and that they are off by default because they cost time and money.
Explain what each produces: RAPTOR adds LLM-written cluster summaries as extra retrievable nodes, and knowledge-graph extraction builds entities and relationships from chunks.
Show you have weighed the operations: cost proportional to corpus size, much longer ingestion, staleness when documents are added, and generated text entering the retrieval surface.
Own the decision framework — characterize the query mix, baseline first, pilot on a subset, price the refresh at real churn, and state how an auditor will tell generated content from source.
## What these toggles actually are Both sit *after* parsing and chunking in RAGFlow's ingestion pipeline, and both are optional. They are the point where the pipeline stops being deterministic extraction and starts generating new content with an LLM. **RAPTOR** takes the chunks a document produced, groups similar ones, asks the LLM to summarize each group, and repeats over the summaries, producing a hierarchy from leaf chunks up to increasingly abstract summaries. The summaries are indexed alongside the original chunks, so retrieval can return either. Its configuration exposes the clustering and summarization knobs — things like a similarity threshold, a cluster limit, a token budget and the summarization prompt. **Knowledge-graph extraction** runs an LLM extraction pass over chunks to identify entities and the relationships between them, and stores the resulting graph for the knowledge base. You configure what entity types matter for your domain, which materially changes what gets extracted. ## The case for enabling them Standard chunk retrieval has a structural blind spot: it can only return text that exists. Two query shapes fall into that gap. - **Corpus-level and thematic questions.** "Summarize the main risks across these forty reports" cannot be answered by top-N chunks, because no chunk contains the synthesis and the top-N will be forty arbitrary fragments. A RAPTOR summary node *is* that synthesis, written once at index time rather than attempted at query time under a context limit. - **Multi-hop questions.** "Which supplier is connected to the component that failed inspection" requires joining a fact in one document to a fact in another. Vector similarity does not traverse; a graph does. If a meaningful share of real user queries has these shapes — and you can check that against actual query logs rather than imagination — the features earn their keep. ## The case against, which is usually stronger **Cost scales with the corpus, not with usage.** Every chunk is read by the LLM at least once, and RAPTOR reads them repeatedly as it builds levels. On a large corpus this exceeds the cost of parsing and embedding combined, and it is paid up front whether anyone asks a thematic question or not. **Ingestion time multiplies.** You have already accepted a heavy parse stage; these add LLM round-trips on top, subject to rate limits. An overnight ingest becomes a multi-day one. **They go stale.** Add a batch of documents and the existing clusters and summaries no longer reflect the corpus; a graph built before those documents has no edges into them. Re-running is not free, so you now own a refresh policy — which is an operational commitment, not a checkbox. **They introduce generated content into the retrieval surface.** This is the part that deserves the most scrutiny in a design review. A RAPTOR summary is the model's paraphrase; an extracted relationship is the model's inference. Both become retrievable, and when one is cited the user sees a citation that looks as authoritative as a quotation from the source PDF. In a regulated or high-stakes setting — the exact setting where layout-accurate parsing and citation grounding attracted you to RAGFlow — silently blending generated text into the evidence base is a real governance problem, and you should be able to distinguish generated nodes from source chunks when auditing an answer. **Quality is not guaranteed.** Summarization loses the specifics that make an answer verifiable; entity extraction is only as good as your entity-type configuration and the model's grasp of your domain, and it produces both spurious and missing relations. Neither is self-evaluating. ## How to decide, concretely 1. **Characterize the queries.** Sample real questions. Count how many are single-fact lookups, how many are thematic, how many are multi-hop. If lookups dominate, stop here — spend the budget on retrieval quality instead. 2. **Establish a baseline.** Measure answer quality with plain chunk retrieval first. Without it you cannot attribute any improvement to the feature. 3. **Pilot on a subset.** Enable on one knowledge base or one document set, not the whole corpus, and compare on the query types that motivated it. 4. **Price the refresh.** Estimate re-run cost at your real document churn rate. A feature you can afford once but not monthly is not deployable. 5. **Decide the audit story.** How will a reviewer tell a generated summary from a source quotation in an answer's citations? If you have no answer, that is a reason to hold off in a compliance-sensitive deployment. ## The principal's framing The honest position is that these features move work from query time to index time, and buy recall on question shapes that chunk retrieval structurally cannot serve — at the price of money, freshness and evidential purity. That trade is excellent for a small, stable, high-value corpus where users ask synthesis questions, and poor for a large, churning corpus of lookup-style documents. Naming that trade, rather than treating the toggles as free quality, is what the question is testing.
- Why is a RAPTOR summary node a governance concern in a citation-grounded deployment?Because it is model-generated text stored beside extracted source chunks and retrieved the same way. When it is cited, the user sees a citation that looks as authoritative as a quotation from the original document, but it is a paraphrase that may have dropped qualifiers or introduced an error. In regulated settings you need to be able to distinguish generated nodes from source chunks when auditing an answer.
- You add 200 new documents to a knowledge base that already has RAPTOR enabled. What is the operational issue?The existing clusters and summaries were computed over the old corpus, so they neither cover nor reflect the new material — retrieval can return a summary that is confidently incomplete. Getting them consistent means re-running the summarization pass, at cost proportional to the corpus. That makes refresh policy a standing commitment you should price before enabling the feature, not after.
- What would you measure before and after enabling knowledge-graph extraction to justify it?Establish a baseline with plain chunk retrieval on a fixed question set that deliberately includes the multi-hop questions that motivated the feature, plus ordinary lookups as a control. Then enable it on a pilot subset and re-run the same set. You want to see multi-hop answers improve without lookup answers degrading, and you weigh that delta against the ingestion cost and the refresh cadence your corpus churn implies.
saying these in an interview costs you the question
- Treats both toggles as free quality improvements
- Enables them on the first bulk ingest with no baseline
- Ignores that generated summaries become citable content
- Forgets they go stale when documents are added
- Assumes a knowledge graph helps ordinary single-fact lookups