When would you choose SummaryIndex or TreeIndex over VectorStoreIndex in LlamaIndex?
answer
- match the structure to the question's scope
- one index type does zero build work
- one index type pays the LLM up front
- derived summaries can go stale
- linear per-query cost versus constant
basics
~20 sChoose by the shape of the question. VectorStoreIndex suits targeted lookups over a large corpus. SummaryIndex keeps every node in order and is for whole-document synthesis over a small set. TreeIndex spends LLM calls at build to create a summary hierarchy for large-scope questions.
solid answer
~50 s`VectorStoreIndex` embeds every node and finds the few most similar ones — the right default when the answer lives in a small slice of a big corpus. `SummaryIndex` builds no embeddings at all: it keeps nodes in a sequential list, and its default retrieval hands back every node, so cost scales with corpus size but nothing is missed. That makes it the correct choice for "summarise this contract" over a handful of documents, and a terrible one over ten thousand. `TreeIndex` spends LLM calls at construction, recursively summarising groups of `num_children` nodes bottom-up into a tree, so queries can traverse summaries instead of raw chunks — useful for broad questions over a corpus too large to stuff into context, at the price of a real build bill and summaries that go stale when data changes. `DocumentSummaryIndex` sits in between, generating one LLM summary per document.
code
python · 6 linesfrom llama_index.core import Document, SummaryIndex, TreeIndex
documents = [Document(text="Quarterly report body ...")]
summary_index = SummaryIndex.from_documents(documents)
tree_index = TreeIndex.from_documents(documents, num_children=10)go deeper
Know that LlamaIndex has more than one index type and that VectorStoreIndex is the default for finding a few relevant chunks. Be able to say SummaryIndex keeps all nodes rather than searching them.
Explain the build-cost and query-cost profile of each type and give a concrete question shape that suits each. Mention num_children and the fact that TreeIndex summarises with the LLM at construction.
Reason about staleness and operating cost: which index type survives a churning corpus, what a rebuild costs, and how you route queries between two indexes rather than compromising on one.
Own the tradeoff as a spend decision — build-time LLM bill versus per-query scan versus recall risk — and set the policy for when a corpus is promoted from an ad-hoc summary index to a maintained vector index with a rebuild pipeline.
## The choice is about scope of the question, not size of the data LlamaIndex ships several index types because "index" means different data structures. Picking one commits you to a build cost, a query cost, and a failure mode. ## VectorStoreIndex — needle in a haystack Every node is embedded and stored; queries embed the question and pull the closest handful. Build cost is one embedding per node, query cost is roughly constant regardless of corpus size, and the failure mode is recall: if the answer is spread across many chunks or phrased unlike the question, similarity search misses it. It is the right default for large corpora and specific questions, and it degrades badly on questions like "what are the themes across all of these?", where the relevant content is everything. ## SummaryIndex — read everything, in order `SummaryIndex` (the old `ListIndex`) stores nodes as a flat, ordered list. Construction does no embedding and no LLM work, so it is instantaneous and free. Its default retriever returns all nodes, and the response synthesizer then walks them, which means nothing is missed and cost scales linearly with the corpus on every single query. It is the correct structure for "summarise this document", "list every obligation in this contract", or any question whose answer genuinely requires reading the whole input — over a small, bounded set. On a large corpus it is ruinous. Its ordered structure also preserves document order, which matters when the source is narrative or sequential. ## TreeIndex — summarise upward at build time `TreeIndex` groups adjacent nodes `num_children` at a time and asks the LLM to summarise each group, then summarises the summaries, until a single root remains. Build cost is therefore LLM calls proportional to the corpus, which is far more expensive than embedding it. In exchange, a query can start at the root and descend only the branches it needs, reading summaries rather than raw text, so broad questions over large corpora become tractable without loading everything into context. The costs are real: the build is slow and expensive, and the summaries are derived data that silently go stale when a source document changes — you must rebuild the affected part of the tree, not just insert a node. ## DocumentSummaryIndex and KeywordTableIndex `DocumentSummaryIndex` writes one LLM-generated summary per document and indexes those summaries, so retrieval first picks documents by summary and then returns their nodes. It is a good fit when documents are internally coherent and the hard part is picking the right document. `KeywordTableIndex` extracts keywords per node and maps keyword to node, useful for exact-term lookup where embeddings blur distinctions. ## How to answer this in an interview State the decision rule rather than listing types: is the answer local to a few chunks, or does it require the whole corpus? Local means vectors. Whole-but-small means `SummaryIndex`. Whole-but-large means paying at build time for a hierarchy, with `TreeIndex` or per-document summaries. Then name the cost you accepted — embedding bill, per-query linear scan, or a build-time LLM bill plus staleness — because the interviewer is testing whether you know that no option is free. ## Composing rather than choosing These are not exclusive. A common production shape is a `VectorStoreIndex` over the full corpus for specific questions plus a `SummaryIndex` or `DocumentSummaryIndex` over a much smaller curated set for synthesis questions, with routing deciding which one a query hits. Indexes share a storage context happily, and each carries its own index id, so keeping two side by side is a wiring detail rather than a duplication of the source data.
- Why is SummaryIndex construction essentially free while TreeIndex construction is not?`SummaryIndex` only stores nodes in a list — no embeddings, no model calls — so building it is bookkeeping. `TreeIndex` calls the LLM to summarise each group of `num_children` nodes and then summarises those summaries recursively, so construction costs LLM tokens proportional to the corpus. You pay `SummaryIndex` at query time, repeatedly; you pay `TreeIndex` once at build time and then query cheaply.
- What breaks when a source document changes under a TreeIndex?The leaf node can be replaced, but every summary above it on the path to the root was generated from the old text and is now wrong. There is no cheap partial fix that preserves correctness of the upper levels, so a changed corpus usually means rebuilding. This staleness is the main reason teams keep TreeIndex for stable reference material and use vector indexes for data that churns.
- Can you keep two index types over the same documents?Yes, and it is a common pattern: a `VectorStoreIndex` for pinpoint questions and a `SummaryIndex` or `DocumentSummaryIndex` over the same or a narrower set for synthesis. Give each one an index id with `set_index_id()` so a shared storage context can hold both and `load_index_from_storage()` can pick the one you want. The extra cost is a second build, not a second copy of the source data.
saying these in an interview costs you the question
- Treats VectorStoreIndex as the answer to every retrieval problem
- Uses SummaryIndex over a large corpus and is surprised by query cost
- Assumes TreeIndex builds cheaply because indexing sounds like a storage step
- Thinks index type is a performance detail rather than a recall decision
- Believes summaries in a TreeIndex update themselves when a document changes