skip to content

In LlamaIndex, how do the compact, refine and tree_summarize response modes differ?

level: middleimportance: must knowfreq 72%

answer

  1. one call per chunk, in sequence
  2. pack first, then do the same thing
  3. bottom-up summaries of summaries
  4. which one can run concurrently
  5. the default is the cheap one

basics

~20 s

refine calls the LLM once per node, each call improving the running answer. compact, the default, packs nodes into as few context-sized prompts as possible before refining, so it makes far fewer calls. tree_summarize summarizes batches and then summarizes the summaries, bottom-up.

solid answer

~60 s

All three are response synthesizers — the half of a query engine that turns retrieved nodes into prose — selected with `response_mode` on `as_query_engine` or via `get_response_synthesizer`. **refine** sends the first node with the QA template, then walks the remaining nodes one at a time through the refine template, each call revising the previous answer: N nodes means N sequential LLM calls, accurate but slow and expensive. **compact** is the default and is refine with packing — it stuffs as many node texts into each prompt as the context window allows, so ten small nodes often collapse into a single call, and only overflow triggers a refine step. **tree_summarize** ignores the refine chain entirely: it summarizes batches of nodes with the summary template, then summarizes those summaries, recursing until one answer remains. It sees every node with equal weight, which suits "summarize all of this" questions, and its batches at each level can run concurrently with `use_async=True`. For ordinary top-k question answering compact is the right default; reach for tree_summarize when the question is genuinely over the whole retrieved set.

code

python · 13 lines
python
from llama_index.core import get_response_synthesizer
from llama_index.core.query_engine import RetrieverQueryEngine
from llama_index.core.response_synthesizers import ResponseMode

synthesizer = get_response_synthesizer(
    response_mode=ResponseMode.TREE_SUMMARIZE,
    use_async=True,
)
engine = RetrieverQueryEngine(
    retriever=index.as_retriever(similarity_top_k=30),
    response_synthesizer=synthesizer,
)
print(engine.query("Summarize the recurring risks across all filings."))

go deeper

for a junior

Be able to name the three modes and say that refine works one chunk at a time, compact packs chunks together first, and tree_summarize builds summaries of summaries. Know compact is the default.

for a middle

Explain the call counts — N sequential calls for refine, usually one for compact, one per batch per level for tree_summarize — and which templates each mode uses.

for a senior

Show that you pick the mode from the question shape and the latency budget, know that only tree_summarize parallelizes with use_async, and can explain packing dilution versus summary lossiness as distinct quality failures.

for a principal

Own the mode as a cost lever: per-query call counts multiplied by traffic often dominate the RAG bill, and a mode change is a spend and latency decision that needs the same review as a model change.

## Where the mode sits A LlamaIndex query engine has two halves: retrieval, which produces nodes, and synthesis, which produces prose. The response mode configures the second half. You set it either inline — `index.as_query_engine(response_mode="compact")` — or explicitly with `get_response_synthesizer(response_mode=ResponseMode.TREE_SUMMARIZE)` and hand the result to a `RetrieverQueryEngine`. The choice does not change *what* was retrieved; it changes how many LLM calls happen, in what order, and how much of each node the model ever sees. ## refine The original strategy. The first node goes into a QA prompt (`text_qa_template`) and yields a draft answer. Each subsequent node goes into a refine prompt (`refine_template`) together with the current answer, and the model is told to improve the answer using the new context, or to repeat it unchanged if the new context adds nothing. Properties: every node gets its own dedicated prompt, so nothing is truncated away by packing. Cost is N calls for N nodes, strictly sequential — call k+1 needs the output of call k, so there is no parallelism to exploit. With `similarity_top_k=10` you pay ten round trips per query. It is also drift-prone: weak models sometimes discard good material during a refine step, or ramble as the answer is rewritten repeatedly. ## compact (the default) Compact is "compact and refine". Before calling anything, it concatenates node texts into the largest chunks that still fit the prompt window, producing a small number of packed contexts. It then runs the refine chain over *those*, not over individual nodes. In practice most top-k retrievals fit into one packed prompt, so compact is a single LLM call where refine would have been ten. That is why it is the library default: same templates, same refine semantics on overflow, a fraction of the latency and cost. Its only real weakness is the flip side of packing — several passages share one prompt, so the model's attention is divided across them, and a strong passage can be crowded by neighbours in the same context block. ## tree_summarize A different shape entirely. Nodes are grouped into batches sized to the context window and each batch is summarized against the query using `summary_template`. The resulting summaries are then themselves batched and summarized, and so on until a single text remains — a bottom-up reduction tree. What this buys: every node influences the final answer through its level of the tree, so no passage is simply refined away, and the mode scales to node counts that would make refine absurd. It also parallelizes — sibling batches at a level are independent, and `use_async=True` runs them concurrently, so wall-clock time grows with tree *depth* (logarithmic) rather than node count. What it costs: more total LLM calls than compact for the same nodes, and information loss at each level, because a summary of a summary is lossy. For a pointed factual question ("what is the refund window?") that lossiness is pure downside — the answer was in one node and compact would have quoted it. For "summarize the risks across these forty filings", it is exactly right. ## Choosing - **Pointed question, small top-k, latency matters** → compact. This covers most chat-style RAG. - **Pointed question, and you suspect packing is diluting a key passage** → refine, accepting N calls; or better, keep compact and fix precision with a reranker so fewer, better nodes are packed. - **Question genuinely spans the whole retrieved set — summaries, comparisons, exhaustive extraction** → tree_summarize with `use_async=True`. Other modes exist for narrower jobs: `simple_summarize` merges and truncates everything into a single call (cheapest, silently drops text), `accumulate` and `compact_accumulate` apply the query to each chunk separately and concatenate the results (good for per-document extraction, not for prose), `no_text` runs retrieval and skips generation entirely (useful for evaluating retrieval or inspecting `source_nodes`), and `generation` ignores context and just asks the LLM. ## Cost model in one line Compact ≈ ceil(total context tokens / window) calls, mostly one. Refine = one call per node, sequential. Tree_summarize ≈ one call per batch per level, parallelizable. Multiply by your per-query rate to see why the mode, not the model, is often the dominant term in a RAG bill — and why quietly switching modes to "improve answers" can double latency without anyone noticing.

  • Which mode does as_query_engine() use if you set nothing, and why that one?
    `compact`. It gives the accuracy semantics of refine while collapsing most retrievals into a single LLM call, so the default is cheap and fast without a quality cliff. Refine as a default would multiply latency by the node count for no gain on the common case where everything fits one prompt.
  • Why is tree_summarize the mode that benefits most from use_async=True?
    Its batches within a tree level are independent — each summarizes a different group of nodes — so they can be issued concurrently, and wall-clock time tracks tree depth rather than node count. Refine cannot parallelize at all, because each call consumes the previous call's answer, and compact usually has only one call to make.
  • You switch from compact to tree_summarize and a simple factual answer gets vaguer. Why?
    Tree summarization is lossy by construction: the exact sentence containing the fact is compressed into a batch summary, then compressed again at the next level. Pointed factual questions want the original passage in the prompt, which compact provides. Use tree_summarize only when the question really is about the whole set.
  • How do you change the wording of the prompts these modes use?
    Pass your own templates to the synthesizer: `text_qa_template` for the initial QA prompt, `refine_template` for the refinement step used by refine and compact, and `summary_template` for tree_summarize. They are set on `get_response_synthesizer(...)` or forwarded through the query-engine constructor, so you can enforce citation or refusal behaviour per mode.

saying these in an interview costs you the question

  • Believing refine is cheaper than compact because it looks simpler
  • Thinking compact truncates nodes rather than packing them
  • Using tree_summarize for single-fact lookups
  • Assuming the response mode changes what gets retrieved
  • Expecting refine calls to run in parallel

context