In LlamaIndex, what is the difference between index.as_retriever() and index.as_query_engine()?
answer
- one finds, the other also writes
- no LLM call on one side
- scored nodes versus a Response
- source_nodes is the bridge
basics
~20 sIn LlamaIndex, as_retriever() returns an object whose retrieve() call gives back scored nodes and makes no LLM call. as_query_engine() wraps a retriever with a response synthesizer, so query() also calls the LLM and returns an answer plus its source nodes.
solid answer
~40 s`index.as_retriever()` gives you a retriever — calling `retriever.retrieve("...")` returns a list of `NodeWithScore` objects, the raw retrieved chunks with their similarity scores. No LLM is involved, so it is cheap and it is the right handle for testing retrieval quality in isolation. `index.as_query_engine()` builds a `RetrieverQueryEngine`: a retriever, an optional list of `node_postprocessors`, and a response synthesizer. `engine.query("...")` runs retrieval, applies the postprocessors, then sends the surviving node text to the LLM and returns a `Response` whose `str()` is the answer and whose `.source_nodes` are the chunks that produced it. Convenience kwargs on `as_query_engine` are split between the two halves — `similarity_top_k` reaches the retriever, `response_mode` and `streaming` reach the synthesizer.
code
python · 13 linesfrom llama_index.core import Document, VectorStoreIndex
index = VectorStoreIndex.from_documents(
[Document(text="Paris is the capital of France.")]
)
retriever = index.as_retriever(similarity_top_k=3)
nodes = retriever.retrieve("What is the capital of France?")
print([(n.score, n.text[:40]) for n in nodes])
engine = index.as_query_engine(similarity_top_k=3)
response = engine.query("What is the capital of France?")
print(str(response), len(response.source_nodes))go deeper
Be able to say plainly that a retriever returns scored chunks with no LLM call, while a query engine retrieves and then asks the LLM to write the answer. Know that the answer object also carries source_nodes.
Explain the three parts a query engine assembles — retriever, node postprocessors, response synthesizer — and which constructor keyword configures which part. Show that you would test retrieval alone before touching prompts.
Demonstrate the debugging fork: measure retrieve() separately from query() to attribute latency, and check whether a bad answer is a missing passage or a mis-synthesized one. Mention using source_nodes for citations and audit trails.
Own the boundary as a design decision: one retrieval configuration shared by several synthesis profiles, evaluated independently so retrieval regressions are caught without generation cost, and cited nodes carried through to the product surface.
## The two halves of a RAG query Answering a question over your own data is two separate jobs. First, *find* the relevant chunks of text. Second, *write* an answer using them. LlamaIndex keeps these as two distinct objects, and knowing which one you are holding is the difference between debugging retrieval and debugging generation. ## as_retriever(): find only `index.as_retriever()` returns a `BaseRetriever`. For a `VectorStoreIndex` the concrete class is `VectorIndexRetriever`. Its only real method is `retrieve(query)` (and `aretrieve` for async), which returns a `list[NodeWithScore]`. A `Node` is a chunk of text plus metadata; a `NodeWithScore` wraps it with a `score` float. For vector retrieval that score is the embedding similarity between the query embedding and the node embedding, as reported by the vector store. Nothing is generated: no prompt is built, no tokens are billed to an LLM, and latency is one embedding call plus one vector-store round trip. This is the object you want when you are asking "is the right passage even coming back?" Print the scores, print `node.text[:200]`, print `node.metadata`. If the answer text is not in that list, no amount of prompt tuning downstream will save you. ## as_query_engine(): find and answer `index.as_query_engine()` returns a `RetrieverQueryEngine`, which is a small assembly of three parts: 1. a retriever (by default the same one `as_retriever()` would give you), 2. `node_postprocessors` — an ordered list of objects that can filter, reorder or rescore the retrieved nodes (rerankers and score cutoffs live here), 3. a response synthesizer — the component that turns surviving node text plus the query into a prompt (or several prompts) and calls the LLM. `engine.query("...")` runs those in order and returns a `Response`. `str(response)` or `response.response` is the answer text. `response.source_nodes` is the post-postprocessing node list that was actually shown to the LLM — this is your citation and attribution surface, and it is the first thing to inspect when an answer looks invented. ## How the keyword arguments split `as_query_engine()` is a convenience constructor, and its kwargs are routed to whichever half owns them. `similarity_top_k` configures the retriever. `response_mode` (`"compact"`, `"refine"`, `"tree_summarize"`, …), `streaming=True` and `llm` configure the synthesizer. `node_postprocessors` is a parameter of the engine itself. When you outgrow the convenience form, build the parts explicitly and pass them to `RetrieverQueryEngine.from_args(retriever, ...)` or to the `RetrieverQueryEngine(...)` constructor — the shortcut and the explicit form produce the same object. ## Why the split matters operationally Because the retriever is usable on its own, you can measure the two stages independently. Time `retriever.retrieve(q)` and compare it with the full `engine.query(q)` to see whether latency lives in the vector store or in synthesis. Run retrieval over a set of labelled questions to compute recall without paying for generation. And when an answer is wrong, the fork is mechanical: if the correct passage is absent from `retrieve()`, fix retrieval (embeddings, chunking, top-k, hybrid, reranking); if it is present but the answer ignores or contradicts it, fix synthesis (response mode, prompt templates, model). The same split also lets one retriever feed several engines with different synthesis behaviour — a fast, cheap engine for chat and a slower summarizing engine for reports — without duplicating the retrieval configuration. ## The common beginner mistake Calling `as_query_engine()` inside a loop over many questions and being surprised by the bill: every call is at least one LLM request, often several depending on the response mode. `as_retriever()` in the same loop costs only embeddings. Reach for the query engine when you want prose; reach for the retriever when you want evidence.
- What exactly is in the list that retrieve() returns?A `list[NodeWithScore]`. Each entry wraps a `Node` — a chunk of text plus its metadata and relationships — with a `score` float from the vector store. Reading `node.text`, `node.metadata` and `score` is the standard way to sanity-check retrieval before any LLM is involved.
- If you already built a retriever, how do you get a query engine from it?Pass it to `RetrieverQueryEngine.from_args(retriever, ...)` or the `RetrieverQueryEngine(retriever=..., response_synthesizer=..., node_postprocessors=[...])` constructor. That is what `as_query_engine()` does for you; building it explicitly is what you do once you need a custom synthesizer or postprocessor chain.
- Where do the source nodes on a Response come from — before or after postprocessing?After. `response.source_nodes` is the list that survived `node_postprocessors`, which is exactly the text the LLM saw. That makes it the right thing to render as citations, and the right thing to check when an answer contains a claim you cannot find in the corpus.
saying these in an interview costs you the question
- Thinking as_retriever() calls the LLM to write an answer
- Assuming query() returns plain text with no node information
- Believing similarity_top_k affects synthesis rather than retrieval
- Debugging a wrong answer by editing prompts without inspecting retrieved nodes
- Calling a query engine when only the evidence chunks were needed