What does similarity_top_k control in a LlamaIndex retriever, and how do you tune it?
answer
- how many neighbours come back
- the default is smaller than you think
- cheap in the store, costly in the prompt
- recall knob upstream, precision knob downstream
- wide net, narrow mouth
basics
~20 ssimilarity_top_k is how many nodes a LlamaIndex retriever asks the vector store for, defaulting to 2 on VectorIndexRetriever. Raising it buys recall at the cost of prompt tokens, latency and distraction, so the usual pattern is a wide top_k narrowed by a reranker.
solid answer
~50 s`similarity_top_k` is the number of nodes the retriever pulls back for a query — `index.as_retriever(similarity_top_k=10)` or the same kwarg on `as_query_engine`. In llama-index-core 0.14.x the default on `VectorIndexRetriever` is 2, which is deliberately small and almost always too small for real corpora, especially after aggressive chunking. Raising it is the cheapest recall fix available: retrieval cost is one vector-store query either way, and the marginal cost is embedding-store bandwidth. The real cost lands downstream — every extra node becomes prompt tokens, and with a compacting synthesizer enough nodes force extra LLM calls, so latency and spend scale with k, not with the query. It also hurts precision: irrelevant nodes crowd the context and the answer drifts. The production shape is therefore retrieve wide, then cut narrow: `similarity_top_k=20–50` into a reranker with `top_n=3–5`, so only the reranked survivors reach the LLM.
code
python · 6 linesretriever = index.as_retriever(similarity_top_k=20)
for k in (2, 5, 10, 20):
r = index.as_retriever(similarity_top_k=k)
hits = r.retrieve("what is the refund window?")
print(k, [round(n.score, 3) for n in hits])go deeper
Know that similarity_top_k is simply how many chunks come back from the search, that it is set on as_retriever or as_query_engine, and that the library default is very small.
Explain where the cost lands: cheap in the vector store, expensive in prompt tokens and extra synthesizer calls. Be able to say why an answer can get worse when k goes up.
Show the retrieve-wide-rerank-narrow pattern and how you would pick k from a recall@k curve on labelled questions rather than by feel, holding one knob fixed while tuning the other.
Frame k as a spend-per-query decision tied to a recall target and a latency budget, including how filtered queries, chunk size and fusion child retrievers change the number, and what regression signal tells you it needs revisiting.
## What the knob actually does `similarity_top_k` is passed straight through to the vector store as the number of nearest neighbours to return for the query embedding. `index.as_retriever(similarity_top_k=10)` builds a `VectorIndexRetriever` with that value; `index.as_query_engine(similarity_top_k=10)` routes the same kwarg to the retriever half of the engine. The retriever returns exactly that many `NodeWithScore` objects (fewer only if the store holds fewer matching nodes, or if metadata filters cut the candidate set). In llama-index-core 0.14.x the default is **2**. That default exists so demos are cheap, not because two chunks is a sensible production setting. With a 512-token chunk size, two nodes is roughly a page of text — enough for a toy corpus, badly insufficient when the answer is spread across several sections or when near-duplicate boilerplate outranks the real passage. ## Why raising it is cheap on the retrieval side An approximate-nearest-neighbour search returning 50 results is not meaningfully more expensive than one returning 2: the same graph or list traversal happens, and the extra work is bookkeeping over the candidate heap plus the bytes to ship the node text back. Going from k=2 to k=20 typically moves retrieval latency by single-digit milliseconds. So if the correct passage is missing from your results, the first experiment is always to raise k and check whether it appears at rank 12 rather than rank 2. If it does, you have a *ranking* problem, not a *coverage* problem, and reranking or hybrid retrieval is the fix. If it never appears at any k, the problem is upstream — embeddings, chunking, or the document never made it into the index. ## Why raising it is expensive on the synthesis side Everything the retriever returns flows into the response synthesizer, and there the costs are real: - **Tokens.** k nodes times chunk size is the prompt budget. Twenty 512-token nodes is ~10k tokens of context per query, paid on every request. - **LLM calls.** Compacting and refining synthesizers batch nodes into as many prompts as the context window requires. More nodes means more batches means more sequential calls, so latency grows in steps, not smoothly. - **Answer quality.** This is the counter-intuitive one. Models attend unevenly across long contexts, and irrelevant passages act as distractors: an answer that was correct at k=5 can become hedged, or can pick up a claim from a plausible-but-wrong chunk, at k=30. Recall and precision trade against each other, and the LLM only sees precision. ## The retrieve-wide-then-cut-narrow pattern Because the two costs sit on opposite sides of the pipeline, the standard production shape separates them: 1. Set `similarity_top_k` high — 20 to 50 — so the right passage is very likely somewhere in the candidate set. This is a recall setting. 2. Put a reranker in `node_postprocessors` with a small `top_n` — 3 to 5. This is a precision setting, and it is what actually reaches the LLM. The LLM then sees a short, high-quality context while the search still had a wide net. Note that `top_n` on a reranker replaces `similarity_top_k` as the number that determines prompt size; `similarity_top_k` becomes purely a candidate-pool setting. ## Tuning it honestly Tune k against a labelled question set, not by feel. For each question record whether the gold passage appears in `retriever.retrieve(q)` at each k — that gives recall@k as a curve. It flattens: somewhere between 10 and 50 the curve stops climbing, and spending beyond the flattening point buys nothing but tokens. Pick k just past the knee, then tune the reranker's `top_n` against answer quality with k held fixed. Also watch the interactions. Small chunks need larger k because the answer is spread over more of them. Metadata filters shrink the candidate pool, so a k that worked unfiltered may under-fill a filtered query. And in a fusion retriever each child retriever carries its own `similarity_top_k` for the candidates it contributes, while the fusion retriever's own `similarity_top_k` decides how many survive the merge — setting only the outer one starves the fusion of candidates. ## What not to do Do not treat a bigger k as a free quality upgrade, and do not tune it in production by eyeballing a handful of queries. Above all, do not raise k to compensate for bad chunking or a weak embedding model; it masks the symptom for the easy questions and inflates cost on every query, including the ones that were already fine.
- You raise similarity_top_k from 5 to 30 and answers get worse. What happened?Precision fell. The extra nodes are distractors: the model now attends over a much longer context in which the relevant passage is a smaller fraction, and plausible-but-wrong chunks can be picked up as support. The fix is not to go back to 5 but to keep the wide candidate pool and add a reranker with a small top_n so only the best few reach the LLM.
- How does similarity_top_k interact with a reranker's top_n?`similarity_top_k` sizes the candidate pool the reranker scores; `top_n` sizes what reaches the synthesizer. Prompt cost tracks `top_n`, reranker cost tracks `similarity_top_k` (a cross-encoder scores every candidate), and recall is capped by `similarity_top_k` — a reranker can only reorder what retrieval already found.
- How would you pick a value with evidence rather than intuition?Build a set of questions with known gold passages, then compute recall@k over `retriever.retrieve(q)` for k in a sweep. Plot it; the curve flattens. Choose k just past the knee, then hold it fixed and tune the reranker's top_n against end-to-end answer quality so the two knobs are never moved together.
saying these in an interview costs you the question
- Leaving the default of 2 in production because it "worked in the demo"
- Assuming a higher top_k always improves answer quality
- Thinking a bigger top_k costs mainly vector-store time
- Using top_k to paper over bad chunking or weak embeddings
- Setting only the fusion retriever's top_k and starving its child retrievers