In RAGFlow, how do top_k and top_n differ in a chat assistant's retrieval config?
answer
- one bounds the pool, one bounds the prompt
- 1024 and 6 by default
- threshold runs between the two
- reranker cost scales with the pool
- more chunks in the prompt is not better
basics
~20 stop_k bounds the candidate pool RAGFlow pulls from the indexes and scores, defaulting to 1024. top_n bounds how many surviving chunks are actually injected into the prompt, defaulting to 6. One controls recall and cost, the other context size.
solid answer
~40 sThey sit at two different stages. `top_k` is the retrieval-side pool: how many candidates RAGFlow pulls from the vector search before scoring, blending and — if a rerank model is set — reranking. Its default is 1024, and it is a recall and cost dial, because every candidate in that pool is work the reranker has to do. `top_n` is the generation-side cut: after the similarity threshold has discarded weak chunks, the best `top_n` remaining chunks are rendered into the `{knowledge}` placeholder. Its default is 6, and it is a context-budget dial that drives prompt tokens and latency at the model. Raising `top_n` without raising quality upstream just pushes more mediocre chunks at the model; raising `top_k` with a reranker configured is the fastest way to make a cheap assistant slow.
go deeper
Know that one setting decides how many candidates are considered and the other decides how many chunks end up in the prompt, and remember the defaults of 1024 and 6.
Explain the funnel order — pool, score, threshold, cut — and why the similarity threshold sits between the two limits rather than after both.
Show that you connect the candidate pool to rerank latency and spend, and that you diagnose with the retrieval-testing panel before turning knobs at random.
Own the cost model: retrieved context is paid on every turn of every conversation, so pool depth and prompt budget are capacity decisions with a per-assistant blast radius, not local tuning preferences.
## Two stages, two limits A RAGFlow retrieval has a funnel shape. Candidates come out of the indexes, get a blended keyword-plus-vector score, are optionally rescored by a rerank model, are filtered by the similarity threshold, and the survivors are cut to a final handful that is rendered into the prompt. `top_k` bounds the top of that funnel; `top_n` bounds the bottom. ## top_k — the candidate pool `top_k` defaults to 1024. It governs how deep the vector search goes before RAGFlow does its own scoring work. A larger pool improves the chance that the genuinely correct chunk is present at all, which matters most on large or heterogeneous knowledge bases where the right passage does not rank near the top on either signal alone. The cost is not free, and it is not linear in the way people expect. Without a reranker, a bigger pool mostly costs search and scoring time. With a rerank model configured, every candidate in the pool is a pair the reranker must score, so pool size becomes the dominant latency and spend term of the whole request. This is why RAGFlow's own guidance flags reranking as time-consuming: the slowness is usually the pool, not the model. If you turn on a reranker and keep a large `top_k`, you have signed up for that cost on every turn. ## top_n — the prompt budget `top_n` defaults to 6. After thresholding, this is how many chunks are formatted into the `{knowledge}` block. It is the direct lever on prompt size: chunk size times `top_n` is roughly your retrieved-context token bill per turn, paid on every message of every conversation. It also interacts with the threshold. The threshold runs first and is a floor on quality; `top_n` runs second and is a cap on quantity. If the threshold already leaves three chunks, a `top_n` of 20 changes nothing. If the threshold is permissive, a large `top_n` fills the prompt with marginal material. ## Tuning them against each other The useful mental model is that `top_k` buys recall and `top_n` spends context. Diagnose which one you need by replaying real queries in the retrieval-testing panel: if the chunk you know is correct never appears at any threshold, the pool or the scoring is wrong and `top_k` or the weight is the lever. If it appears but sits below the cut, `top_n` or the threshold is the lever. If it appears at rank one and the answer is still bad, neither knob is your problem — look at the template and the model. A reasonable production shape is a generous `top_k` with no reranker, or a deliberately reduced `top_k` once a reranker is on, paired with a small `top_n` — because past a handful of chunks the marginal one rarely helps and reliably costs tokens. ## Things people get wrong Treating `top_n` as the thing that controls retrieval quality is the classic error: it only decides how many of the already-selected chunks are used. Assuming a bigger `top_n` is safer is the second: more context is more tokens, more latency, and more chance the model leans on a weak passage. And forgetting that these are per-assistant settings means a knowledge base tuned in the testing panel can behave differently in the assistant that uses it, because the assistant carries its own copies of the values.
- You enable a rerank model and p95 latency triples. What do you change first?Reduce the candidate pool. With a reranker configured, every candidate that retrieval pulled has to be scored by the rerank model, so pool size is the dominant latency term. Cut `top_k` to a few hundred, confirm in the retrieval-testing panel that the chunks you care about still appear, and only then look at the reranker itself or at hosting it closer to the service. Leaving a 1024-candidate pool in place and blaming the model is the common mistake.
- Does raising top_n ever hurt answer quality rather than just cost?Yes. Once the genuinely relevant chunks are in the prompt, additional ones are by definition weaker matches, and they compete for the model's attention while inflating the context. On a permissive threshold a large top_n can pull in near-duplicates or off-topic passages that the model then cites. The safer combination is a threshold that removes weak material and a small top_n that keeps the prompt tight.
- If the correct chunk never shows up in the retrieval-testing panel at any threshold, which knob is at fault?Not top_n — that only trims what already survived. Look upstream: the candidate pool may be too shallow for a large knowledge base, or the keyword-versus-vector weighting may be scoring the wrong signal for that query shape. Adjust the pool and the weight, re-test, and if the chunk still never appears, question whether the document was parsed into the chunk you think it was.
saying these in an interview costs you the question
- Says top_n controls how many chunks are retrieved
- Thinks a bigger top_n always improves answers
- Enables a reranker without reducing the candidate pool
- Assumes the threshold runs after the top_n cut
- Believes the two settings are the same value under different names