skip to content

In RAGFlow, what do the similarity threshold and keyword similarity weight control?

level: middleimportance: must knowfreq 75%

answer

  1. two knobs, one blended score
  2. term overlap plus vector cosine
  3. the weights are complements summing to one
  4. threshold is a floor, not a count
  5. defaults 0.2 and 0.7

basics

~20 s

RAGFlow scores each candidate chunk as a weighted blend of keyword term similarity and vector cosine similarity. The keyword similarity weight sets that blend; the similarity threshold is the minimum blended score a chunk must reach to be kept.

solid answer

~50 s

RAGFlow retrieval is hybrid by default. Each candidate chunk gets a term-overlap score from the full-text index and a cosine score from the embedding index, and they are combined as `(1 - vector_similarity_weight) * term_similarity + vector_similarity_weight * vector_similarity`. The UI exposes the keyword end of that dial as **Keywords similarity weight** (default 0.7); the API parameter is its complement, `vector_similarity_weight` (default 0.3). They are one dial from two ends, not two independent settings. The **similarity threshold** (default 0.2) is then a floor on the blended score: anything below it is dropped before the top-N cut, so it governs recall, not ranking. Raising the threshold never improves answer quality on its own — it only removes chunks, and set high enough it returns nothing and the assistant answers unaided or falls back to its empty response.

code

python · 6 lines
python
def ragflow_hybrid_score(term_similarity, vector_similarity, vector_similarity_weight=0.3):
    return (1 - vector_similarity_weight) * term_similarity + vector_similarity_weight * vector_similarity

# weak literal overlap, strong semantic match
print(round(ragflow_hybrid_score(0.10, 0.62), 3))  # 0.256 -> clears a 0.2 threshold
print(round(ragflow_hybrid_score(0.10, 0.62, 0.9), 3))  # 0.568 -> vector-heavy weighting

go deeper

for a junior

Know that RAGFlow retrieval mixes keyword matching with vector similarity, and that the threshold is the cutoff below which a chunk is discarded. Being able to point at both controls in the retrieval settings is enough here.

for a middle

Be ready to write the blend formula, state that the UI keyword weight and the API vector weight are complements, and explain that the threshold filters before the top-N cut rather than limiting how many chunks are used.

for a senior

Show that you tune these per knowledge base against a fixed query set rather than guessing, and that you recognise a too-high threshold as a common cause of an assistant silently answering without documents.

for a principal

Own the calibration policy: who sets thresholds, how they are re-validated after a re-parse or an embedding-model change, and why enabling a reranker invalidates previously tuned values across every assistant that touches the knowledge base.

## Two indexes, one score A RAGFlow knowledge base is indexed twice: into a full-text index for term matching and into a vector index for embedding similarity. Every retrieval path — the knowledge base's retrieval-testing panel, a chat assistant, and the retrieval API — runs both and merges the results. Two settings govern that merge and they appear side by side in all three places. **Keywords similarity weight** decides how much of the final score comes from literal term overlap versus semantic closeness. **Similarity threshold** decides how good the blended score has to be for the chunk to survive at all. ## The blend The combined score is a straight linear mix: `score = (1 - vector_similarity_weight) * term_similarity + vector_similarity_weight * vector_similarity` With the defaults, `vector_similarity_weight` is 0.3, so 70% of the score is term similarity and 30% is vector cosine. The term component is not raw word counting: it is a weighted match over the analysed tokens of the query against the full-text index, so rare and content-bearing words carry more of it than stopwords. ## The naming trap The web UI labels the control **Keywords similarity weight** and defaults it to 0.7. The retrieval API and SDK expose `vector_similarity_weight` and default it to 0.3. Candidates who have only used one surface often assume there are two knobs that could disagree. There is one knob; the two numbers always sum to 1. Setting the API value to 0.9 is the same act as dragging the UI slider down to 0.1. ## Choosing the weight Keyword-heavy settings win where the corpus is full of exact tokens the user will type: part numbers, error codes, statute references, API names, drug names. Vector-heavy settings win where users paraphrase and the documents use different words than the question — policy prose, support articles, meeting notes. Most production knowledge bases end up somewhere in between, and the honest way to pick is to run a fixed set of real questions through the retrieval-testing panel at two or three weights and look at whether the chunk you know is correct actually comes back and at what rank. ## Choosing the threshold The threshold is a floor, not a count. It exists to suppress the long tail of weakly-matching chunks so the model is not handed noise, and to make "nothing relevant here" an observable outcome rather than six bad chunks. Because scores are a blend of two heterogeneous similarities over one particular corpus, a given threshold value is not portable: 0.2 may be permissive in one knowledge base and brutal in another with different chunk sizes or a different embedding model. Tune it per knowledge base against real queries, and treat "returns nothing" as the signal that you have gone too far. ## What a reranker changes When a rerank model is configured, RAGFlow keeps the same weighted structure but the reranker's relevance score takes the place of the vector-similarity term. The keyword weight still applies, and the threshold still applies to the resulting blend — which means switching a reranker on shifts the score distribution and usually invalidates the threshold you had tuned without one. Retune after enabling it. ## Common failure shapes An assistant that suddenly answers from general model knowledge is often a threshold that is too high for the corpus, not a prompt problem. An assistant that cites plausible-looking but off-topic chunks is often a weight problem: heavy keyword weighting on a paraphrasing user base, or heavy vector weighting on a corpus whose value is exact identifiers. Both are diagnosed the same way — replay the query in the retrieval-testing panel with the assistant's exact threshold and weight and read the per-chunk scores.

  • You raise the similarity threshold to 0.5 and the assistant starts saying it does not know. Is that a bug?
    No — it is the threshold doing exactly what it is for. At 0.5 few chunks in a typical knowledge base clear the blended floor, so retrieval returns an empty set and the assistant either emits its configured empty response or answers unaided. Confirm by replaying the query in the retrieval-testing panel at 0.5 and at 0.2 and comparing what comes back. Tune the threshold per knowledge base against real queries rather than picking a round number.
  • Why can the same threshold behave completely differently on two knowledge bases?
    Because the blended score depends on the corpus. Chunk size changes term-similarity density, the embedding model changes the cosine distribution, and the chunk template changes what a chunk even contains. A short-chunk knowledge base parsed with a table-oriented template produces a different score spread than long prose chunks. Thresholds are therefore per-knowledge-base calibration, not a global constant to copy between projects.
  • If you configure a rerank model, does the keyword similarity weight stop mattering?
    No. RAGFlow keeps the weighted blend and substitutes the reranker's score for the vector-similarity term, so the keyword component still contributes at whatever weight you set. What does change is the score scale: reranker outputs distribute differently from cosine similarity, so a threshold tuned without a reranker is usually wrong once one is enabled, and needs retuning.

saying these in an interview costs you the question

  • Says the threshold filters on vector cosine only
  • Assumes keyword weight and vector weight are set independently
  • Thinks raising the threshold makes answers more accurate
  • Treats 0.2 as a universal value that ports across knowledge bases
  • Believes the threshold caps how many chunks reach the prompt

context