skip to content

In RAGFlow, what constrains a chat assistant that spans multiple knowledge bases?

level: seniorimportance: should knowfreq 42%

answer

  1. one ranking, so one vector space
  2. embedding model must match across bases
  3. settings belong to the assistant, not the base
  4. no per-corpus quota in the cut
  5. a big corpus can crowd out a small one

basics

~20 s

All knowledge bases attached to one RAGFlow assistant must share an embedding model, since their vectors are compared in a single ranking. Beyond that hard rule, one threshold, one weight and one top-N are applied across corpora whose score distributions differ.

solid answer

~50 s

The hard constraint is the embedding model: RAGFlow ranks chunks from every attached knowledge base in one blended ordering, so the vectors have to live in the same space, and you cannot mix knowledge bases built with different embedding models in one assistant. The soft constraints are the ones that bite in production. The assistant carries a single similarity threshold, keyword weight and top N, applied to every corpus at once — but a large chatty knowledge base and a small precise one produce different score distributions, so one threshold is a compromise and a big corpus can occupy every top-N slot while the small one is never cited. Different chunk templates across the bases make it worse by mixing chunk granularity in the same ranking. When those distributions genuinely diverge, split into separate assistants rather than tuning a threshold that fits neither.

go deeper

for a junior

Know that one assistant can be attached to several knowledge bases and that they must all use the same embedding model for their chunks to be ranked together.

for a middle

Explain that the retrieval settings live on the assistant and are applied to every attached corpus at once, and that top N is a single global cut over the merged ranking.

for a senior

Show the production judgment: diagnose a never-cited corpus by testing it in isolation, and reach for splitting the assistant or aligning chunk templates rather than endlessly tuning one threshold.

for a principal

Own the topology decision — how many assistants, which corpora combine, what routes between them — and treat the embedding model as a standardised, coordinated choice because it determines what can ever be merged.

## The hard rule first Attaching several knowledge bases to one chat assistant means their chunks compete in a single ranked list. That only makes sense if the vector scores are comparable, which means every attached knowledge base must have been built with the same embedding model. RAGFlow enforces this at selection time rather than letting you assemble an incoherent assistant. The practical consequence is that the embedding model is an early, sticky decision: changing it later means re-indexing every knowledge base that must stay combinable, and knowledge bases built on different models can never be merged into one assistant without a rebuild. ## One set of knobs, several corpora The prompt-engine settings — similarity threshold, keyword similarity weight, top N, rerank model — belong to the assistant, not to each knowledge base. They are applied uniformly across everything attached. That uniformity is where multi-knowledge-base assistants go wrong, because the blended score distribution is a property of a corpus. Short structured chunks from a tabular source score differently from long prose paragraphs. A corpus written in the user's vocabulary scores high on the keyword component; a corpus written in internal jargon does not. A threshold calibrated on the first will silently exclude the second, and the assistant will look like it simply does not know about half its material. ## Crowding in the top-N budget There is no per-knowledge-base quota. Top N is a global cut over the merged ranking, so a 50,000-chunk knowledge base with many near-matching passages will routinely fill all the slots and a 200-chunk one may never surface, even when it holds the authoritative answer. The bigger corpus is not better; it just has more tickets in the draw. Symptoms are answers that are always sourced from the same one or two documents, and a knowledge base that is attached but never appears in any citation. ## Mitigations, roughly in order of preference **Split the assistant.** Two assistants with settings tuned to their own corpora, chosen by the caller or by an upstream router, beats one assistant tuned to neither. This is usually the right answer and candidates too rarely say it. **Make the corpora comparable before merging.** Aligning chunk templates and chunk sizes across the attached knowledge bases narrows the score-distribution gap that made one threshold unworkable. **Add a rerank model.** A reranker rescores candidates against the query directly, which makes chunks from different corpora more comparable than raw hybrid scores are — at the cost of reranking the whole candidate pool on every turn. **Tune to the weaker corpus, then raise top N slightly.** A lower threshold plus a modestly larger cut gives the small corpus a chance to appear, at the price of more marginal context. ## Watch the citations Citation coverage is the cheap diagnostic. Run the assistant over a set of questions whose answers you know live in the smaller knowledge base and check whether its chunks ever appear in the references. If they never do, the problem is not the model or the prompt; it is that this corpus loses the ranking on every query, and no amount of prompt engineering fixes a chunk that was never injected.

  • Why does one similarity threshold rarely fit two different corpora?
    Because the blended score is a property of the corpus, not an absolute quality measure. Chunk length, chunk template, and how closely the documents' vocabulary matches the users' all shift the score distribution. A threshold tuned where users echo the documents' wording will sit above most scores from a corpus written in internal jargon, so that corpus is silently excluded even though its chunks are relevant.
  • Your smaller knowledge base is never cited even though it holds the authoritative answers. What do you do?
    First confirm it in the retrieval-testing panel by querying that knowledge base alone — if the chunks score well in isolation, the problem is that they lose the merged ranking. Then either split it into its own assistant with its own settings, align its chunk template with the larger corpus so scores are comparable, or add a reranker so relevance rather than raw hybrid score decides the order. Prompt changes cannot rescue an uninjected chunk.
  • What does the embedding-model constraint imply for how you plan knowledge bases up front?
    That the embedding model is an architectural decision with a wide blast radius, not a per-knowledge-base preference. Any two corpora that might ever be served by one assistant must be built on the same model, and switching models later means re-indexing all of them together. Standardise the model per environment early and treat a change as a coordinated re-index, not an incremental one.

saying these in an interview costs you the question

  • Thinks knowledge bases with different embedding models can be mixed
  • Assumes each attached knowledge base gets its own share of the top-N slots
  • Believes similarity scores are comparable across corpora by default
  • Tunes the threshold on the largest corpus and calls it done
  • Says a better prompt can surface a chunk that was never retrieved

context