skip to content

A RAGFlow assistant answers from model knowledge, not your documents — how do you debug it?

level: seniorimportance: must knowfreq 58%

answer

  1. bisect before tuning
  2. replay in the retrieval-testing panel
  3. use the assistant's own threshold and weight
  4. empty result versus good chunks, bad answer
  5. blank empty response hides the failure

basics

~20 s

Bisect retrieval from generation. Replay the question in the knowledge base's retrieval-testing panel using the assistant's exact threshold and weight: if good chunks come back there, the fault is prompt-side; if nothing comes back, it is retrieval-side.

solid answer

~50 s

The panel is the bisection tool. Run the failing question against the knowledge base with the assistant's own similarity threshold, keyword weight and rerank setting, and read the per-chunk scores. **Nothing returned** points at retrieval: threshold too high for this corpus, the wrong knowledge base attached, documents that were never fully parsed, or a query whose vocabulary the current keyword-versus-vector weighting scores badly. **Good chunks returned but a bad answer** points at the assistant: a system prompt that lost its `{knowledge}` placeholder, a top-N cut too small to include the passage, or a template that never forbids outside knowledge. Then close the observability gap — an assistant with a blank empty response calls the model even when retrieval found nothing, so the failure looks like hallucination instead of an empty result. Setting an explicit empty response makes that state visible.

go deeper

for a junior

Know that the knowledge base has a retrieval-testing panel and that running the failing question there tells you whether any chunks were found at all before you touch the prompt.

for a middle

Be able to list the two branches and their causes: empty result means threshold, parsing or weighting; good chunks with a bad answer means the placeholder, the top-N cut, or a permissive template.

for a senior

Demonstrate the bisection discipline and the observability fix — copy the assistant's settings into the panel, and configure an empty response so an ungrounded turn is a visible event rather than fluent prose.

for a principal

Own the regression story: a fixed set of questions with known correct chunks, replayed after every re-parse, threshold change and reranker rollout, because retrieval regressions are otherwise discovered by users.

## Why this is the archetypal RAGFlow question The symptom — fluent, confident, unsourced answers — has at least six distinct causes spread across two subsystems. Candidates who start turning knobs lose; candidates who bisect win. RAGFlow gives you the bisection instrument for free in the knowledge base's retrieval-testing panel, which runs the same retrieval path an assistant uses and shows what came back and why. ## Step one: reproduce in the panel with the assistant's settings The panel has its own threshold, keyword weight and rerank controls, so a test at defaults proves nothing about an assistant configured differently. Copy the assistant's prompt-engine values across, run the failing question verbatim, and look at the returned chunks, the source document for each, and the score breakdown. ## Branch A: the panel returns nothing The fault is upstream of the prompt. - **Threshold too high for this corpus.** Blended scores are corpus-dependent; a value that worked on one knowledge base can empty another. Lower it and watch what appears. - **Wrong or unparsed documents.** A document that was uploaded but never finished parsing has no chunks to retrieve. Confirm the file is in a parsed state and that its chunks exist. - **Wrong knowledge base attached.** An assistant can point at a different knowledge base than the one you tested, which is the single most embarrassing cause. - **Weighting mismatched to the query shape.** Users typing exact identifiers on a vector-heavy configuration, or paraphrasing on a keyword-heavy one, both produce plausible-looking scores that never clear the floor. Try two or three weights. ## Branch B: the panel returns exactly the chunk you wanted The fault is between retrieval and the model. - **The `{knowledge}` placeholder is missing from the system prompt.** Retrieval runs, chunks are found, nothing is injected, and the model answers unaided. Nothing errors. Check this first — it is fast and it is common. - **Top N is too small.** The passage is at rank eight and the assistant keeps six. - **The template permits outside knowledge.** If the prompt never says to answer only from the supplied material, a model with strong priors will happily blend them. - **Model settings.** A very high sampling temperature makes a weakly-grounded answer drift further from the retrieved text. ## Make the failure observable The deeper problem is that the ungrounded path is silent. Two settings fix that. The **empty response** turns "retrieval found nothing" into a fixed, recognisable reply instead of a model call — with it blank, the model is invoked anyway and produces prose that reads exactly like a normal answer. **Show quote** surfaces the references, so an answer with no citations attached becomes visually distinct from a grounded one. Together they convert a silent quality bug into something support and QA can spot without reading logs. ## After you fix it Add the failing question to a fixed test set and re-run it in the panel after every change to the threshold, weight, top N, rerank model, or chunk template. Retrieval regressions are invisible until someone asks the right question, so the only durable defence is a small set of questions whose correct chunks you know, replayed on every configuration change and after every re-parse.

  • The panel returns the right chunk but the assistant still ignores it. What are your two prime suspects?
    First, a system prompt that no longer contains the `{knowledge}` placeholder — retrieval runs, chunks are found, and nothing is injected, with no error anywhere. Second, a top-N cut smaller than the rank at which the chunk sits, so it is retrieved but discarded before the prompt. Check the template, then compare the chunk's rank in the panel against the assistant's top N.
  • How does configuring an empty response change the failure mode rather than just the wording?
    It changes whether the model is called at all. With text configured, a retrieval that returns nothing short-circuits into that fixed reply, so the failure is loud and countable. Left blank, RAGFlow calls the model anyway with an empty knowledge block and you get a fluent unsourced answer that is indistinguishable from a good one at a glance. It converts a silent quality bug into an observable event.
  • Why can this bug appear weeks after the assistant was working fine?
    Because the inputs drift. Someone re-parses a knowledge base with a different chunk template and the chunk boundaries move; someone raises the threshold to cut noise for a different query class; a new batch of documents shifts the score distribution; or a reranker is enabled and the previously tuned threshold no longer fits the new score scale. Any of those can quietly empty the result set for a query family that used to work.

saying these in an interview costs you the question

  • Starts by rewriting the prompt without checking what retrieval returned
  • Tests in the panel at default settings rather than the assistant's
  • Blames the embedding model before confirming documents were parsed
  • Assumes a fluent answer with no citations is still grounded
  • Leaves the empty response blank and calls the result hallucination

context