How would you turn RAGFlow's retrieval-testing panel into a real evaluation loop?
answer
- a spot check is not an evaluation
- fixed question set with known-good chunks
- replay after every parse or setting change
- sweep one variable at a time
- record rank and margin above the threshold
basics
~20 sReplace ad-hoc queries with a fixed question set whose correct chunks are known, replay it after every configuration or parsing change, change one variable at a time, and record the settings alongside the results so a regression is attributable.
solid answer
~50 sThe panel gives you a single-query view with per-chunk scores; an evaluation loop is that view made repeatable and attributable. Build a question set drawn from real user traffic, and for each question record which chunk or document is the right answer. Replay the whole set whenever anything upstream changes — threshold, keyword weight, candidate pool, rerank model, chunk template, embedding model, or a document re-parse — and record for each question whether the known-good chunk came back and at what rank. Sweep one variable at a time, because moving the threshold and the weight together tells you nothing about either. Keep the configuration snapshot with the results so a later regression can be traced to a change. Then measure the retrieval stage separately from answer quality: an assistant can answer well on chunks that barely made the cut, and that is fragility, not success.
go deeper
Know that the retrieval-testing panel lets you run a query against a knowledge base with explicit settings and see which chunks came back and with what scores.
Be able to describe a repeatable loop: a fixed set of questions with known-good chunks, replayed after configuration changes, with one variable moved at a time.
Show that you separate retrieval grading from answer grading, track rank and margin above the threshold rather than pass or fail, and re-run after every re-parse.
Own the process: who owns the question set, what changes require a replay before rollout, how configuration snapshots are recorded for attribution, and when the loop graduates from a human panel to an automated run.
## What the panel gives you, and what it does not The knowledge base's retrieval-testing panel runs a query against that knowledge base with an explicit similarity threshold, keyword weight and optional rerank model, and shows the chunks that came back with their scores and source documents. That is a genuinely good instrument: it isolates retrieval from generation and it exposes the numbers that drive the ranking. What it does not give you is memory. Each run is a spot check, so a team using it casually learns whether one query works today and nothing about whether last week's queries still work. The gap between the two is the entire evaluation story an interviewer is listening for. ## Build the question set from traffic, not imagination Harvest real questions: the ones users actually asked, especially the ones that failed. For each, record the ground truth as the document and passage that answers it. Include deliberate negatives — questions the corpus genuinely cannot answer — because a configuration that answers those is worse than one that returns nothing. Keep the set small enough to replay by hand at first; fifty well-chosen questions beats a thousand nobody runs. ## Replay on the changes that actually move retrieval The triggers are not only the obvious knobs. Threshold, keyword weight, candidate pool and rerank model change ranking directly. A chunk-template change or a document re-parse changes the chunks themselves, so previously good results can vanish while every setting looks untouched. An embedding-model change invalidates everything. Ingesting a large new batch shifts the score distribution and can push a previously safe chunk under the threshold. Every one of those deserves a replay. ## Sweep one variable at a time The discipline that separates evaluation from fiddling. Hold everything fixed, move the threshold across a few values, and record the outcome per question. Then reset and do the weight. Then decide on the reranker as a separate question, because it changes both the ranking and the score scale, so any threshold you had tuned must be re-tuned after it goes on. Record the whole configuration with each result set, or you will have numbers you cannot attribute. ## What to record per question Whether the known-good chunk was returned at all; at what rank; at what score; and how close that score sat to the threshold. That last number is the one teams skip and the one that predicts future breakage — a chunk clearing the floor by a hair will fall below it on the next re-parse. Reading the term-similarity and vector-similarity components separately also tells you which signal is carrying the query, which is what makes weight tuning evidence-based rather than superstitious. ## Keep retrieval and answer quality separate A good answer produced from a chunk that barely survived is a fragile pass, and a bad answer from a perfectly retrieved chunk is a prompt or model problem that no retrieval tuning will fix. Evaluating the two stages together hides both. Grade retrieval first on the panel's evidence, then grade answers only on questions whose retrieval already passed. ## Scaling past the panel At some point clicking stops being viable, and the loop moves to driving the same retrieval path programmatically with the same parameters, so the set can run in CI or on a schedule. The panel stays as the human debugging surface for the individual failures the automated run flags. The governance layer is the last piece: configuration changes reviewed rather than applied live, a replay required before any knowledge base is re-parsed in production, and an owner for the question set — otherwise it rots and everyone quietly stops trusting it.
- Why record how far a chunk's score sat above the threshold rather than just pass or fail?Because margin predicts breakage. A chunk clearing the floor by a hair passes today and disappears after a re-parse, a new document batch, or any change that nudges the score distribution. Tracking the margin turns a binary pass into a fragility signal, and it tells you which questions to watch when the corpus grows. It also shows whether a threshold change bought real headroom or just moved the boundary.
- Why replay the whole question set after re-parsing a knowledge base with a new chunk template?Because re-parsing regenerates the chunks. Boundaries move, chunk lengths change, and the term and vector similarities are recomputed over different text, so scores shift even though every retrieval setting is untouched. Questions that passed can silently fall under the threshold and a previously irrelevant passage can start winning. The settings look identical, which is exactly why the regression goes unnoticed without a replay.
- How do you decide whether a rerank model is worth enabling, using this loop?Run the set with the reranker off and on, holding everything else fixed, and compare rank and margin for the known-good chunks, not just pass rate. Then measure the added latency at the candidate-pool size you would actually run. If reranking mainly reorders chunks that were already inside the cut, you are buying little for real per-turn cost. Retune the threshold afterwards, since the score scale has changed.
- What makes a question set rot, and how do you prevent it?Drift and neglect: the corpus changes, the product's vocabulary changes, users ask new things, and the recorded ground truth points at documents that no longer exist. Prevent it by refreshing from live traffic on a cadence, retiring questions whose source documents were removed, keeping negatives current, and giving the set a named owner. An unowned suite is trusted for a while and then quietly ignored.
saying these in an interview costs you the question
- Calls one-off panel queries an evaluation
- Changes the threshold and the weight in the same run
- Grades answer quality without checking what was retrieved
- Builds the question set from imagination rather than real traffic
- Assumes results stay valid after re-parsing a knowledge base