skip to content

Where does a Ragas-generated synthetic testset mislead you about quality?

level: principalimportance: should knowfreq 30%

answer

  1. Answerable by construction
  2. No unanswerable questions in it
  3. Questions inherit corpus vocabulary
  4. Ground truth was never verified
  5. Good for deltas, bad for absolutes

basics

~20 s

Every question is written from chunks already in the corpus, so retrieval is being tested on queries constructed to be answerable — no unanswerable questions, no vocabulary your corpus lacks, and reference answers written by the same model family that will be judged.

solid answer

~50 s

Three structural biases. **Answerability by construction**: each sample is generated *from* nodes in the graph, so the retrievable evidence provably exists. Real traffic contains questions your corpus simply cannot answer, and a synthetic set contains none of them — so it cannot tell you whether the system refuses gracefully. **Lexical closeness**: questions inherit the corpus' vocabulary, which makes embedding retrieval look better than it will on user phrasing; personas mitigate this but do not remove it. **Generated ground truth**: the `reference` answer was written by an LLM from those same contexts, so it is a plausible-looking answer, not a verified one, and it may encode the generator's own errors. The practical consequence is that a synthetic set is an excellent *relative* instrument — it detects regressions and ranks candidate pipelines — and a poor *absolute* one. Treat generated scores as a baseline to curate and compare against, never as a published accuracy figure.

go deeper

for a junior

Know that questions generated from your own documents are always answerable from those documents, so they cannot tell you how the system behaves when it has no relevant content.

for a middle

Explain the three structural biases — answerability by construction, corpus vocabulary in the questions, unverified generated reference answers — and why they push retrieval scores optimistic.

for a senior

Describe the curation workflow you would actually run: read every row, fix or delete bad references, add hand-written unanswerable and adversarial cases, freeze and version the artefact, and use it for deltas rather than absolute claims.

for a principal

Own the policy question of what a synthetic suite may and may not authorise — comparison and regression yes, published accuracy no — and plan the migration from generated rows toward real logged queries as traffic accumulates.

## The honest framing A generated testset solves a real problem — the cold start where you have a corpus, a pipeline, and no way to measure anything. It is genuinely better than shipping blind. But its construction imposes biases that are invisible unless you name them, and a lead is expected to name them before the number gets quoted to anyone. ## Bias one: answerability by construction Ragas writes each question *from* nodes in the knowledge graph. The evidence for every question therefore exists in the corpus, is retrievable in principle, and is recorded in `reference_contexts` as the ground truth. Nothing in the set is unanswerable. Real traffic is not like this. A meaningful fraction of production queries concern things the corpus does not cover, is out of date about, or covers only partially. Handling those well — retrieving nothing and declining, rather than retrieving the nearest irrelevant chunk and confabulating over it — is one of the most important behaviours a RAG system has, and a purely synthetic set gives you exactly zero signal on it. If your evaluation consists only of generated questions, you have not tested the failure mode most likely to embarrass you. The fix is to author a small set of deliberately unanswerable questions by hand and mix them in. It is a handful of rows and it covers a hole the generator structurally cannot fill. ## Bias two: lexical closeness The generator reads a chunk and writes a question about it, so the question tends to use the chunk's own terminology. Embedding retrieval is then being asked to match text against text it was derived from — an easier problem than matching a user's paraphrase against the same chunk. Retrieval metrics come out optimistic, sometimes substantially, and the gap between eval-set performance and production performance is exactly this. Personas reduce this by pushing the generator into a different register, especially personas whose descriptions explicitly say the user does not know the product's vocabulary. Grounding those persona descriptions in real query logs reduces it further. Neither eliminates it, because the generator still has the chunk in front of it while writing. ## Bias three: generated ground truth The `reference` field is an LLM's answer to its own question, given the contexts it selected. Nobody verified it. This has two consequences worth stating separately. First, the reference can simply be wrong — a subtly incorrect reading of the source, a conflation of two facts, an over-confident claim. Scoring a correct system answer against an incorrect reference marks a good pipeline down, and because it happens per-row rather than systematically, it looks like noise. Second, if the model that generates the references and the model that later judges the answers share a family, they share failure modes and stylistic preferences. The set can quietly reward answers that resemble the generator's own style rather than answers that are correct. Whether that judge relationship invalidates a comparison is broader eval methodology, but the *provenance* of the reference is a fact about the generated set that you own. ## Bias four: coverage is not what it looks like Fifty questions over 5,000 documents touch perhaps a hundred nodes. Which hundred is determined by how the synthesizers sampled the graph, not by what your users care about. A generated set can be systematically silent about the busiest part of your corpus while thoroughly exercising a corner nobody reads. Cross-referencing generated question coverage against real query-volume distribution — even roughly, by topic — is the check most teams skip. ## What the set is actually good for Stated positively, because the answer should not be pure scepticism: - **Relative comparison.** Ranking two retrieval configurations against the same frozen generated set is valid; the biases apply equally to both and cancel. - **Regression detection.** If the score on a frozen set drops after a change, something changed for the worse. That is real signal even if the absolute level is inflated. - **A seed for curation.** The strongest workflow is generate, then have a human read every row: delete the incoherent ones, fix wrong references, keep the good ones. A hundred generated rows curated down to sixty verified ones is worth far more than a thousand unread ones, and it costs an afternoon rather than a week of writing from scratch. ## The operating policy Generate once, curate by hand, freeze the artefact, and version it. Add hand-written unanswerable and adversarial rows the generator cannot produce. Use it for deltas, not for absolute claims. And as production traffic accumulates, migrate the suite toward real queries with human-checked answers — the synthetic set is scaffolding for the period before you have those, and it should shrink as a share of your evaluation over time, not calcify as the permanent definition of quality.

  • What kind of question can a Ragas-generated set never contain, and why does that matter?
    One the corpus cannot answer. Every sample is synthesized from nodes that exist, so the evidence is always there. That leaves the graceful-refusal path completely untested — precisely the behaviour that fails loudly in production when a user asks about something undocumented. Hand-write a few unanswerable rows and mix them in; it is a small effort covering a structural blind spot.
  • How do you curate a generated testset into something you would gate on?
    Read every row. Delete incoherent or ambiguous questions, correct references that misread the source, and drop rows whose reference_contexts do not actually support the answer. Then add hand-written unanswerable and adversarial cases, freeze the file, and version it next to the code. Sixty verified rows beat a thousand unread ones.
  • Why is a generated set still legitimate for comparing two retrieval configurations?
    Because the biases are properties of the set, not of the systems under test, so both candidates face the same inflated-but-consistent instrument and the biases largely cancel in the comparison. The absolute number is untrustworthy; the ordering and the delta are informative, provided the set is frozen and identical across both runs.
  • How should the synthetic share of an evaluation suite change over the first year in production?
    It should shrink. Synthetic generation is scaffolding for the cold start; once real traffic accumulates you have actual queries, actual failures and actual human judgments, which are strictly better material. Keep the generated rows for corpus areas traffic does not yet reach, and let real logged queries with checked answers take over the core of the suite.

saying these in an interview costs you the question

  • Quoting a synthetic testset score as the system's accuracy
  • Assuming the generated reference answer has been verified
  • Believing the generated set covers the corpus proportionally to traffic
  • Testing graceful refusal with questions built from existing chunks
  • Freezing a synthetic suite as the permanent definition of quality

context