skip to content

Which Ragas metrics run on production traces with no reference answers?

level: seniorimportance: must knowfreq 58%

answer

  1. ground truth is the dividing line
  2. two suites, two cadences
  3. one precision class per judging target
  4. never synthesise reference from response
  5. recall questions need labels, always

basics

~10 s

Faithfulness, ResponseRelevancy and LLMContextPrecisionWithoutReference run on unlabelled traces, as do AspectCritic and RubricsScore when their criteria judge the response against retrieved context. LLMContextRecall, FactualCorrectness, NoiseSensitivity and LLMContextPrecisionWithReference all read reference and cannot.

solid answer

~40 s

Ragas splits cleanly into two families. Reference-free: `Faithfulness`, `ResponseRelevancy`, `LLMContextPrecisionWithoutReference`, and `AspectCritic` / `RubricsScore` when you write their criteria to judge the answer against the retrieved context. Reference-requiring: `LLMContextPrecisionWithReference`, `LLMContextRecall`, `FactualCorrectness` and `NoiseSensitivity`. The paired precision classes exist exactly for this problem — `WithoutReference` uses your own `response` as the judging target instead of a ground-truth answer, which buys you a runnable metric at the cost of a self-referential one. So the production setup is two suites: the reference-free set over sampled live traces, continuously, and the reference-requiring set over a curated labelled dataset on whatever cadence you can maintain labels. The failure mode to name is synthesising a `reference` from the response and then running the reference-requiring metrics on it — that measures the model's agreement with itself.

go deeper

for a junior

Know that some Ragas metrics need a ground-truth answer and some do not, and be able to name Faithfulness as one that runs without one.

for a middle

Be ready to sort the built-in metric classes into the reference-free and reference-requiring families, and to explain what LLMContextPrecisionWithoutReference substitutes for the missing reference.

for a senior

Show the two-suite design — reference-free metrics sampled continuously over live traces, reference-requiring metrics gated over a curated labelled set — and name the self-referential trap of synthesising references from responses.

for a principal

Own the annotation budget as the real constraint. The size and freshness of the labelled set determines which questions your organisation can ever answer, so argue for it as infrastructure rather than treating metric choice as a library decision.

## Why the split matters at all Production traffic gives you three things for free: the question the user asked, the chunks your retriever returned, and the answer your model produced. It never gives you the correct answer, because if you had that you would have shipped it. Ragas's metric catalogue is therefore partitioned by a hard data requirement, and the first thing a senior engineer does with the catalogue is sort it. ## The reference-free set **Faithfulness** takes `user_input`, `response` and `retrieved_contexts`. It asks whether the answer is supported by the chunks that were actually retrieved — a purely internal consistency question, so no ground truth is needed. **ResponseRelevancy** takes `user_input` and `response` plus an embeddings object. It has the judge reverse-generate questions from the answer and measures how close they are to the real question. Again, self-contained. **LLMContextPrecisionWithoutReference** takes `user_input`, `response` and `retrieved_contexts`. This is the deliberate substitution: where the `WithReference` class asks "was this chunk useful for arriving at the ground-truth answer?", the `WithoutReference` class asks "was this chunk useful for arriving at the answer we actually produced?". Same ranking-aware precision computation, different judging target. **AspectCritic** and **RubricsScore** are reference-free or not depending on you. An AspectCritic whose `definition` reads "Does the response contain any claim not present in the retrieved context?" runs on live traffic. One that reads "Does the response match the ground truth?" does not, because it names a field that is empty. ## The reference-requiring set `LLMContextPrecisionWithReference`, `LLMContextRecall`, `FactualCorrectness` and `NoiseSensitivity` all read `reference`. These are your offline metrics. They answer questions the reference-free set structurally cannot — most importantly recall-shaped questions, since "did we retrieve everything needed" is unanswerable without knowing what was needed. ## The two-suite pattern The arrangement that survives contact with production is two distinct suites with different cadences. One suite runs the reference-free metrics over a sample of live traces. It is continuous, cheap per sample because you control the sample rate, and it is a monitor: it tells you *something changed*, not *the change is bad*. This suite catches retrieval drift, prompt regressions and a model swap that broke grounding. The other suite runs the full catalogue over a curated labelled dataset — a few hundred examples where a human wrote the `reference`. It is a gate: you run it before a release, compare against the previous run, and it is the only place recall-shaped and correctness-shaped numbers exist. Its size is bounded by annotation budget, not by compute. ## The trap, stated plainly The move that looks clever and is not: take the production `response`, write it into `reference`, and unlock the reference-requiring metrics on live traffic. Every one of those metrics then compares the answer to itself. `FactualCorrectness` returns approximately a perfect score forever. `LLMContextRecall` reports that retrieval covered everything the answer needed — which is trivially true, because the answer was written from what was retrieved. You have manufactured a dashboard that cannot go red, which is strictly worse than having no dashboard, because someone will trust it. A weaker variant of the same trap is generating `reference` with a stronger model and treating it as ground truth without ever auditing it. That can be legitimate — but only if you have sampled and hand-checked the generated references, and only if you are honest in the reporting that the number is model-versus-model agreement, not correctness. ## Practical selection Given a leaf-level metric choice, the decision procedure is short. Do I have `reference` for this data? If yes, the full catalogue is open. If no, take Faithfulness for grounding, LLMContextPrecisionWithoutReference for retrieval usefulness, ResponseRelevancy for on-topic-ness, and an AspectCritic or two for the domain-specific properties you care about — refusal behaviour, tone, presence of a citation. Then accept that no reference-free metric can tell you the answer was *right*, only that it was consistent with what you retrieved, and route the correctness question to the labelled suite.

  • What exactly does LLMContextPrecisionWithoutReference use in place of the missing reference?
    Your application's own `response`. It asks the judge, chunk by chunk, whether that chunk was useful in arriving at the response that was produced, and combines the verdicts with the same rank-weighted precision computation as the WithReference class. The number is therefore relative to the answer you actually gave, so it degrades in a specific way: if the model ignored good context and answered from parametric memory, the chunks look useless.
  • Why can't you get any recall-shaped signal on unlabelled production traffic?
    Recall asks whether everything required was retrieved, and "required" is defined by the correct answer, which unlabelled traffic does not contain. LLMContextRecall therefore stays in the offline suite. The nearest reference-free proxies are indirect — refusal rate, a fall-back-to-parametric-knowledge critic, or user thumbs-down feedback — and none of them is the metric; they are symptoms you correlate with it.
  • How do you keep the labelled suite from going stale as production traffic shifts?
    Treat it as a living dataset rather than a fixture. Periodically sample real production questions, cluster them, and check that each cluster is represented in the labelled set; when a new query shape appears at volume with no representative, that is the annotation backlog. Also re-check labels when the corpus changes, because a reference written against last quarter's documents can silently become wrong.

saying these in an interview costs you the question

  • Writing the response into reference to unlock more metrics
  • Claiming Faithfulness proves the answer is correct
  • Believing there is a reference-free recall metric
  • Running one suite at one cadence for both jobs
  • Treating a generated reference as audited ground truth

context