skip to content

Metric Classes

The concrete metric objects in ragas.metrics and what each one needs on a sample. The classic trap is reaching for a reference-requiring metric when your production traces have no ground-truth answers.

on this pageshow

questions

6

Which SingleTurnSample fields does each Ragas metric class require?

level: middleimportance: must knowfreq 68%

answer

  1. the sample decides the metric
  2. fields, not taste, pick your metrics
  3. some metrics need ground truth
  4. one metric also needs embeddings
  5. user_input, retrieved_contexts, response, reference

basics

~20 s

Each Ragas metric reads only the SingleTurnSample fields it declares. Faithfulness uses user_input, response and retrieved_contexts; LLMContextRecall and FactualCorrectness also need a human-written reference; ResponseRelevancy needs an embedding model in addition to a judge LLM.

solid answer

~40 s

In Ragas a metric is scored against a `SingleTurnSample`, whose fields are `user_input`, `retrieved_contexts`, `reference_contexts`, `response`, `reference` and `rubrics`. Every metric class declares the subset it consumes. `Faithfulness` and `LLMContextPrecisionWithoutReference` need `user_input`, `response` and `retrieved_contexts`. `LLMContextPrecisionWithReference` swaps the response for `reference`. `LLMContextRecall`, `FactualCorrectness` and `NoiseSensitivity` all require `reference`, which is why they are offline-only metrics. `ResponseRelevancy` needs `user_input` and `response` plus an embeddings object, because it generates candidate questions from the answer and compares them to the real one by cosine similarity. `AspectCritic` and `RubricsScore` judge whatever fields your definition or rubric mentions. The practical consequence is that your data shape, not your taste, decides which metrics you can run: build the sample first, then pick the metrics whose declared fields you actually have.

code

python · 17 lines
python
import asyncio

from langchain_openai import ChatOpenAI
from ragas import SingleTurnSample
from ragas.llms import LangchainLLMWrapper
from ragas.metrics import Faithfulness

evaluator_llm = LangchainLLMWrapper(ChatOpenAI(model="gpt-4o-mini"))

sample = SingleTurnSample(
    user_input="When was the Eiffel Tower completed?",
    response="It was completed in 1889.",
    retrieved_contexts=["The Eiffel Tower was completed in March 1889."],
)

metric = Faithfulness(llm=evaluator_llm)
print(asyncio.run(metric.single_turn_ascore(sample)))

go deeper

for a junior

Know that Ragas scores a SingleTurnSample, and be able to name its main fields: user_input, retrieved_contexts, response and reference. Say plainly that different metrics read different fields.

for a middle

Be ready to map each built-in metric to its required fields on the spot, and to explain that ResponseRelevancy additionally needs an embeddings object because it compares generated questions by cosine similarity.

for a senior

Show that you plan the sample shape before the metric list: production traces give you three fields for free, so reference-requiring metrics belong to a curated offline set you can afford to label.

for a principal

Own the tradeoff of how much ground truth your org will actually maintain. The field inventory sets a hard ceiling on which metrics can ever run continuously, and that ceiling, not the metric catalogue, is what your quality programme is built on.

## The sample is the contract Ragas does not hand a metric a free-form prompt and answer. It hands it a `SingleTurnSample`, a typed record with a fixed vocabulary of optional fields: - `user_input` — the end user's question. - `retrieved_contexts` — the list of chunks your retriever actually returned for that question. - `reference_contexts` — the chunks that *should* have been retrieved, when you have that label. - `response` — what your application generated. - `reference` — the human-written ground-truth answer. - `rubrics` — a per-sample rubric mapping, used when you want a sample to carry its own grading scale. Every field is optional on the object. Nothing validates that a sample is complete at construction time. The validation happens when a metric tries to read a field it needs. ## Who needs what **Faithfulness** reads `user_input`, `response` and `retrieved_contexts`. It never looks at `reference`, which is the single most important fact about it operationally: it is scoreable on raw production traffic. **ResponseRelevancy** reads `user_input` and `response`. Beyond fields, it also needs an embeddings object passed to the constructor alongside the judge LLM, because its mechanism is to have the LLM write `strictness` candidate questions from the response and then embed them to measure similarity against the real `user_input`. A metric that silently needs a second model is the classic wiring surprise here. **LLMContextPrecisionWithReference** reads `user_input`, `reference` and `retrieved_contexts`. **LLMContextPrecisionWithoutReference** reads `user_input`, `response` and `retrieved_contexts` — the class pair exists precisely so you can trade the ground-truth answer for the model's own answer as the judging target. **LLMContextRecall** requires `reference` together with `retrieved_contexts`, since its whole method is decomposing the ground-truth answer and asking whether each piece is attributable to something that was retrieved. No reference, no metric. **FactualCorrectness** compares `response` against `reference` by claim decomposition, so both must be present. **NoiseSensitivity** needs `user_input`, `response`, `reference` and `retrieved_contexts` at once — it is the most data-hungry of the built-ins, which is why it usually shows up on a small curated set rather than a large one. **AspectCritic** and **RubricsScore** are the open-ended pair. They read the fields your natural-language `definition` or rubric descriptions refer to, so they can be reference-free or reference-requiring depending on how you write them. ## What happens when a field is missing A metric that cannot find a required field does not quietly return zero. Scored directly, it fails for that sample; run through a dataset evaluation you see the failure surface as an error or a missing value in the results rather than a real score. That is the behaviour you want — a silent zero would look like a quality regression instead of a plumbing bug — but it means a half-populated dataset shows up as a broken run, and juniors often read the traceback as "the metric is broken" rather than "my sample is incomplete". ## Read the fields backwards The useful habit is to enumerate what your pipeline can actually emit before choosing metrics. A live application emits `user_input`, `retrieved_contexts` and `response` for free — they are already in the request path. `reference` is a human annotation that exists only where somebody wrote it, and `reference_contexts` only where somebody labelled the corpus. So the field inventory partitions the metric catalogue automatically: the three-field metrics run everywhere and continuously; the reference-requiring ones run on your curated set on the cadence at which you can afford to maintain labels. A final practical note: field names changed across Ragas major versions (the older column-style names such as `question`, `answer`, `contexts` and `ground_truth` are what you will find in pre-0.2 blog posts). On 0.4.x the names above are the current ones, and `ResponseRelevancy` is the current class name for what older material calls `AnswerRelevancy`. Copying a two-year-old snippet is the most common way to get a field error on your first run.

  • What is reference_contexts for, given that retrieved_contexts already exists?
    `retrieved_contexts` is what your retriever actually returned; `reference_contexts` is what it should have returned, a human or generator-supplied label about the corpus rather than about the answer. The LLM-judged context metrics work from `retrieved_contexts`, while `reference_contexts` is the field you populate when you have gold retrieval labels and want to compare against them without spending judge calls.
  • Why does ResponseRelevancy need an embedding model when the other metrics do not?
    Its mechanism is different. Most Ragas metrics ask the judge LLM for verdicts on claims. ResponseRelevancy instead asks the LLM to reverse-engineer questions from the response, then measures how close those questions are to the real `user_input` — and that closeness is a cosine similarity, which needs vectors. So it takes both an LLM and an embeddings object, and a run configured with only an LLM fails on that metric alone.
  • A dataset has every field populated except reference. Which built-in metrics still run?
    Faithfulness, ResponseRelevancy and LLMContextPrecisionWithoutReference all run, plus AspectCritic or RubricsScore if you wrote their criteria to judge the response against the retrieved context rather than against a ground truth. LLMContextRecall, FactualCorrectness, NoiseSensitivity and LLMContextPrecisionWithReference cannot run — they read `reference` directly.

saying these in an interview costs you the question

  • Assuming every Ragas metric needs a ground-truth answer
  • Using pre-0.2 field names like question, answer, contexts, ground_truth
  • Thinking a missing field just yields a score of zero
  • Wiring only an LLM and expecting ResponseRelevancy to run
  • Believing retrieved_contexts and reference_contexts are interchangeable

context

open as a page

Which Ragas metrics run on production traces with no reference answers?

level: seniorimportance: must knowfreq 58%

basics

~10 s

Faithfulness, ResponseRelevancy and LLMContextPrecisionWithoutReference run on unlabelled traces, as do AspectCritic and RubricsScore when their criteria judge the response against retrieved context. LLMContextRecall, FactualCorrectness, NoiseSensitivity and LLMContextPrecisionWithReference all read reference and cannot.

open as a page

In Ragas, what does a metric's single_turn_ascore() return and why await it?

level: juniorimportance: should knowfreq 36%

basics

~20 s

single_turn_ascore() scores one SingleTurnSample with one metric and resolves to a float. It is asynchronous, so calling it without await or asyncio.run() hands back a coroutine object that never executes and never contacts the judge model.

open as a page

In Ragas, when do you use AspectCritic instead of RubricsScore?

level: middleimportance: should knowfreq 42%

basics

~20 s

AspectCritic returns a binary 0 or 1 verdict against a natural-language definition, so it suits pass/fail gates and safety checks. RubricsScore returns a graded score against named band descriptions, so it suits tracking gradual quality change where a hard boundary would be arbitrary.

open as a page

In Ragas FactualCorrectness, what do mode='precision' and mode='recall' change?

level: middleimportance: should knowfreq 38%

basics

~20 s

FactualCorrectness breaks both the response and the reference into claims. Precision mode scores what fraction of the response's claims the reference supports, punishing invented extras. Recall mode scores what fraction of the reference's claims the response contains, punishing omissions. The default f1 combines both.

open as a page

Why does each Ragas metric cost more than one judge LLM call per sample?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Ragas metrics are multi-step judge pipelines, not single prompts. Faithfulness extracts statements then verifies each against context; the context precision metrics issue a verdict per retrieved chunk; FactualCorrectness decomposes two texts and compares claim sets. Cost scales with metrics times samples times chunks.

open as a page