In LlamaIndex, how do FaithfulnessEvaluator, RelevancyEvaluator and CorrectnessEvaluator differ?
answer
- three judges, different inputs
- one of them needs ground truth
- grounded vs on-topic vs right
- score, passing, feedback on every result
basics
~20 sFaithfulnessEvaluator checks whether the answer is supported by the retrieved context. RelevancyEvaluator checks whether the retrieved context and answer address the query. CorrectnessEvaluator scores the answer against a labelled reference answer, so only it needs ground truth.
solid answer
~50 sAll three are LLM-as-judge evaluators in `llama_index.core.evaluation`, but they take different inputs and answer different questions. `FaithfulnessEvaluator` takes the response plus its retrieved contexts and asks "is this answer supported by what was retrieved?" — a hallucination check that needs no labels, so you can run it on production traffic. `RelevancyEvaluator` additionally takes the query and asks whether the retrieved context and the answer are actually on-topic for that question — it catches retrieval that pulled the wrong chunks. `CorrectnessEvaluator` compares the answer to a `reference` answer you supply and returns a 1–5 score with a default passing threshold of 4.0, so it requires a labelled eval set. Each returns an `EvaluationResult` carrying `passing`, `score` and `feedback`; call `evaluate_response(response=...)` with a query engine's `Response`, or `evaluate(query=..., response=..., contexts=[...])` with raw strings, or the async `aevaluate*` variants.
code
python · 26 linesfrom llama_index.core import Document, VectorStoreIndex
from llama_index.core.evaluation import (
CorrectnessEvaluator,
FaithfulnessEvaluator,
RelevancyEvaluator,
)
from llama_index.llms.openai import OpenAI
judge = OpenAI(model="gpt-4o", temperature=0)
index = VectorStoreIndex.from_documents(
[Document(text="Refunds are accepted within 30 days of purchase.")]
)
query_engine = index.as_query_engine()
query = "What is the refund window?"
response = query_engine.query(query)
f = FaithfulnessEvaluator(llm=judge).evaluate_response(response=response)
r = RelevancyEvaluator(llm=judge).evaluate_response(query=query, response=response)
c = CorrectnessEvaluator(llm=judge).evaluate(
query=query,
response=str(response),
reference="Refunds are accepted within 30 days.",
)
print(f.passing, r.passing, c.score, c.feedback)go deeper
Be able to name the three evaluators and say in one line what each checks: grounded in context, on-topic for the query, and matching a reference answer.
Explain which inputs each evaluator needs, that they are LLM-as-judge calls returning passing/score/feedback, and why only correctness requires a labelled set.
Show how you use them in production: faithfulness on sampled live traffic because it needs no labels, relevancy to localise a failure to retrieval, correctness on a curated regression set.
Own the judge policy — which model judges, why it differs from the generator, what judge drift does to historical comparability, and what evaluation costs per release.
## What these classes are LlamaIndex ships a family of evaluators under `llama_index.core.evaluation`. They are not statistical metrics; they are **LLM-as-judge** components. Each one builds a prompt from your query, your answer and (usually) the retrieved chunks, sends it to a judge LLM, and parses the reply into an `EvaluationResult`. That object carries `query`, `contexts`, `response`, a boolean `passing`, a numeric `score`, and a free-text `feedback` string explaining the verdict. The feedback field is the part engineers under-use — it is what turns a red number into a diagnosis. Because they are LLM calls, every evaluation costs money and latency, and results carry judge variance. Pin the judge model and set `temperature=0` so a re-run of the same eval set is comparable. ## FaithfulnessEvaluator — grounding `FaithfulnessEvaluator` asks a single question: is the generated answer supported by the retrieved context? It needs the response and the contexts, and nothing else — no query is required, and crucially **no ground-truth answer**. That property is what makes it the workhorse of production monitoring: you can sample live traffic and score it, because live traffic has no labels. It returns a binary verdict (`passing` True/False, score 1.0/0.0 by default). A failure means the model asserted something the retrieved chunks do not say — the classic RAG hallucination. It says nothing about whether the answer is *useful*: a faithful answer can be perfectly grounded in irrelevant chunks and still be worthless to the user. ## RelevancyEvaluator — on-topic-ness `RelevancyEvaluator` takes the query as well, and judges whether the retrieved context and the response actually address that query. A failure here usually points upstream at retrieval — wrong chunks came back, so even a well-behaved generator had nothing useful to work with. Running faithfulness and relevancy together is informative precisely because they fail in different places: faithful-but-irrelevant means retrieval missed; relevant-but-unfaithful means the generator embroidered. There are also finer-grained variants in the same package, `AnswerRelevancyEvaluator` and `ContextRelevancyEvaluator`, which split the judgement into "does the answer address the query" and "is the retrieved context relevant to the query" respectively. ## CorrectnessEvaluator — needs labels `CorrectnessEvaluator` is the only one of the three that compares against ground truth. You pass `query`, `response` and `reference`, and it returns a score on a 1–5 scale plus feedback; `passing` is derived from a `score_threshold` that defaults to 4.0. Because it needs a `reference`, it can only run against a curated eval set — you cannot point it at unlabelled production traffic. This is the metric that answers the question a stakeholder actually asks ("is the answer right?"), which is why teams build a golden set of a few dozen to a few hundred question/reference pairs and treat it as a regression suite. ## Calling them Two call shapes exist. `evaluate_response(query=..., response=...)` takes the `Response` object a query engine returned and pulls the source nodes out of it for you — the convenient path. `evaluate(query=..., response=<str>, contexts=[<str>, ...])` takes plain strings, which is what you want when the answer came from somewhere else, or when you are replaying stored traces. Every evaluator also exposes async `aevaluate` / `aevaluate_response`, which is what batch evaluation is built on. ## Choosing between them - No labels, need a hallucination guard: faithfulness. - Suspect retrieval is pulling the wrong chunks: relevancy (and look at the retrieved context in the trace). - Comparing two pipeline configurations on a curated set: correctness, plus faithfulness and relevancy alongside so you can see *why* a score moved. A further evaluator, `PairwiseComparisonEvaluator`, sidesteps absolute scoring altogether and asks the judge which of two answers is better — useful for A/B comparisons where absolute 1–5 scores are noisy. ## Common failure modes Using the same model as both generator and judge invites self-preference bias; a stronger, separate judge is the usual mitigation. Reading only `passing` and ignoring `feedback` throws away the diagnostic. And treating a single 20-question run as evidence is the biggest one: with a small set and a non-deterministic judge, a two-point move is noise, not a result.
- A response scores as faithful but the user says it is useless. What does that tell you?Faithfulness only asks whether the answer is supported by the retrieved chunks, so a perfectly grounded answer built on the wrong chunks passes. That pattern points at retrieval, not generation. Run `RelevancyEvaluator`, which brings the query into the judgement, and inspect the retrieved nodes — if the supporting fact never came back, the fix is in chunking, top-k or the embedding model, not the prompt.
- Why would you use a different model for judging than for generating?A model asked to grade its own output tends toward self-preference, so scores drift optimistic. Using a stronger, independent judge decouples the score from the system under test, which matters most when you are comparing two generator configurations. The cost is that changing the judge invalidates comparison with earlier runs, so the judge model and temperature should be pinned and versioned alongside the eval set.
- How do you turn CorrectnessEvaluator's 1-5 score into a pass/fail gate?It already does that: `CorrectnessEvaluator` has a `score_threshold` defaulting to 4.0, and sets `passing` on the `EvaluationResult` from it. You can raise or lower the threshold for a stricter or looser gate. In practice teams gate on the aggregate — mean score plus pass rate across the whole eval set — rather than on any single row, because a single judge call is too noisy to block a release on.
saying these in an interview costs you the question
- Claiming faithfulness needs a reference answer
- Thinking these are deterministic metrics, not LLM judgements
- Using faithfulness alone and calling retrieval validated
- Ignoring the feedback field and reading only pass/fail
- Judging with the same model that generated the answer