skip to content

Ragas

The RAG-focused evaluation library: metrics that separate retrieval quality from generation quality, so you learn whether the context was wrong or the answer was. Interviewers like it because it forces a precise statement of what "the RAG is bad" actually means.

on this pageshow

explore

questions

24

In Ragas, how do you run evaluate() over an EvaluationDataset of samples?

level: juniorimportance: must knowfreq 70%

answer

  1. one call over a whole dataset
  2. each record becomes a typed sample
  3. wrap the list before you score
  4. evaluate takes dataset, metrics, judge model
  5. result is an object, not a dict

basics

~20 s

Ragas scores a whole dataset in one call. Put each record into a SingleTurnSample, wrap the list in an EvaluationDataset, then call evaluate(dataset=..., metrics=[...], llm=...). It returns an EvaluationResult carrying per-metric averages and a to_pandas() view.

solid answer

~40 s

A Ragas run has three pieces. First the **samples**: a `SingleTurnSample` holds the fields a metric may read — `user_input`, `retrieved_contexts`, `response`, and `reference` when the metric needs ground truth. Second the **dataset**: `EvaluationDataset(samples=[...])`, or `EvaluationDataset.from_list(list_of_dicts)` when the data already exists as records. Third the **run**: `evaluate(dataset=ds, metrics=[Faithfulness()], llm=evaluator_llm)`, which fans the sample × metric jobs out concurrently rather than looping in Python. It hands back an `EvaluationResult`; printing it shows one averaged number per metric, and `result.to_pandas()` gives one row per sample with a column per metric so you can look at individual failures. Most metrics call a judge model, so an evaluator LLM must be supplied — either on `evaluate()` or on each metric.

code

python · 19 lines
python
from ragas import evaluate, EvaluationDataset
from ragas.dataset_schema import SingleTurnSample
from ragas.metrics import Faithfulness
from ragas.llms import LangchainLLMWrapper
from langchain_openai import ChatOpenAI

sample = SingleTurnSample(
    user_input="When was the Eiffel Tower completed?",
    retrieved_contexts=["The Eiffel Tower was completed in 1889."],
    response="It was completed in 1889.",
    reference="1889",
)

dataset = EvaluationDataset(samples=[sample])
evaluator_llm = LangchainLLMWrapper(ChatOpenAI(model="gpt-4o-mini"))

result = evaluate(dataset=dataset, metrics=[Faithfulness()], llm=evaluator_llm)
print(result)
df = result.to_pandas()

go deeper

for a junior

Be able to name the three pieces without notes: a sample per record, an EvaluationDataset holding them, and one evaluate() call with metrics and a judge model. Mention that to_pandas() shows per-sample scores.

for a middle

Explain that evaluate() fans sample-by-metric jobs out concurrently rather than looping, that each metric declares the sample fields it requires, and that reference-requiring metrics are unusable when you have no ground truth.

for a senior

Show where the run fits in a real workflow: capturing records from production or a golden set, choosing metrics against the fields you actually have, and reading the worst rows of the dataframe rather than reporting the mean.

for a principal

Own the framing that Ragas scores captured records, not a live pipeline — so the value of any run is bounded by how representative the captured dataset is and by what you can afford to score at your traffic volume.

## What a Ragas run actually is Ragas does not score "a RAG pipeline". It scores **records that a pipeline already produced**. You run your application first, capture what it did for each question, and then hand that captured data to Ragas. That separation is why the API has three distinct objects: a sample (one record), a dataset (a homogeneous collection of records), and a run (`evaluate`) that returns a result object. ## The sample `SingleTurnSample` is a typed container for one question-and-answer interaction. Its fields map onto what RAG metrics need: - `user_input` — the question that was asked. - `retrieved_contexts` — the list of context strings your retriever returned. - `response` — what your generator produced. - `reference` — the ground-truth answer, present only when you have one. - `reference_contexts` — the contexts that *should* have been retrieved, for recall-style metrics. Not every field is needed for every metric; each metric declares which fields it reads, and a dataset missing a required field fails validation rather than silently scoring badly. For conversations there is `MultiTurnSample`, whose `user_input` is a list of message objects instead of a string; a dataset must be all one sample type, and only multi-turn-capable metrics accept it. ## The dataset `EvaluationDataset(samples=[sample1, sample2, ...])` is the direct constructor. When your data is already a list of dictionaries — the usual shape after pulling rows out of a warehouse, a JSONL file, or a dataframe converted with `to_dict("records")` — `EvaluationDataset.from_list(records)` builds the samples for you, and the dictionary keys must match the sample field names. The dataset is deliberately dumb: it holds records and knows its sample type. It contains no scoring logic and no configuration. ## The run `evaluate(dataset=ds, metrics=[...], llm=evaluator_llm)` is the single entry point. Two things about it surprise people the first time: **It is not a loop.** Internally Ragas builds one job per (sample, metric) pair and executes them concurrently through an executor, with a progress bar. A hundred samples and three metrics is three hundred jobs, not three hundred sequential HTTP round trips. Concurrency, per-call timeouts and retries are configured through a `RunConfig` passed to `evaluate`. **It needs a judge model.** Almost every metric in Ragas is LLM-based: it prompts a model to extract claims, check them against context, or produce a verdict. You supply that evaluator model either once on `evaluate(llm=...)` or individually on each metric object. Forgetting it is the most common first-run error. Some metrics also need an embedding model. ## The result `evaluate` returns an `EvaluationResult`. Printing it renders a dictionary of one averaged number per metric — the headline. `result.to_pandas()` returns a dataframe with one row per sample: the sample's own fields as columns, plus a column per metric holding that sample's individual score. That per-row view is where the useful information is; a single mean tells you a number moved but never which query broke. ## Why running it this way matters Calling `single_turn_ascore` on one sample at a time in a Python loop technically works and is how you debug a single case, but it gives up the executor, the shared run configuration, the progress reporting, the aggregate result object, and the dataframe. For anything beyond a handful of records, the batch call is the API you want. ## The shape of a real workflow Capture production or golden-set records → build `SingleTurnSample`s → wrap in `EvaluationDataset` → choose the metrics that match the fields you have (reference-requiring metrics are unusable on raw production traffic with no ground truth) → wire an evaluator model → `evaluate` → read `to_pandas()`, sort ascending by the metric column, and read the worst rows. The last step is the one candidates skip and interviewers ask about.

  • What changes if the records are multi-turn conversations rather than single question-answer pairs?
    You build `MultiTurnSample` objects instead, whose `user_input` is a list of message objects rather than a single string. A dataset must be homogeneous — all single-turn or all multi-turn — and only metrics that declare multi-turn support will accept it. Mixing the two sample types in one `EvaluationDataset` is not a supported run.
  • You already have the data in a pandas dataframe. How do you get it into an EvaluationDataset?
    Convert the dataframe to records with `df.to_dict("records")` and pass that list to `EvaluationDataset.from_list(...)`. The dictionary keys must match the sample field names exactly — `user_input`, `retrieved_contexts`, `response`, `reference` — and `retrieved_contexts` must be a list of strings, not a single concatenated blob, or context-level metrics score nonsense.
  • When would you call a metric on one sample directly instead of using evaluate()?
    When debugging a single case — you want the score for one record and the metric's reasoning without spinning up a run. `single_turn_ascore` on the metric object does exactly that. For anything at scale you lose the concurrency executor, shared run configuration, progress reporting, the aggregate result object and the dataframe, so batch evaluation is the default.

Think of it like a test runner: the samples are the test cases, the metrics are the assertions, and evaluate() is the runner that executes every case-assertion pair and returns a report.

saying these in an interview costs you the question

  • Scoring rows one by one in a Python loop instead of evaluate()
  • Passing a bare list of dicts straight into evaluate()
  • Assuming evaluate() runs samples sequentially
  • Forgetting that most metrics need an evaluator LLM configured
  • Thinking the returned result is a plain dict of averages

context

open as a page

Which Ragas calls turn your own documents into a synthetic testset?

level: juniorimportance: must knowfreq 55%

basics

~10 s

Two calls. TestsetGenerator.from_langchain(llm, embedding_model) builds the generator from a LangChain model pair, then generator.generate_with_langchain_docs(docs, testset_size=N) returns a Testset of N synthetic samples with questions, reference contexts and reference answers.

open as a page

In Ragas, what do RunConfig's max_workers, timeout and max_retries control?

level: middleimportance: must knowfreq 58%

basics

~20 s

RunConfig tunes how a Ragas run executes: max_workers caps how many metric calls run concurrently (default 16), timeout bounds a single call (default 180 seconds), and max_retries is how many times a failed call is retried with backoff (default 10). Pass it as evaluate(run_config=RunConfig(...)).

open as a page

In Ragas, how do you give metrics an evaluator LLM and embedding model?

level: middleimportance: must knowfreq 78%

basics

~20 s

Wrap a LangChain chat model in LangchainLLMWrapper and a LangChain embeddings object in LangchainEmbeddingsWrapper, then pass them to evaluate() as llm= and embeddings=, or set them on individual metric objects. Ragas ships no model of its own.

open as a page

Which SingleTurnSample fields does each Ragas metric class require?

level: middleimportance: must knowfreq 68%

basics

~20 s

Each Ragas metric reads only the SingleTurnSample fields it declares. Faithfulness uses user_input, response and retrieved_contexts; LLMContextRecall and FactualCorrectness also need a human-written reference; ResponseRelevancy needs an embedding model in addition to a judge LLM.

open as a page

In Ragas, how do you control the single-hop vs multi-hop query mix?

level: middleimportance: must knowfreq 48%

basics

~10 s

Pass query_distribution: a list of (synthesizer, weight) tuples whose weights sum to 1, using SingleHopSpecificQuerySynthesizer, MultiHopAbstractQuerySynthesizer and MultiHopSpecificQuerySynthesizer. default_query_distribution(llm) supplies a starting mix you can reweight.

open as a page

Which Ragas metrics run on production traces with no reference answers?

level: seniorimportance: must knowfreq 58%

basics

~10 s

Faithfulness, ResponseRelevancy and LLMContextPrecisionWithoutReference run on unlabelled traces, as do AspectCritic and RubricsScore when their criteria judge the response against retrieved context. LLMContextRecall, FactualCorrectness, NoiseSensitivity and LLMContextPrecisionWithReference all read reference and cannot.

open as a page

In Ragas, what does a metric's single_turn_ascore() return and why await it?

level: juniorimportance: should knowfreq 36%

basics

~20 s

single_turn_ascore() scores one SingleTurnSample with one metric and resolves to a float. It is asynchronous, so calling it without await or asyncio.run() hands back a coroutine object that never executes and never contacts the judge model.

open as a page

In Ragas, what does EvaluationResult.to_pandas() return after an evaluation run?

level: middleimportance: should knowfreq 45%

basics

~20 s

to_pandas() returns a dataframe with one row per evaluated sample: the sample's own fields as columns plus one column per metric holding that sample's individual score. Printing the result object instead shows only the per-metric average across samples.

open as a page

Which Ragas metrics need an embedding model, not just a judge LLM?

level: middleimportance: should knowfreq 40%

basics

~20 s

Most ragas metrics are pure LLM-as-judge and need only an evaluator LLM. Similarity-based ones — ResponseRelevancy above all — additionally embed text and compare vectors, so a run containing them fails or misbehaves if you wired an LLM but no embeddings.

open as a page

In Ragas, when do you use AspectCritic instead of RubricsScore?

level: middleimportance: should knowfreq 42%

basics

~20 s

AspectCritic returns a binary 0 or 1 verdict against a natural-language definition, so it suits pass/fail gates and safety checks. RubricsScore returns a graded score against named band descriptions, so it suits tracking gradual quality change where a hard boundary would be arbitrary.

open as a page

In Ragas FactualCorrectness, what do mode='precision' and mode='recall' change?

level: middleimportance: should knowfreq 38%

basics

~20 s

FactualCorrectness breaks both the response and the reference into claims. Precision mode scores what fraction of the response's claims the reference supports, punishing invented extras. Recall mode scores what fraction of the reference's claims the response contains, punishing omissions. The default f1 combines both.

open as a page

In Ragas testset generation, what do knowledge-graph transforms do?

level: middleimportance: should knowfreq 42%

basics

~20 s

Transforms enrich the KnowledgeGraph before any question is written. Splitters break documents into smaller nodes, extractors attach properties like headlines, summaries, keyphrases and entities, and relationship builders draw edges between related nodes so multi-hop synthesizers have something to traverse.

open as a page

How do you stop a Ragas re-run from re-paying for identical judge calls?

level: seniorimportance: should knowfreq 28%

basics

~20 s

Attach a cache backend to the evaluator model — Ragas ships DiskCacheBackend in ragas.cache, passed as cache= when constructing the evaluator LLM wrapper. Identical judge requests then hit a local cache directory instead of the provider, so re-runs over unchanged data are fast and free.

open as a page

A Ragas run finishes but some metric scores are NaN — what happened, and what do you do?

level: seniorimportance: should knowfreq 40%

basics

~20 s

By default a Ragas run does not raise on a failing metric call — the exception is swallowed and that sample's score is recorded as NaN. Causes are exhausted retries, timeouts, provider errors, or unparseable judge output. Count the NaNs before trusting the averages.

open as a page

How do you measure judge-model token usage and cost for a Ragas evaluation run?

level: seniorimportance: should knowfreq 33%

basics

~10 s

Pass a token usage parser to evaluate() — for OpenAI-style judges, token_usage_parser=get_token_usage_for_openai from ragas.cost. The result then answers result.total_tokens() and result.total_cost(cost_per_input_token=..., cost_per_output_token=...). Without a parser, no usage is collected.

open as a page

How do you point Ragas at a judge model different from the app's model?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Build a second wrapper instance around a different chat model and hand it to ragas — the evaluator LLM is a wholly separate object from anything your pipeline uses. Pin a dated model snapshot, give it its own API key, and record it with the scores.

open as a page

In Ragas, how do you export judge LLM calls to Langfuse or LangSmith?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Pass a callback handler to evaluate()'s callbacks argument: because the evaluator LLM is a wrapped LangChain model, a LangChain-compatible handler sees every judge call. Ragas also ships langsmith and tracing.langfuse integration modules for the same purpose.

open as a page

Why does each Ragas metric cost more than one judge LLM call per sample?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Ragas metrics are multi-step judge pipelines, not single prompts. Faithfulness extracts statements then verifies each against context; the context precision metrics issue a verdict per retrieved chunk; FactualCorrectness decomposes two texts and compares claim sets. Cost scales with metrics times samples times chunks.

open as a page

Your Ragas testset run over 5,000 documents takes hours — what do you change?

level: seniorimportance: should knowfreq 36%

basics

~20 s

The cost is in the transforms, which run per node over the whole corpus, not in the questions, which scale with testset_size. Build and enrich the KnowledgeGraph once, save it with kg.save(), reload it for later runs, and use a cheaper transforms_llm.

open as a page

How do you turn a Ragas evaluation into a CI gate on a pull request?

level: principalimportance: should knowfreq 45%

basics

~20 s

Run evaluate() over a small pinned dataset, read aggregate scores, and fail the job when a metric falls below its floor. The hard part is not the assert — it is choosing thresholds that catch regressions without going red at random, and keeping the run cheap enough to sit on every pull request.

open as a page

Where does a Ragas-generated synthetic testset mislead you about quality?

level: principalimportance: should knowfreq 30%

basics

~20 s

Every question is written from chunks already in the corpus, so retrieval is being tested on queries constructed to be answerable — no unanswerable questions, no vocabulary your corpus lacks, and reference answers written by the same model family that will be judged.

open as a page

How do you run Ragas against a LlamaIndex or Haystack pipeline?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Ragas ships a first-party llama_index integration module, but no Haystack one — Haystack support is the separate deepset-published ragas-haystack distribution providing a RagasEvaluator component. Either way, the universal fallback is to build SingleTurnSample objects yourself from whatever your pipeline returns.

open as a page

What are personas in Ragas testset generation, and why set them?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

A Persona has a name and a role_description, and it tells the query synthesizers whose voice to write in. Passing persona_list to TestsetGenerator makes generated questions read like your real user segments; leave it out and Ragas invents personas from the knowledge graph.

open as a page