In Ragas, how do you run evaluate() over an EvaluationDataset of samples?
answer
- one call over a whole dataset
- each record becomes a typed sample
- wrap the list before you score
- evaluate takes dataset, metrics, judge model
- result is an object, not a dict
basics
~20 sRagas scores a whole dataset in one call. Put each record into a SingleTurnSample, wrap the list in an EvaluationDataset, then call evaluate(dataset=..., metrics=[...], llm=...). It returns an EvaluationResult carrying per-metric averages and a to_pandas() view.
solid answer
~40 sA Ragas run has three pieces. First the **samples**: a `SingleTurnSample` holds the fields a metric may read — `user_input`, `retrieved_contexts`, `response`, and `reference` when the metric needs ground truth. Second the **dataset**: `EvaluationDataset(samples=[...])`, or `EvaluationDataset.from_list(list_of_dicts)` when the data already exists as records. Third the **run**: `evaluate(dataset=ds, metrics=[Faithfulness()], llm=evaluator_llm)`, which fans the sample × metric jobs out concurrently rather than looping in Python. It hands back an `EvaluationResult`; printing it shows one averaged number per metric, and `result.to_pandas()` gives one row per sample with a column per metric so you can look at individual failures. Most metrics call a judge model, so an evaluator LLM must be supplied — either on `evaluate()` or on each metric.
code
python · 19 linesfrom ragas import evaluate, EvaluationDataset
from ragas.dataset_schema import SingleTurnSample
from ragas.metrics import Faithfulness
from ragas.llms import LangchainLLMWrapper
from langchain_openai import ChatOpenAI
sample = SingleTurnSample(
user_input="When was the Eiffel Tower completed?",
retrieved_contexts=["The Eiffel Tower was completed in 1889."],
response="It was completed in 1889.",
reference="1889",
)
dataset = EvaluationDataset(samples=[sample])
evaluator_llm = LangchainLLMWrapper(ChatOpenAI(model="gpt-4o-mini"))
result = evaluate(dataset=dataset, metrics=[Faithfulness()], llm=evaluator_llm)
print(result)
df = result.to_pandas()go deeper
Be able to name the three pieces without notes: a sample per record, an EvaluationDataset holding them, and one evaluate() call with metrics and a judge model. Mention that to_pandas() shows per-sample scores.
Explain that evaluate() fans sample-by-metric jobs out concurrently rather than looping, that each metric declares the sample fields it requires, and that reference-requiring metrics are unusable when you have no ground truth.
Show where the run fits in a real workflow: capturing records from production or a golden set, choosing metrics against the fields you actually have, and reading the worst rows of the dataframe rather than reporting the mean.
Own the framing that Ragas scores captured records, not a live pipeline — so the value of any run is bounded by how representative the captured dataset is and by what you can afford to score at your traffic volume.
## What a Ragas run actually is Ragas does not score "a RAG pipeline". It scores **records that a pipeline already produced**. You run your application first, capture what it did for each question, and then hand that captured data to Ragas. That separation is why the API has three distinct objects: a sample (one record), a dataset (a homogeneous collection of records), and a run (`evaluate`) that returns a result object. ## The sample `SingleTurnSample` is a typed container for one question-and-answer interaction. Its fields map onto what RAG metrics need: - `user_input` — the question that was asked. - `retrieved_contexts` — the list of context strings your retriever returned. - `response` — what your generator produced. - `reference` — the ground-truth answer, present only when you have one. - `reference_contexts` — the contexts that *should* have been retrieved, for recall-style metrics. Not every field is needed for every metric; each metric declares which fields it reads, and a dataset missing a required field fails validation rather than silently scoring badly. For conversations there is `MultiTurnSample`, whose `user_input` is a list of message objects instead of a string; a dataset must be all one sample type, and only multi-turn-capable metrics accept it. ## The dataset `EvaluationDataset(samples=[sample1, sample2, ...])` is the direct constructor. When your data is already a list of dictionaries — the usual shape after pulling rows out of a warehouse, a JSONL file, or a dataframe converted with `to_dict("records")` — `EvaluationDataset.from_list(records)` builds the samples for you, and the dictionary keys must match the sample field names. The dataset is deliberately dumb: it holds records and knows its sample type. It contains no scoring logic and no configuration. ## The run `evaluate(dataset=ds, metrics=[...], llm=evaluator_llm)` is the single entry point. Two things about it surprise people the first time: **It is not a loop.** Internally Ragas builds one job per (sample, metric) pair and executes them concurrently through an executor, with a progress bar. A hundred samples and three metrics is three hundred jobs, not three hundred sequential HTTP round trips. Concurrency, per-call timeouts and retries are configured through a `RunConfig` passed to `evaluate`. **It needs a judge model.** Almost every metric in Ragas is LLM-based: it prompts a model to extract claims, check them against context, or produce a verdict. You supply that evaluator model either once on `evaluate(llm=...)` or individually on each metric object. Forgetting it is the most common first-run error. Some metrics also need an embedding model. ## The result `evaluate` returns an `EvaluationResult`. Printing it renders a dictionary of one averaged number per metric — the headline. `result.to_pandas()` returns a dataframe with one row per sample: the sample's own fields as columns, plus a column per metric holding that sample's individual score. That per-row view is where the useful information is; a single mean tells you a number moved but never which query broke. ## Why running it this way matters Calling `single_turn_ascore` on one sample at a time in a Python loop technically works and is how you debug a single case, but it gives up the executor, the shared run configuration, the progress reporting, the aggregate result object, and the dataframe. For anything beyond a handful of records, the batch call is the API you want. ## The shape of a real workflow Capture production or golden-set records → build `SingleTurnSample`s → wrap in `EvaluationDataset` → choose the metrics that match the fields you have (reference-requiring metrics are unusable on raw production traffic with no ground truth) → wire an evaluator model → `evaluate` → read `to_pandas()`, sort ascending by the metric column, and read the worst rows. The last step is the one candidates skip and interviewers ask about.
- What changes if the records are multi-turn conversations rather than single question-answer pairs?You build `MultiTurnSample` objects instead, whose `user_input` is a list of message objects rather than a single string. A dataset must be homogeneous — all single-turn or all multi-turn — and only metrics that declare multi-turn support will accept it. Mixing the two sample types in one `EvaluationDataset` is not a supported run.
- You already have the data in a pandas dataframe. How do you get it into an EvaluationDataset?Convert the dataframe to records with `df.to_dict("records")` and pass that list to `EvaluationDataset.from_list(...)`. The dictionary keys must match the sample field names exactly — `user_input`, `retrieved_contexts`, `response`, `reference` — and `retrieved_contexts` must be a list of strings, not a single concatenated blob, or context-level metrics score nonsense.
- When would you call a metric on one sample directly instead of using evaluate()?When debugging a single case — you want the score for one record and the metric's reasoning without spinning up a run. `single_turn_ascore` on the metric object does exactly that. For anything at scale you lose the concurrency executor, shared run configuration, progress reporting, the aggregate result object and the dataframe, so batch evaluation is the default.
Think of it like a test runner: the samples are the test cases, the metrics are the assertions, and evaluate() is the runner that executes every case-assertion pair and returns a report.
saying these in an interview costs you the question
- Scoring rows one by one in a Python loop instead of evaluate()
- Passing a bare list of dicts straight into evaluate()
- Assuming evaluate() runs samples sequentially
- Forgetting that most metrics need an evaluator LLM configured
- Thinking the returned result is a plain dict of averages