In Ragas, what does EvaluationResult.to_pandas() return after an evaluation run?
answer
- two resolutions, not one number
- the mean hides the distribution
- one row per sample, one column per metric
- inputs sit next to their scores
- sort ascending, read the worst rows
basics
~20 sto_pandas() returns a dataframe with one row per evaluated sample: the sample's own fields as columns plus one column per metric holding that sample's individual score. Printing the result object instead shows only the per-metric average across samples.
solid answer
~40 sAn `EvaluationResult` carries two views of the same run. The **aggregate** view is what `print(result)` renders: a dictionary mapping each metric name to its score averaged over the samples. The **per-sample** view is `result.to_pandas()`: a dataframe whose rows are the samples, whose columns include `user_input`, `retrieved_contexts`, `response` and `reference`, and which gains one extra column per metric with that row's score. `result.scores` exposes the same per-sample scores as a list of dicts if you would rather not use pandas. The practical habit is to never stop at the aggregate — sort the dataframe ascending by a metric column and read the bottom rows, because a mean of 0.82 can equally mean "uniformly mediocre" or "excellent except for eleven catastrophic queries", and those two situations need completely different fixes.
go deeper
Know that the run returns an object with two views: a printed average per metric, and to_pandas() giving one row per sample with a score column per metric. Say that you look at the rows, not just the average.
Explain what the columns are and why an average hides the distribution — uniform mediocrity, a bimodal split, and a partially failed run can all produce the same mean, and only the per-row view separates them.
Demonstrate the diagnostic workflow: NaN counts first, then sort ascending on the metric column, then cross-tabulate context metrics against answer metrics to place the blame on retrieval or generation, then slice by segment.
Own the artifact policy — per-row results persisted with model, prompt, dataset and metric versions, so runs are comparable over months and a regression can be diffed row by row instead of argued about from two averages.
## Two views of one run When `evaluate` finishes it returns an `EvaluationResult` object, not a number and not a dataframe. That object exposes the run at two different resolutions, and knowing both is the whole point of this question. **The aggregate.** Printing the result renders a mapping of metric name to a single score, averaged across every sample that produced a score. This is the number that ends up in a slide, a PR comment, or a dashboard. **The per-sample detail.** `result.to_pandas()` materialises a dataframe. Each row is one sample. The columns are the sample's own fields — `user_input`, `retrieved_contexts`, `response`, `reference` and any other populated field — plus one column per metric, named after the metric, containing that sample's score. `result.scores` gives you the same per-sample numbers as a list of dictionaries if pandas is not in play. ## Why the aggregate alone is a trap Most RAG metrics produce a value in the 0–1 range, and averaging them destroys the distribution. Three genuinely different failure shapes all average to roughly the same place: - **Uniform mediocrity** — every row scores around 0.8. Usually a prompt or a chunking problem affecting everything. - **Bimodal** — most rows are 1.0 and a minority are 0.0. Usually a specific class of query the retriever cannot serve at all. - **Silent partial failure** — many rows never scored (they hold NaN) and the printed average was computed over the rest. The number looks fine because it describes a subset you did not choose. The dataframe distinguishes them in seconds and the aggregate never can. `df["faithfulness"].isna().sum()` catches the third; a histogram or a `value_counts` catches the second; sorting ascending and reading the bottom twenty rows tells you what the first actually is. ## Reading it like an engineer The dataframe keeps the inputs alongside the scores on purpose. A low faithfulness score in isolation is just a number; a low faithfulness score sitting next to the exact `retrieved_contexts` and `response` that produced it is a bug report. The usual moves: - Sort ascending on the metric column and read the worst rows end-to-end. - Cross-tabulate two metric columns. Rows with good context metrics but a poor answer metric point at generation; rows with poor context metrics point at retrieval. That split is the reason people reach for this library in the first place. - Join a category or tenant column you carried through on the sample and group by it, so a regression concentrated in one segment is visible rather than diluted. - Persist the dataframe with the run's identifying details — model, prompt version, dataset revision, date — because a score is meaningless without knowing what produced it, and a stored per-row artifact lets you diff two runs row by row instead of comparing two means. ## What the result object is not It is not a live handle on the pipeline; the run is over and the object holds recorded numbers. It is not a plain dictionary, so code that treats it like one and indexes it blindly is relying on convenience behaviour rather than the documented surface — `to_pandas()` and `scores` are the two accessors to reach for. And it does not by itself tell you whether a difference between two runs is meaningful; deciding that is statistics, not an API call, and it depends on sample size and judge variance rather than on anything the dataframe can show you. ## The interview answer in one line The aggregate is the headline, the dataframe is the evidence, and an engineer who only ever reports the headline cannot tell you which queries broke or whether half the run failed to score at all.
- Your run prints an average faithfulness of 0.86 but a colleague says the run is broken. What would you check in the dataframe first?The count of NaN values in the metric column. Ragas records a sample whose metric call errored as NaN rather than failing the run, and the printed average is computed over the samples that did score. A run where a third of the rows never scored can still print a healthy-looking mean, so `isna().sum()` per metric column comes before any interpretation of the number.
- How would you use the dataframe to separate a retrieval problem from a generation problem?Score both a context-quality metric and an answer-quality metric in the same run, then cross-tabulate their columns. Rows where the context metrics are strong but the answer metric is weak indict generation — the model had what it needed and still went wrong. Rows where the context metrics are weak indict retrieval, and the answer metric there is largely uninformative.
- What should you store alongside the dataframe for a run to stay useful next month?The identifying context the numbers depend on: the generator and judge model names and versions, the prompt version, the dataset revision, the metric configuration, and the timestamp. Without those a stored score cannot be compared to a later one, because you cannot tell whether a difference came from your change, a different judge, or a different dataset.
saying these in an interview costs you the question
- Reporting only the aggregate and never opening the per-sample rows
- Assuming every row in the dataframe actually produced a score
- Treating the result object as a plain dictionary of numbers
- Reading a low answer score without looking at the retrieved contexts beside it
- Comparing two runs by their means with no record of what produced each