skip to content

A Ragas run finishes but some metric scores are NaN — what happened, and what do you do?

level: seniorimportance: should knowfreq 40%

answer

  1. failures are swallowed, not raised
  2. a gap in the column, not an error
  3. averages cover only the survivors
  4. the failing rows are not random
  5. count them before you read the mean

basics

~20 s

By default a Ragas run does not raise on a failing metric call — the exception is swallowed and that sample's score is recorded as NaN. Causes are exhausted retries, timeouts, provider errors, or unparseable judge output. Count the NaNs before trusting the averages.

solid answer

~50 s

`evaluate()` runs with `raise_exceptions=False` by default, which means a scoring job that fails after its retries does not abort the run — the sample gets a missing score and everything else continues. The usual causes are a call that exhausted `max_retries`, a call that blew the per-call `timeout`, a provider error such as a content filter or context-length rejection on a long context, or a judge whose output could not be parsed into the structure the metric expects. The danger is that the printed averages are computed over the rows that *did* score, so a run where a third of the samples silently failed still prints a plausible number. The workflow is: call `to_pandas()`, check `isna().sum()` per metric column before reading any average, then re-run the failing subset with `raise_exceptions=True` to see the real error, and fix the cause — usually concurrency, timeout, or oversized contexts.

go deeper

for a junior

Know that a Ragas run does not stop when one metric call fails — the sample gets a missing score instead — so a completed run is not automatically a fully scored run.

for a middle

Explain the causes: exhausted retries, per-call timeouts, provider rejections on huge contexts, and judge output the metric could not parse. Say that averages are computed over the rows that did score.

for a senior

Show the diagnostic sequence — count missing scores per metric column, look for what the failing rows have in common, reproduce a small subset with exceptions raised, then fix the concurrency, timeout or context size that caused it.

for a principal

Own the reporting rule: every automated evaluation number ships with its scored-row count, because a silently shrinking survivor set can hold a dashboard flat while the pipeline degrades.

## What a NaN actually means Ragas treats an evaluation as a batch of independent jobs, and it is deliberately tolerant of individual job failure. `evaluate(..., raise_exceptions=False)` is the default: when a scoring job raises after its retry budget is spent, the framework records no score for that sample-and-metric pair and moves on. In the dataframe from `to_pandas()`, that shows up as NaN in the metric column. This is a reasonable default — one poisoned row should not destroy a forty-minute run — but it converts a loud failure into a quiet one, and quiet failures in evaluation are dangerous because the output still looks like a result. ## The realistic causes, roughly in order **Retries exhausted.** The provider was throttling or erroring for long enough that backoff ran out. Symptom: NaNs cluster in time, and often across every metric at once, because the whole executor was hitting the same wall. **Per-call timeout.** A slow judge, or a metric issuing several sequential calls, ran past the deadline. Symptom: NaNs concentrate on the samples with the longest contexts or the most expensive metrics. **Provider-side rejection.** Context-length errors from a sample whose `retrieved_contexts` are enormous, or a safety filter refusing to process the content. Symptom: NaNs correlate with a content property, not with time. **Unparseable judge output.** LLM metrics ask the judge for a structured answer and then parse it. A judge that returns prose, truncates, or emits malformed JSON produces a parse failure. Symptom: a weaker or smaller judge model has far more NaNs than a stronger one for the same dataset — this is the classic reason a cost-saving judge swap goes wrong. ## Why the average lies Aggregation is over the samples that produced a score. If 300 of 1,000 rows failed, the headline is the mean of the 700 that survived — and those 700 are not a random sample. If the failures were context-length rejections, you have quietly excluded exactly the hardest, longest-context queries and made the system look better than it is. This selection effect is the reason NaN counting comes before interpretation, always. ## The diagnostic sequence 1. `df = result.to_pandas()`, then `df.isna().sum()`. Any non-zero count invalidates the headline until explained. 2. Look at *which* rows failed. Are they the longest ones? All from one time window? All on one metric? 3. Re-run a small failing subset with `raise_exceptions=True`. The run now stops on the first real exception and you see the actual error rather than a silent gap. 4. Fix by cause: reduce concurrency for throttling, raise the timeout for a slow judge, truncate or chunk oversized contexts for length rejections, use a stronger or better-constrained judge for parse failures. 5. Re-run and confirm the NaN count is zero before comparing to anything. ## Distinguish this from an upfront validation error A different and louder failure happens *before* any scoring: a metric declares the sample fields it requires, and a dataset that does not carry them fails validation rather than producing NaNs. Handing reference-free production traffic to a metric that requires `reference` is that case — you get an explicit error about the missing field, not silent gaps. Knowing which of the two you are looking at tells you whether the problem is your dataset shape or your run configuration. ## Reporting discipline Any automated report over a Ragas run should carry the scored-row count next to the score, in the same way a survey reports its response rate. "Faithfulness 0.86" and "faithfulness 0.86 over 702 of 1,000 samples" are different claims, and only the second one can be safely compared to last week's run. If a run cannot be repaired in time, reporting it with the completeness figure is honest; reporting the mean alone is not. ## The failure mode to name in an interview The worst version of this is a scheduled evaluation whose judge provider degrades, whose rows quietly stop scoring, and whose dashboard keeps drawing a flat, healthy line computed over a shrinking survivor set. A completeness check is the guard, and it belongs in the same job that computes the score.

  • Why is a run with 30% NaN scores worse than a run that crashed outright?
    A crash is a signal; a partially scored run produces a plausible number that gets reported and compared. The surviving rows are also not a random sample — if the failures were context-length rejections or timeouts, you have systematically dropped the hardest queries, so the average is biased upward by exactly the cases you most wanted to measure.
  • You swapped to a smaller, cheaper judge model and NaNs appeared. What is the likely mechanism?
    Parse failures. LLM metrics prompt the judge for a structured verdict and then parse it; a weaker model is more likely to return prose, truncate, or emit malformed JSON, so the metric raises and the sample goes unscored. The cost saving is illusory if it costs you scored rows — measure the completeness of the cheaper judge before adopting it.
  • When would you deliberately set raise_exceptions=True?
    When debugging, and when the run's completeness matters more than its completion — for example a gating run whose output should never be a partially scored average. It stops on the first genuine exception so you see the real error instead of a gap. For long exploratory runs the default tolerance is better, provided you check the missing-score count afterwards.

saying these in an interview costs you the question

  • Reading the average without checking how many rows scored
  • Assuming a run that completed scored every sample
  • Treating the failed rows as a random subset of the dataset
  • Blaming the model for a low score that is really a parse failure
  • Ignoring that a cheaper judge can raise the unscored-row count

context