In Ragas FactualCorrectness, what do mode='precision' and mode='recall' change?
answer
- two directions of claim comparison
- denominator tells you the mode
- extras versus omissions
- f1 hides which side broke
- decomposition settings move the baseline
basics
~20 sFactualCorrectness breaks both the response and the reference into claims. Precision mode scores what fraction of the response's claims the reference supports, punishing invented extras. Recall mode scores what fraction of the reference's claims the response contains, punishing omissions. The default f1 combines both.
solid answer
~40 s`FactualCorrectness` decomposes the `response` and the `reference` into atomic claims and runs entailment checks between the two sets, so the `mode` argument picks which direction of that comparison becomes the score. With `mode="precision"` the denominator is the response's claims: a model that adds true-but-unreferenced detail is penalised, which is what you want when spurious additions are the risk. With `mode="recall"` the denominator is the reference's claims: a model that answers correctly but leaves things out is penalised, which is what you want when completeness is the requirement. The default, `mode="f1"`, is the harmonic mean and hides which of the two failed. The `atomicity` and `coverage` arguments tune how finely the claim decomposition splits the text, which changes the denominators and therefore the absolute score. Because it reads `reference`, this is an offline metric only.
code
python · 9 linesfrom langchain_openai import ChatOpenAI
from ragas.llms import LangchainLLMWrapper
from ragas.metrics import FactualCorrectness
evaluator_llm = LangchainLLMWrapper(ChatOpenAI(model="gpt-4o-mini"))
precision = FactualCorrectness(llm=evaluator_llm, mode="precision")
recall = FactualCorrectness(llm=evaluator_llm, mode="recall")
combined = FactualCorrectness(llm=evaluator_llm) # mode="f1" by defaultgo deeper
Know that FactualCorrectness compares the response to a written reference answer by breaking both into claims, and that it therefore needs ground truth to run at all.
Be ready to state which side each mode puts in the denominator — precision over the response's claims, recall over the reference's — and to pick the mode that matches the failure you care about.
Show that you report both directions rather than f1 alone, because extras and omissions have opposite fixes, and that you freeze the decomposition settings so the number stays comparable across releases.
Own which direction the product should be optimised for. In a regulated domain unsupported additions are the liability and precision is the number that gets governed; in a completeness-critical workflow it is recall, and that choice belongs in the quality policy, not in a config file.
## What the metric does before it scores anything `FactualCorrectness` is not a similarity measure. It first asks the judge LLM to break the `response` into atomic claims, then to break the `reference` into atomic claims, and then it runs natural-language-inference checks between the two sets: is each response claim entailed by the reference, and is each reference claim entailed by the response. Those two directions produce two different counts, and `mode` selects which one you see. ## Precision mode The denominator is the number of claims in the response. The numerator is how many of them the reference supports. A response that says three things, two of which appear in the reference, scores two-thirds regardless of how much the reference also contained. The behaviour this creates is specific and often surprising: a model that produces a correct answer plus extra correct detail that simply is not in your reference gets marked down. That is not a bug, it is the metric doing its job — in a domain where unsupported additions are dangerous (medical, legal, financial), an extra claim you did not verify is exactly the thing to punish, and precision mode is how you make the number reflect that. ## Recall mode The denominator flips to the number of claims in the reference. The numerator is how many of them appear in the response. Now extras are free and omissions are expensive. This is the mode for questions whose answer has several required parts — a checklist, the steps of a procedure, all the eligibility criteria — where a fluent answer that mentions two of four requirements is a real failure that precision mode would score perfectly. ## Why f1 as the default is a reporting problem `mode="f1"` is the harmonic mean of the two and is the default. It is a reasonable single number and a poor diagnostic: a 0.6 could be a terse answer that omitted half the reference, or a verbose answer full of unverified additions, and those two failures have opposite fixes. Prompt for brevity and you make the second better and the first worse. Teams that only track f1 tend to oscillate. Report both directions during development and keep f1 for the summary row. ## Atomicity and coverage The `atomicity` and `coverage` arguments control the decomposition step — how aggressively a sentence is split into separate claims and how much of the source text the claims are required to account for. They matter because they move the denominators. Finer decomposition produces more, smaller claims, and the score changes even though neither your model nor your data did. The operational rule that follows: fix these settings when you establish a baseline and treat a change to them as a break in the time series, exactly as you would treat a change of judge model. A score comparison across different decomposition settings is not a comparison. ## Where it fits FactualCorrectness reads `reference`, so it lives entirely in the labelled offline suite — it cannot run on unlabelled production traffic. It is also one of the more expensive built-ins per sample, because a single score requires two decomposition calls plus entailment checks over the resulting claim sets, and both sets scale with answer length. Long-form answers cost noticeably more than short factual ones, which is worth knowing before you point it at a dataset of essay-length outputs. ## Reading a result In review, the useful framing is a sentence per direction. "Precision fell from 0.91 to 0.78 while recall held, so the new prompt is adding claims we cannot trace to the reference" is a diagnosis. "FactualCorrectness fell from 0.88 to 0.81" is a fire alarm with no address on it.
- A model's precision score falls while recall holds steady. What changed?The response is asserting claims the reference does not support — it got longer, more speculative, or started adding detail from parametric knowledge rather than the source. Recall staying flat says it did not drop anything it used to say. The fix is on the generation side: tighten the prompt against unsupported elaboration, or check whether a model swap traded terseness for verbosity.
- Why must the atomicity and coverage settings stay fixed across runs?They govern how text is split into claims, which sets the denominator on both sides. Finer decomposition yields more, smaller claims and shifts the score even with identical model output and identical data. Changing them mid-programme silently rebaselines your metric, so treat a change as starting a new time series — the same discipline you apply to swapping the judge model.
- Can FactualCorrectness be used to monitor live production traffic?No. It compares the response against `reference`, and unlabelled production traces have no reference. It belongs to the curated labelled suite you run before a release. On live traffic the reference-free options are grounding and relevancy metrics, or a custom binary critic — none of which measures correctness, only consistency with what was retrieved.
saying these in an interview costs you the question
- Thinking FactualCorrectness is a text similarity score
- Reporting only f1 and losing the diagnosis
- Calling extra correct detail a metric bug
- Changing decomposition settings mid-baseline
- Trying to run it on unlabelled production traces