In Ragas, when do you use AspectCritic instead of RubricsScore?
answer
- binary verdict versus graded band
- what will you do with the number?
- gates want pass rates
- band descriptions must be qualitatively distinct
- majority vote against judge flakiness
basics
~20 sAspectCritic returns a binary 0 or 1 verdict against a natural-language definition, so it suits pass/fail gates and safety checks. RubricsScore returns a graded score against named band descriptions, so it suits tracking gradual quality change where a hard boundary would be arbitrary.
solid answer
~60 sBoth are the open-ended metrics in Ragas — you supply the criterion instead of using a built-in RAG definition. `AspectCritic` takes a `name` and a natural-language `definition` and returns a binary verdict: 1 if the sample satisfies the definition, 0 otherwise. That makes it the right tool for properties with a genuine yes/no answer — did the response refuse, does it contain PII, did it cite a source, is it harmful — and it is what you gate a deploy on, because a pass rate is a number you can threshold honestly. `RubricsScore` takes a `rubrics` mapping of band descriptions (`score1_description` through `score5_description`) and asks the judge to place the sample on that scale. Use it when quality is genuinely ordinal and you want to see movement rather than a cliff. The tradeoff: binary verdicts are stable and cheap to reason about but throw away magnitude; graded scores keep magnitude but drift, because a judge's idea of a 3 versus a 4 is far less reproducible than its idea of yes versus no.
code
python · 21 linesfrom langchain_openai import ChatOpenAI
from ragas.llms import LangchainLLMWrapper
from ragas.metrics import AspectCritic, RubricsScore
evaluator_llm = LangchainLLMWrapper(ChatOpenAI(model="gpt-4o-mini"))
unsupported = AspectCritic(
name="unsupported_claim",
definition="Does the response make any claim that is not supported by the retrieved context?",
llm=evaluator_llm,
)
rubrics = {
"score1_description": "The response contradicts the retrieved context.",
"score2_description": "The response is mostly unsupported by the retrieved context.",
"score3_description": "The response is supported but omits a key fact.",
"score4_description": "The response is supported and complete.",
"score5_description": "The response is supported, complete and cites its source.",
}
quality = RubricsScore(rubrics=rubrics, llm=evaluator_llm)go deeper
Know that both metrics let you define your own criterion in plain English, and that AspectCritic gives a yes/no result while RubricsScore gives a graded one.
Be ready to construct both — a name plus definition for AspectCritic, a mapping of score band descriptions for RubricsScore — and to justify which output shape fits the question being asked.
Demonstrate the operational split: binary critics as release gates because a pass rate is thresholdable, a rubric alongside as a trend instrument, and criteria worded so they run on unlabelled traffic.
Own how many custom criteria the organisation maintains. Every critic is a definition someone must keep true as the product changes, so argue for a small set of non-negotiable gates rather than a sprawling scorecard nobody re-reads.
## Two shapes of custom criterion Every other metric class in Ragas encodes a fixed definition — faithfulness, context recall, factual correctness. `AspectCritic` and `RubricsScore` are the two escape hatches where the criterion is yours. They differ only in the shape of the output, and that shape drives almost every practical decision about them. ## AspectCritic: a binary verdict You construct it with a `name` (which becomes the column the score appears under) and a `definition`, a plain-English yes/no question about the sample. The judge answers it and the metric returns 0 or 1. The definition has to be genuinely binary to work well. "Does the response contain any statement not supported by the retrieved context?" is binary. "Is the response high quality?" is not — the judge will still answer, but its threshold moves between runs and you have built a graded metric wearing a binary costume. AspectCritic also exposes `strictness`, which runs the judgement multiple times and takes the majority verdict. An odd value avoids ties. This is the direct lever against judge flakiness: a criterion that flips between runs on the same sample gets more stable at higher strictness, at proportionally more judge calls per sample. What you get out is a per-sample 0 or 1, which aggregates over a dataset into a pass rate. A pass rate is easy to gate on and easy to explain in a review: "98.2% of samples had no unsupported claim, down from 99.4%" is an actionable sentence. It is also easy to drill into, because the failures are a discrete list of samples, not a tail of the distribution. ## RubricsScore: a graded band You construct it with a `rubrics` dictionary whose keys are `score1_description` through `score5_description` and whose values describe what a sample at that level looks like. The judge picks a band. The quality of the metric is entirely the quality of those descriptions. Bands that are qualitatively distinct — "contradicts the source", "omits the key fact", "correct but incomplete", "correct and complete", "correct, complete and well-cited" — produce judgements you can reproduce. Bands that differ only by adverbs — "somewhat good", "quite good", "very good" — produce a metric whose mean drifts whenever you change judge model, and you will not be able to tell drift from a real regression. Rubrics can also be attached per sample through the sample's `rubrics` field, which is the mechanism when different questions in a dataset deserve different grading scales. ## Choosing between them Ask what you will do with the number. If the answer is *block a release*, use AspectCritic. Thresholding a mean of an ordinal scale is a category error: the difference between a mean of 3.8 and 3.6 has no defensible interpretation, and everyone in the room knows it, so the gate gets overridden the first time it fires. If the answer is *watch a trend during development*, RubricsScore is the more informative instrument, because it distinguishes "slightly worse" from "catastrophically worse" where a binary just flips. The usual mature setup runs a handful of AspectCritics as hard gates on the properties that are non-negotiable — safety, citation presence, refusal correctness — plus one RubricsScore for overall answer quality that nobody gates on but everybody looks at. ## Cost and failure modes Both are one judge call per sample at default settings, which makes them cheaper than the multi-step decomposition metrics. AspectCritic with `strictness` above 1 multiplies that. The shared failure mode is a criterion that reaches for information the sample does not carry. If your definition or your band descriptions talk about the correct answer, the metric needs `reference` populated, and it will be useless on production traces. Write criteria that judge the response against `retrieved_contexts` if you want them to run on live traffic.
- Why is gating a deploy on the mean of a 1-5 rubric score a bad idea?Because the scale is ordinal, not interval — the gap between a 3 and a 4 is not the same quantity as between a 4 and a 5, so the mean has no defensible unit. A threshold on it cannot be justified when it fires, and judge-side drift moves the mean without any change in your application. Gate on a binary pass rate and keep the rubric as a trend line.
- What does raising strictness on an AspectCritic actually do?It repeats the judgement and takes the majority verdict, so an odd value avoids ties. It buys stability on criteria where the judge flips between runs on identical input, and it costs proportionally more calls per sample. If a criterion needs high strictness to stabilise, that is usually a signal the definition is ambiguous and should be rewritten rather than voted on.
- How would you write an AspectCritic definition that still works on unlabelled production traffic?Word it so it only refers to fields a live trace carries — the question, the retrieved chunks and the response. "Does the response make any claim not supported by the retrieved context?" runs anywhere. "Does the response match the correct answer?" silently requires a reference and is useless online. The wording of the definition is what decides whether the metric is reference-free.
saying these in an interview costs you the question
- Writing a vague non-binary definition into AspectCritic
- Rubric bands that differ only by adverbs
- Gating releases on a mean rubric score
- Assuming AspectCritic returns a probability rather than 0 or 1
- Criteria that mention the correct answer on unlabelled traffic