In DeepEval, is a high HallucinationMetric score good or bad?
answer
- two families, opposite directions
- zero is the ideal for one family
- threshold is a tolerance, not a bar
- same default hides the inversion
- a green safety metric may prove nothing
basics
~20 sBad. HallucinationMetric and ToxicityMetric are lower-is-better: the score measures how much hallucination or toxicity was found, and the case passes when the score is at or below the threshold — the inverse of AnswerRelevancyMetric or FaithfulnessMetric.
solid answer
~40 sMost DeepEval metrics are higher-is-better: `AnswerRelevancyMetric`, `FaithfulnessMetric` and the contextual metrics all score quality, so a case succeeds when the score is at or above the threshold. `HallucinationMetric`, `ToxicityMetric` and `BiasMetric` invert that — they score the amount of a bad thing, so 0.0 is the ideal result and success means the score is at or **below** the threshold. The reason this matters is that both families take a `threshold` argument with the same default of 0.5, so nothing in the constructor signals the direction. A team that sets `HallucinationMetric(threshold=0.8)` intending "be strict" has actually authorised up to 80% hallucination. Read the direction per metric family and, in review, treat any threshold copied between a quality metric and a safety metric as a bug.
code
python · 12 linesfrom deepeval.metrics import AnswerRelevancyMetric, ToxicityMetric
from deepeval.test_case import LLMTestCase
case = LLMTestCase(input="Summarise the policy.", actual_output="...")
quality = AnswerRelevancyMetric(threshold=0.8) # passes when score >= 0.8
safety = ToxicityMetric(threshold=0.0) # passes when score <= 0.0
quality.measure(case)
safety.measure(case)
print(quality.score, quality.is_successful())
print(safety.score, safety.is_successful())go deeper
Remember that hallucination, toxicity and bias scores measure the bad thing, so 0.0 is the good result, while relevancy and faithfulness scores measure quality and 1.0 is the good result.
Explain that success compares score to threshold in opposite directions for the two families, and that the shared default of 0.5 gives no hint which one you are in.
Show the operational consequence: a permissive safety threshold produces a permanently green check that proves nothing, so treat any never-failing safety metric as suspect until you have seen it fail on a planted bad case.
Own thresholds as policy rather than configuration — safety tolerances near zero, quality bars set with measured margin, and a review rule that a threshold may never be copied between the two families.
## Two families, one argument name DeepEval's built-ins split into two groups that behave in opposite directions but share an identical-looking constructor. **Higher is better (quality metrics).** `AnswerRelevancyMetric`, `FaithfulnessMetric`, `ContextualPrecisionMetric`, `ContextualRecallMetric`, `ContextualRelevancyMetric`. The score is a proportion of goodness — how much of the answer was relevant, how many claims were supported, how much of the retrieved context was on topic. 1.0 is perfect, and the case succeeds when the score is at or above the threshold. **Lower is better (safety / error metrics).** `HallucinationMetric`, `ToxicityMetric`, `BiasMetric`. The score is a proportion of badness — how much of the output contradicts the supplied context, how much of it is toxic or biased. 0.0 is perfect, and the case succeeds when the score is at or below the threshold. Both families expose `threshold`, both default it to 0.5, and both report through `is_successful()`. Nothing in the API surface tells you which direction you are in; you have to know the metric. ## The mistake this produces The failure is not exotic, it is the natural one. Someone has internalised "higher threshold means stricter" from the relevancy metrics and applies it uniformly. Writing `AnswerRelevancyMetric(threshold=0.9)` genuinely is strict — it demands 90% relevance. Writing `ToxicityMetric(threshold=0.9)` is the opposite: it tolerates a toxicity score up to 0.9 before failing anything. The second line reads like tightening and is in fact almost complete permissiveness. Worse, it produces no visible symptom: the suite goes green, and it goes green precisely on the cases you built the metric to catch. A safety metric that never fires looks exactly like a system that is behaving. The mirror-image mistake is setting a hallucination threshold very low as a *loosening* gesture and then wondering why every case fails. ## How to reason about the number For the lower-is-better metrics, read the threshold as a **tolerance**, not a bar. `ToxicityMetric(threshold=0.0)` means "any detected toxicity fails this case", which for a user-facing product is often the right setting and is a perfectly defensible answer in an interview. `HallucinationMetric(threshold=0.2)` means "up to a fifth of the supplied contexts may be contradicted before I call this a failure" — sometimes reasonable on noisy multi-document cases, rarely reasonable on a single authoritative context. For the higher-is-better metrics, read the threshold as a **bar**, and set it below the score you actually observe rather than at it, because the judge is an LLM and its output moves between runs. A threshold pinned to a single observed score turns normal judge variance into random CI failures. ## Reading the score itself Because the score is a proportion in both families, it carries information beyond pass/fail. A hallucination score of 0.5 on a case with two contexts means one of them was contradicted — that is a much more actionable statement than "failed". Turning on the reason string makes that concrete, giving you the judge's account of which claim conflicted with which context. When you are triaging a failing safety case, the reason plus the raw proportion usually locates the problem faster than re-running with a different threshold. ## What to say in the interview Three beats: name the two families, state that threshold semantics invert between them, and give the concrete consequence — that a threshold copied from a quality metric to a safety metric silently disables it. If you can add that safety thresholds are usually set at or very near 0.0 while quality thresholds are set with margin below observed scores, you have covered both the mechanic and the judgment.
- What does ToxicityMetric(threshold=0.9) actually do to a suite?It effectively disables the check. Because toxicity is lower-is-better, the case only fails when the score exceeds 0.9, so almost any output passes. It is the most common form of this bug because 0.9 reads as strict by analogy with the relevancy metrics, and the suite goes green rather than erroring.
- What is a sensible default threshold for the safety metrics?Usually 0.0 or very close to it for user-facing output: any detected toxicity or bias should fail the case and be looked at. Hallucination sometimes warrants a small tolerance on cases with many supplied contexts, where one contradicted document out of ten is noise, but on a single authoritative context 0.0 is the honest setting.
- Why shouldn't you set a quality threshold exactly at the score you observed once?Because the judge is an LLM and its scores move slightly between runs. A threshold pinned to a single observation converts normal variance into random red builds. Set it below the observed level with a margin, sized from the spread you see across a few repeat runs of the same suite.
saying these in an interview costs you the question
- Assuming a higher threshold is always stricter
- Reading a high HallucinationMetric score as a good result
- Copying a relevancy threshold onto a toxicity metric
- Thinking DeepEval warns you about an inverted threshold
- Believing 0.5 is a safe default for a safety metric