What does strict_mode=True do to a DeepEval GEval metric's score?
answer
- all-or-nothing scoring
- the threshold you set stops mattering
- fine for guardrails, hostile to quality gradients
- perfection is a narrow target
- judge wobble becomes a binary flip
basics
~20 sstrict_mode turns the metric into a binary one: the score becomes 1 for a perfect judgement and 0 otherwise, and the threshold is overridden to 1. Partial credit disappears, so anything less than flawless fails.
solid answer
~50 sBy default a `GEval` metric returns a continuous score between 0 and 1 and passes when it clears `threshold` (0.5 unless you set it). Setting `strict_mode=True` collapses that: the metric reports 1 only for a perfect judgement and 0 for everything else, and it overrides whatever threshold you configured, setting it to 1. Any imperfection is a failure. That is the right shape for a genuine guardrail — a metric asserting "no personally identifying data appears in the answer" has no meaningful 0.6. It is the wrong shape for a quality metric like helpfulness or tone, where the interesting signal lives in the gradient, and where binarising throws away the information you would use to see a change that is real but small. The operational caveat is that a strict metric inherits all of the judge's variance with none of the slack. One borderline case that the judge scores harshly on a rerun flips the whole metric red.
go deeper
Remember that strict_mode makes a DeepEval metric all-or-nothing: score 1 for perfect, 0 otherwise, with the threshold forced to 1. Without it the score is a continuous value between 0 and 1.
Explain that strict_mode overrides whatever threshold you configured, and say which criteria deserve binary treatment — hard rules like no-PII or schema validity — versus gradient criteria like tone where you lose sensitivity and triage order.
Talk about what binarising does to a running suite: judge variance turns into flips rather than a wobble, so a strict metric on an ambiguous criterion becomes a flaky blocker. Argue for a deterministic coded metric when the rule is mechanical.
Decide which metrics in the portfolio are guardrails that may block and which are trend indicators that may only report, and hold the line that quality gradients are never expressed as strict binaries just to make a dashboard look decisive.
## Default scoring `GEval` produces a score in the range 0 to 1 along with a natural-language `reason`. The metric is successful when the score meets its `threshold`, which defaults to 0.5, and `is_successful()` reports that verdict. This continuous shape is what lets you track "our correctness metric averaged 0.81 this week, 0.78 last week" and what lets a threshold express a tolerance. ## What strict_mode changes Setting `strict_mode=True` on the metric makes it binary. The score becomes 1 when the judge finds the output perfect against the rubric, and 0 in every other case, and the metric's threshold is overridden to 1 — so passing requires the perfect verdict. The `reason` still explains the verdict, which matters, because with a binary score the reason is the only diagnostic you have left. ```python safety = GEval( name="NoPII", evaluation_steps=[ "Scan 'actual output' for names, addresses, phone numbers, emails or account identifiers.", "The output is acceptable only if it contains none of them.", ], evaluation_params=[LLMTestCaseParams.ACTUAL_OUTPUT], strict_mode=True, ) ``` ## When binary is the honest shape Some criteria genuinely have no middle. "Did it leak a customer's phone number" is yes or no. "Did it output valid JSON conforming to the schema" is yes or no. "Did it refuse the request it was required to refuse" is yes or no. Encoding those as a 0-to-1 score with a 0.5 threshold is a lie about the underlying question, and it invites someone to "fix" a failure by nudging the threshold down. For these, `strict_mode=True` states the intent in the metric definition itself, so nobody has to reason about what 0.7 meant. ## When binary is the wrong shape Quality criteria — helpfulness, tone, completeness, style adherence — are gradients, and the gradient is the product. Binarising them costs you three things: 1. **Sensitivity.** A prompt change that lifts average helpfulness from 0.62 to 0.71 is invisible if every case is already reported as 0. 2. **Triage.** With continuous scores you sort by score and read the worst ten. With binary scores you have a pile of zeros in no order. 3. **Stability.** Perfection is the narrowest possible target, so a strict quality metric fails constantly, and a metric that is always red is a metric people stop reading. ## The variance problem A judge model's verdict is sampled, so borderline cases can land differently on repeat runs. Under continuous scoring that shows up as a small wobble in the average — annoying but absorbable. Under `strict_mode` the same wobble is a binary flip, and if the metric is wired into a pipeline that halts on failure, one flipped case halts it. That is why strict metrics work best when the criterion is one the judge can decide unambiguously. "Contains an email address" is decidable. "Fully answers the user's question" is not, and dressing it in `strict_mode` does not make it so — it just converts fuzziness into flakiness. If a criterion is truly a hard rule and you can express it in code — a regex, a schema validation, a lookup — a deterministic custom metric beats a strict judged one outright, because it has no variance at all. ## The sibling flag: verbose_mode `verbose_mode=True` is unrelated to scoring. It prints the metric's intermediate steps and reasoning to the console as it measures, which is how you inspect what a judge actually did on a case that scored surprisingly. It is a development affordance; leaving it on in a large suite buries the run output in noise, so it is normal to switch it on for a handful of cases and off again. ## Summary - Default: continuous 0-1 score, `threshold` decides pass or fail. - `strict_mode=True`: score is 1 or 0, threshold forced to 1, only perfection passes. - Use it for hard rules; keep quality criteria continuous. - Expect flips on borderline cases, and prefer a coded metric when the rule is truly mechanical.
- If strict_mode overrides threshold, is there any point setting both?No — setting threshold alongside strict_mode is misleading, because the strict flag forces the threshold to 1 regardless of what you wrote. A reader of the config sees a 0.8 that never applies. Pick one: either a continuous metric with a deliberate threshold, or a strict binary metric with no threshold argument at all.
- Your strict NoPII metric flips red on one case per run. What do you do?Read the reason on the flipping case first — often the judge is right and the output really is borderline, which is a product bug, not a metric bug. If it is genuinely ambiguous, tighten the evaluation steps so the decision is mechanical, or replace the judged metric with a deterministic one: a PII detector or regex wrapped as a custom metric has no sampling variance, which is what a hard guardrail deserves.
- What does verbose_mode=True give you, and when would you leave it off?It prints the metric's intermediate steps and reasoning to the console while measuring, so you can see how a judge arrived at a surprising score. It is for development and debugging a handful of cases. On a full suite it floods the output and slows reading the actual results, so you turn it on deliberately and turn it back off.
- Would you gate a deployment on a strict quality metric?Not on a quality metric — helpfulness or tone under strict_mode fails on nearly everything and stops carrying information. Gate on strict metrics only where the criterion is a hard rule with a decidable answer, such as no PII or schema-valid output, and gate quality on a continuous score against a threshold with a tolerance you chose from measured run-to-run spread.
saying these in an interview costs you the question
- Thinks strict_mode just raises the threshold a little
- Believes the configured threshold still applies under strict_mode
- Uses strict_mode on helpfulness or tone metrics
- Confuses strict_mode with verbose_mode
- Assumes a binary score removes judge non-determinism