Why does an eval that scores blank answers like wrong answers push a model to guess?
answer
- compare expected value of guess versus silence
- zero for wrong equals zero for silent
- the eval is a selection pressure
- state the threshold in the prompt
- report coverage beside accuracy
basics
~20 sUnder scoring where a wrong answer and an abstention both earn zero, guessing has positive expected value and saying "I don't know" has none. Optimizing against that rule selects for confident guessing, so the eval manufactures the hallucination it then reports.
solid answer
~50 sWith binary accuracy scoring, an abstention is worth exactly as much as a wrong answer — nothing — while a guess carries whatever probability of being right the model can muster. Expected score is therefore maximized by always answering, at any confidence. Because that scoring shape dominates public leaderboards and most internal eval sets as of mid-2026, every selection pressure applied through them — checkpoint choice, post-training reward, prompt iteration — favours a model that never abstains. The fix is to price abstention into the rubric: state a confidence threshold in the prompt and score against it, for example plus one for correct, minus three for wrong, zero for "I don't know", so that guessing below 75% confidence is negative expected value. Then report coverage and accuracy-on-answered separately, so a model that answers 60% of questions almost perfectly is distinguishable from one that answers everything at 70%.
go deeper
Understand that if a wrong answer and a refusal both score zero, guessing can only help the score. Know that this is about how the test is graded, not about the model wanting to please anyone.
Work the expected-value comparison explicitly and propose a concrete rubric — points for correct, a penalty for wrong, zero for abstaining — with the implied confidence threshold stated in the prompt.
Show that the eval is a selection pressure across checkpoint choice, reward and prompt iteration, and report coverage alongside accuracy-on-answered. Include unanswerable items deliberately and name over-abstention as the opposing regression.
Own the threshold as a product risk decision. Derive the penalty from the real cost ratio per surface, defend it with a risk-coverage curve, and set the coverage floor below which the refusal behaviour is itself an incident.
## The arithmetic that creates the behaviour Take the standard scoring rule: one point for a correct answer, zero for a wrong one, zero for an abstention. Now put the model in front of a question it half-knows, with subjective probability p of getting it right. Answering has expected value p. Abstaining has expected value 0. For any p above zero — including a wild guess — answering wins. There is no confidence so low that silence is the better play. That is the whole mechanism. It requires no bad intent, no faulty training, no exotic failure mode. Any process that selects for score under that rule selects for a model that always answers. ## Where the pressure actually enters This matters because the scoring rule is not confined to a report at the end. It is the selection signal at several stages: which post-training checkpoint gets shipped, which reward the model is optimized against on tasks with a checkable answer, which prompt variant wins your A/B, which fine-tune you keep. Each of those is a filter, and the filter has the same shape. A candidate model that correctly says "the handbook I was given does not cover that" loses to a candidate that invents a plausible clause and happens to be right some of the time. The concrete case is worth holding onto. An internal HR policy assistant is asked about a parental-leave rule that genuinely is not in its document snapshot. The correct behaviour — "I don't have that policy" — scores zero on an accuracy-only eval, identical to fabricating a leave entitlement. Run that eval as your selection criterion and you will systematically ship the bot that invents entitlements, then be surprised by the incident. ## Repricing the rubric The fix is to make the scoring reflect what you actually want, which is almost never raw accuracy. **Put a price on being wrong, and state it in the prompt.** A rule such as plus one for correct, minus three for wrong, zero for abstaining implies a threshold: answering is worth it only above 75% confidence. Crucially, the threshold must be stated to the model as part of the task, not just applied silently at grading time — otherwise you are asking it to guess your risk appetite. This is behavioural calibration: you are not asking for a probability, you are asking for a decision under a stated payoff, which is far easier to elicit and to verify. **Report two numbers instead of one.** Coverage — the share of questions answered — and accuracy on answered items. A single accuracy figure collapses two very different systems: one answering 60% of questions at 97%, another answering 100% at 70%. For a policy bot or a clinical summarizer the first is obviously better, and the collapsed metric cannot see it. Sweeping the abstention threshold and plotting accuracy against coverage gives you a risk-coverage curve, which is the honest way to compare and the right artefact to take to a product decision about where the threshold should sit. **Keep unanswerable items in the set.** An eval built only from questions your corpus covers cannot measure abstention at all. Deliberately include questions outside the snapshot, and score the refusal as the correct answer. This is the single cheapest change most teams have not made. ## The honest limits Repricing your own eval does not fix the models you buy. Public leaderboards still overwhelmingly report plain accuracy as of mid-2026, so the pretrained candidates you choose among have already been shaped by that pressure, and no rubric of yours undoes it. What you can do is select among them on your rubric, and make abstention safe in your own product surface. There is also a real cost to over-correcting. A penalty set too high produces a model that refuses constantly, which users experience as useless and which quietly pushes them to a less careful tool. Over-abstention is a failure mode with its own incident report; it is just a quieter one. The threshold is a product decision — how bad is a wrong parental-leave answer versus a shrug — and should be argued in those terms, with the risk-coverage curve on the table, rather than picked by whoever is writing the eval that week. Finally, an abstention is only as good as what surrounds it. "I don't have that policy" is a correct answer and a poor experience if it dead-ends; pair it with the escalation path, the document it searched, or the person to ask. ## What an interviewer is listening for They want the expected-value argument stated plainly, the recognition that the eval is a selection pressure rather than a passive measurement, and a concrete rubric change with numbers in it. Mentioning coverage and accuracy-on-answered as a pair, and naming over-abstention as the opposing failure, is what separates the senior answer from a recital of "the model is rewarded for guessing".
- How do you set the wrong-answer penalty rather than picking a number arbitrarily?Derive it from the cost ratio you actually face. A penalty of minus k implies answering only above k/(k+1) confidence, so choose k from how much worse a confident wrong answer is than a refusal on that surface — a wrong drug-interaction claim justifies a far higher k than a wrong restaurant recommendation. Then sanity-check it by sweeping the threshold and looking at where the risk-coverage curve puts coverage, because a defensible ratio that leaves you answering 5% of questions is not shippable.
- What is the failure mode of over-penalizing wrong answers?A model that refuses constantly. Users read it as useless, route around it to a less careful tool, and you lose both the utility and the safety you were buying. Over-abstention also hides quality problems, because a system that answers nothing is never observably wrong. Watch coverage as a first-class metric, set a floor on it, and treat a sharp drop after a rubric change as a regression rather than a success.
- Your eval set contains only questions your document corpus covers. What is that hiding?Everything about abstention. If every question is answerable, refusing is always wrong and answering always has a chance, so the set cannot distinguish a model that knows its limits from one that guesses well. Deliberately include out-of-scope questions — the parental-leave rule that is genuinely not in the snapshot — and score the refusal as correct. Without those items your accuracy number is measuring recall on a rigged sample.
saying these in an interview costs you the question
- Says the model chooses to bluff to please the user
- Treats a single accuracy number as sufficient
- Applies a confidence threshold without telling the model
- Builds an eval set where every question is answerable
- Assumes maximizing abstention is always the safer setting