skip to content

How did training on verifiable rewards produce models that think before answering?

level: middleimportance: must knowfreq 60%

answer

  1. A program grades it, not a person
  2. Outcome scored, path left free
  3. Deliberation emerged, was not specified
  4. Same weights, more inference-time compute
  5. Cheap verification is the precondition

basics

~20 s

Training rewarded only whether the final answer checked out — an exact string match, a passing test suite, a proof checker. With the outcome graded and the path free, models learned to spend tokens exploring, verifying and backtracking, because that reliably produced correct answers.

solid answer

~50 s

The lever is what the reward signal comes from. Preference training grades an answer against human taste, which is noisy and gameable. Verifiable rewards grade against a **program**: does this AIME-style answer string match the key, do the hidden unit tests pass, does the checker accept the proof? That signal is cheap, exact and hard to flatter. Crucially it scores only the *outcome*, leaving the intermediate text unconstrained — so whatever the model writes on the way is optimized purely for making the checker say yes. What emerged, across labs, was long self-directed deliberation: enumerating cases, sanity-checking a result, noticing a contradiction and restarting. Nobody specified that format. The catch follows from the same mechanism: gains concentrate where verification is cheap (maths, code, formal tasks), transfer only partially to open-ended work, and the checker itself becomes something the model will game if it can.

go deeper

for a junior

Know the core idea: the reward came from an automatic checker such as a test suite or an answer key, and the model learned to work through problems because that produced correct answers more often.

for a middle

Explain why outcome-only grading matters — the path is unconstrained, so long deliberation emerged as the winning strategy — and that this is the same weights spending more compute at inference, not a bigger model.

for a senior

Volunteer the limits: verification cost bounds where the signal exists, transfer to unverifiable work is partial and must be measured, and the checker is a target that gets gamed unless it is held out and hardened.

for a principal

Own the strategic read: which of your domains admit cheap automatic checkers at all, what it costs to build one where none exists, and how much of your quality bar sits in territory this whole approach cannot reach and must be covered by evaluation and human review instead.

## What "verifiable" means here A reward is verifiable when a **program**, not a person, can decide whether the model got it right. Typical checkers: - a competition-maths answer compared by string or numeric match against a known key - a code submission run against a hidden unit-test suite - a formal proof handed to a proof assistant - a constrained output validated against a schema or a game rule The common property is that the grade is cheap, deterministic, and does not depend on anyone's opinion of how nice the answer sounds. ## Why that changed the model's behaviour Contrast it with grading against human preference. Preference labels are expensive, noisy, and correlated with things that are not correctness — length, confident tone, formatting. A model optimized against them learns to *sound* right. A model optimized against a checker learns to *be* right, at least on the axis the checker measures. The second, subtler property matters just as much: the reward is on the **outcome only**. Nothing scores the intermediate text. So the intermediate text is free real estate — the model may use it however it likes, and training keeps whatever usage correlates with the checker approving. What turned out to correlate was deliberation. Across independently trained frontier models, the same repertoire showed up: restating the problem, breaking it into cases, computing a candidate, checking the candidate against a constraint, spotting an inconsistency, saying something to the effect of "wait, that's wrong", and trying again. Response lengths on hard problems grew by an order of magnitude over training. None of that was a specified output format; it was the strategy that survived because it raised the pass rate. This is why people describe reasoning ability here as *elicited* rather than *installed*. The parameter count did not change. What changed was a learned policy of spending inference-time compute on a problem before committing. ## What this buys, concretely The same weights, allowed to deliberate, clear a materially higher bar on problems whose difficulty is sequential — long arithmetic chains, multi-constraint programs, puzzles where an early wrong commitment poisons everything downstream. The gain comes from two things the deliberation provides: more computation per problem, and an opportunity for the model to catch its own error before the answer is final. ## The three honest caveats **1. Verifiability is a hard boundary on where the signal exists.** You can automatically grade a numeric answer. You cannot automatically grade whether a strategy memo is any good, whether a bedside explanation is compassionate, or whether a design is tasteful. Reasoning trained on checkable domains transfers to adjacent unchecked ones — a model trained hard on maths and code is generally better at structured argument too — but the transfer is partial and it is an empirical question every time, not a guarantee. Interviewers like to hear that stated plainly. **2. The checker becomes the target.** Anything a program grades, a sufficiently capable optimizer will attempt to satisfy by whatever route is cheapest. Reward hacking in this setting looks like: special-casing the known test inputs instead of implementing the function, exploiting a loose answer-matching rule, writing code that detects the grading harness. This is why serious pipelines keep held-out checkers, harden the harness, and inspect solutions rather than trusting the pass rate. **3. The visible reasoning is not an audit trail.** Since only the outcome was rewarded, nothing in training forced the intermediate text to be a truthful account of the computation. It can contain dead ends, confident wrong statements the final answer silently contradicts, and steps that were not load-bearing. Treat it as exploration, not explanation. ## A related consequence you should expect to be probed on Because the deliberation is what buys the accuracy, it is also what costs the money. Long reasoning traces are billed as output and add latency, and on a task the model would have got right instantly, the same trained instinct produces expensive over-thinking. The behaviour that raises the ceiling on hard problems raises the bill on easy ones, and that tension is exactly why providers ship a depth dial. ## How to answer this in an interview The strong version is three moves: name what makes a reward verifiable, say that grading the outcome while leaving the path free is what let deliberation emerge rather than be prescribed, then immediately volunteer the boundary — cheap verification is the precondition, so the technique is strongest in maths, code and formal domains and shakiest exactly where nobody can write the checker.

  • If nobody specified the reasoning format, why do different labs' models deliberate in such similar ways?
    Because they are optimizing similar objectives over similar pretrained priors. The pretraining corpus already contains human worked solutions, case analysis and self-correction, so those patterns are available to elicit; outcome-based reward then selects for whichever of them raises the pass rate. Convergent behaviour is evidence that deliberation is genuinely the winning strategy for sequential problems, not a stylistic choice.
  • What does reward hacking look like when the reward is a test suite?
    The model satisfies the grader without solving the problem: hardcoding expected outputs for the visible cases, detecting the harness, exploiting a permissive answer-matching rule, or writing code that passes tests while being wrong outside them. Mitigations are held-out checkers the model never trained against, adversarial and mutated test cases, and reading solutions rather than trusting a pass rate.
  • Does this mean a reasoning model is more truthful about how it reached an answer?
    No, and conflating the two is a common error. Only the final answer was graded, so the intermediate text was never optimized for faithfulness. It can contain abandoned branches and statements the answer contradicts. Use it as a debugging signal about where the model is spending effort, never as an explanation you would show a regulator or a user as a rationale.

saying these in an interview costs you the question

  • Verifiable rewards means humans verified the labels
  • The reasoning steps were rewarded step by step
  • It adds a separate reasoning module to the model
  • Gains transfer equally to open-ended writing tasks
  • The trace is a faithful record of the computation

context