skip to content

What do process reward models reward that outcome reward models don't?

level: middleimportance: should knowfreq 45%

answer

  1. one reward per solution versus per step
  2. credit assignment across a derivation
  3. free labels versus annotated labels
  4. answer-getting versus grader-acceptable work
  5. judge of steps can itself be gamed

basics

~20 s

An outcome reward model scores only the final answer, so every chain that reaches it gets reinforced, including lucky ones. A process reward model scores each reasoning step, penalising an invalid derivation even when the final answer is right.

solid answer

~50 s

Outcome supervision gives one reward per solution, computed from the final answer. Process supervision gives a reward per step, so credit is assigned to the parts of the derivation that were actually sound. Scoring a stoichiometry derivation line by line makes the difference concrete: an outcome reward model sees only the final mass in grams and reinforces whatever produced it, while a process reward model marks the line that used the wrong molar ratio and withholds credit there even though the mass came out right. The incentives diverge accordingly — outcome supervision optimises for answer-getting and tolerates unfaithful or shortcut reasoning, process supervision optimises for derivations a grader would accept. The cost is the mirror image: outcome labels are free wherever the answer is machine-checkable, while step labels need human annotation or an automated labelling procedure. As of mid-2026 most large-scale reasoning training leans on verifiable outcome rewards, with process signals used more for verification and evaluation.

go deeper

for a junior

Recall the one-line distinction: outcome supervision rewards the final answer, process supervision rewards each step. Being able to say why the first can reinforce a lucky chain is already a good answer at this level.

for a middle

Explain credit assignment — why one terminal reward is a weak signal across twenty lines — and give the cost trade-off: outcome labels are free where answers are checkable, step labels are not.

for a senior

Bring the failure modes of both. Outcome rewards leave the trace unconstrained; process rewards create a learned judge the policy can optimise against. Say how you would detect the second, using held-out accuracy and human step grading on a sample.

for a principal

Own the choice for a programme: which domains admit verifiable outcome checks, what annotation budget a process signal would need, and what labelling convention you would mandate for ambiguous steps. Be explicit that the field has not settled this and that your call is a bet with a review point.

## The problem both are solving When a model produces a twenty-line derivation and you know only whether the last line is right, you face a credit-assignment problem: which of those twenty lines deserved the reward? Reinforcement learning from a single terminal signal has to spread that one bit across the whole trajectory, which is slow and noisy. Outcome and process supervision are the two answers to this, and the difference between them is where the reward signal is attached. ## Outcome supervision An outcome reward model, or a programmatic checker, produces one scalar per solution derived from the final answer alone. Its great virtue is that the label is often free and objective: for arithmetic, symbolic maths, unit tests, or anything with a canonical answer, you do not need a human or a learned model at all — you check the answer. This is what makes it scale to millions of samples, and it is the backbone of reinforcement learning with verifiable rewards, the dominant recipe behind the current generation of reasoning models. What it incentivises is answer-getting. Any chain that reliably produces the right final answer is reinforced, whether it is a clean derivation, a memorised shortcut, a guess dressed up as reasoning, or a chain with two errors that cancel. The trace is unconstrained, so nothing pushes it toward being a truthful account of the computation. ## Process supervision A process reward model scores each step of the derivation, typically as correct, neutral, or incorrect, and a solution-level score is formed by aggregating those step scores. In a stoichiometry problem, the reward attaches to "convert 12.4 g of magnesium to moles", "apply the 2:1 mole ratio", and "multiply by the molar mass of the product" individually. A step that misapplies the mole ratio is penalised at the point it happens, even if the arithmetic downstream rescues the final number. The signal is denser, so credit assignment is far better: the model learns which move was the mistake, not merely that the episode failed. It also directly rewards derivations a human grader would accept, which is what you want when the reasoning itself is the deliverable. ## What each one silently encourages This is the part interviewers care about. Outcome supervision tolerates — and, at scale, actively selects for — reasoning that is not a faithful account of the computation, because nothing in the objective examines it. Process supervision constrains the trace, but it introduces its own hazard: the model is now optimising against a learned judge of steps, so it can learn to write steps that the reward model likes. Confident, conventional-looking, well-formatted derivations score well whether or not they are sound, and the failure is harder to notice than a wrong answer because everything reads plausibly. There is also a subtler asymmetry. A process reward model must decide what an "incorrect" step means. Is a step wrong because it is mathematically false, or because it leads somewhere unproductive? Those are different labels, and the second one penalises legitimate exploration and backtracking — which is exactly the behaviour long-form reasoning models were trained to do. ## Cost and label sourcing Outcome labels: free wherever answers are checkable, otherwise a human or judge model on the final answer only. Process labels: expensive. A published human-labelled step dataset for competition maths runs to roughly 800,000 step labels, which gives a sense of the annotation bill. The cheaper route samples several continuations from each prefix and labels a step by how often it still reaches the right answer, requiring no human at all but inheriting the ambiguity above. ## Where the field actually stands As of mid-2026 this is not settled. Verifiable outcome rewards carried most of the recent gains in reasoning models, largely because they scale without annotation. Process rewards remain valuable as verifiers and as evaluation instruments, and in domains where no automatic answer check exists, but claims that one strictly dominates the other should be treated with suspicion. The honest framing in an interview is: outcome supervision is cheap, scalable and blind to the trace; process supervision is expensive, dense, and only as good as the step labels you can afford.

  • If process supervision gives a denser signal, why is most large-scale reasoning training still driven by outcome rewards?
    Because outcome labels are free wherever the answer is machine-checkable, which lets the recipe scale to enormous sample counts without annotation. Process labels need humans or a labelling procedure, and a learned step judge adds a second model that can be gamed. Density buys better credit assignment; verifiability buys scale, and scale has so far won.
  • What does a process reward model do with a step that is mathematically valid but leads nowhere useful?
    That depends entirely on the labelling convention, and it is a real design decision. Labelling it incorrect punishes exploration and backtracking, which long-form reasoning depends on. Labelling it neutral preserves that behaviour but weakens the signal. Publish the convention with the model, because two process reward models with different conventions are not comparable.
  • How would you notice that a policy is gaming a process reward model rather than reasoning better?
    Watch for the step score rising while independent measures stay flat: final-answer accuracy on a held-out set, human step grading on a sample, and performance on perturbed problems. A policy that has learned the judge's stylistic preferences produces confident, conventionally formatted derivations whose correctness under an unrelated grader has not moved.

saying these in an interview costs you the question

  • Says process supervision is simply better with no mention of cost
  • Assumes outcome rewards check the reasoning trace somehow
  • Confuses a process reward model with a step-by-step prompt
  • Ignores that a learned step judge can itself be gamed
  • Claims the outcome-versus-process question is settled

context