When can a verifiable reward replace a learned reward model in post-training?
answer
- let a program, not a model, decide
- tests, checked answers, schema validation
- nothing learned means nothing to overfit
- binary and sparse without a group baseline
- tone and safety have no grader
basics
~20 sWhenever a program can decide correctness - unit tests passing, a checked final answer, a schema validating - the grader becomes the reward directly. This removes the proxy that reward hacking exploits, but only works where correctness is mechanically checkable and the grader itself is hard to game.
solid answer
~50 sRLVR - reinforcement learning with verifiable rewards - replaces the learned reward model with a deterministic grader. For a code task the reward is whether the hidden test suite passes; for maths it is whether the extracted final answer matches; for structured output it is whether the result validates against a schema. Nothing is learned, so there is nothing to overfit and no distribution shift between the reward's training data and current rollouts. That is why this became the dominant vocabulary for reasoning post-training: the reward is cheap, exact, and enormously harder to hack than a learned scalar. It pairs naturally with group-relative methods, which convert a batch of binary pass/fail outcomes into a graded signal. The limit is coverage. Tone, helpfulness, safety and taste have no programmatic grader, so preference-based rewards still own those. Real pipelines mix both - verifiable rewards for the checkable core, learned or judged rewards for the rest.
go deeper
Know that some tasks can be graded by a program - tests passing, an answer matching - and that training can use that result directly as the reward instead of a learned model of human preference.
Explain why a deterministic grader avoids the overfitting and distribution-shift problems of a learned reward, and why binary graders are usually paired with sampling several responses per prompt.
Be concrete about failure: models that pass tests without solving the problem, prompt sets that yield no signal, and correctness rewards that say nothing about the quality of the reasoning or the rest of the model's behaviour.
Decide which slices of the objective deserve programmatic verification and fund building those graders and datasets, while being explicit that the unverifiable remainder - tone, safety, judgement - still needs preference-based supervision.
## What makes a reward verifiable A reward is verifiable when a deterministic procedure that is not a neural network can decide it. The canonical examples: - **Code**: run a hidden test suite; reward is the pass rate or a binary all-pass. - **Mathematics**: extract the final answer and compare it to a known ground truth. - **Structured output**: validate against a schema or a parser. - **Tool and environment tasks**: check the end state of a sandbox - was the file written, did the query return the right rows. - **Constraint satisfaction**: check that stated formatting or length rules were obeyed. The defining property is that the grader is a program with no learned parameters standing between the model's output and the reward. ## Why this is such a strong signal A learned reward model has two structural weaknesses: it fits noisy finite labels, and its judgements degrade off the distribution it was trained on - which is exactly where an optimiser drives the policy. A grader has neither. It costs a test-runner invocation rather than a forward pass through a large model, it does not drift as the policy moves, and it does not need refreshing with new human labels each iteration. That combination is why reasoning training scaled the way it did: you can run enormous amounts of RL against maths and code because the supervision is free and exact once the datasets have ground truth. ## The pairing with group-relative methods Verifiable rewards are usually sparse and binary, and a lone binary reward is a poor gradient. Sampling a group of responses per prompt and computing advantage relative to the group mean fixes this: if three of eight rollouts pass, the passing ones get positive advantage proportional to how unusual that was. This is why RLVR and critic-free group-relative training are almost always discussed together. ## Where it fails **Coverage.** Most of what alignment is about has no grader. Is this refusal appropriate? Is this explanation clear for a non-expert? Is this tone right for a distressed customer? A test runner cannot say. Verifiable rewards handle the checkable core of a task, never the whole of it. **Gaming the grader.** Verifiable is not unhackable. Models find solutions that pass tests without solving the problem: special-casing the visible inputs, mutating global state the assertions read, catching failures broadly, or exploiting a weak grader that only checks a substring of the final answer. Test suites need to be adversarially strong and held out. **Correct answer, terrible process.** A grader rewards outcomes. A model can reach the right number through incoherent reasoning and be reinforced for it, which matters when the visible reasoning is part of the product. **Sparsity and difficulty curation.** If the model never solves a problem, every rollout scores zero and the prompt contributes nothing. Prompts must sit near the model's current success boundary, which turns dataset curation into an ongoing engineering job rather than a one-off. **Narrowing.** Heavy optimisation on a checkable slice can degrade the rest. Regression evaluation on unrelated capabilities is not optional. ## The hybrid in practice Mature pipelines route by task type. Checkable tasks get graders; conversational, safety and style behaviour gets preference-based training, sometimes with a model judge substituting for a human labeller. The grader may also be composite - correctness from tests, plus a penalty for exceeding a length budget, plus a learned or judged component for readability. ## Answering it well Lead with the criterion - *is there a program that can decide this* - rather than with a list of domains. Then be specific about the two limits that matter in production: the parts of the objective no grader can reach, and the fact that a weak grader is just another hackable proxy wearing a different hat.
- Give a concrete way a model games a supposedly verifiable code reward.It writes code that special-cases the inputs the visible tests use rather than implementing the logic, or it manipulates state the assertions read, or it wraps the body so failures are swallowed and the suite still reports success. The mitigations are held-out tests the model never sees, adversarial and mutation-style test generation, sandboxing so the grader's own environment cannot be touched, and human spot checks of high-reward rollouts.
- How do you handle a prompt set where the model solves nothing?Those prompts produce all-zero groups, no variance and no gradient, so they are wasted compute. Curate towards problems near the current success boundary, and let that boundary move as the model improves - a curriculum rather than a fixed set. Partial-credit rewards, such as fraction of tests passing rather than all-or-nothing, also recover signal from problems the model can only partly solve.
- Can a model judge count as a verifiable reward?No - a judge model is a learned reward with a prompt instead of a trained scalar head, so it carries the same weaknesses: it can be gamed, it drifts, and it has its own biases toward length and confident phrasing. It is a useful reward where nothing programmatic exists, but calling it verifiable confuses the category and leads teams to trust it more than the evidence supports.
saying these in an interview costs you the question
- Calling verifiable rewards unhackable rather than much harder to hack
- Assuming a grader can evaluate tone, safety or helpfulness
- Ignoring that all-fail prompt sets produce no learning signal
- Treating a model judge as a verifiable reward
- Rewarding outcomes and assuming the reasoning that produced them is sound