skip to content

When does reinforcement fine-tuning beat prompting on a task with an automatic grader?

level: seniorimportance: should knowfreq 34%

answer

  1. a grader instead of gold answers
  2. correctness must be machine-checkable
  3. needs a non-zero base success rate
  4. the proxy is what gets optimized
  5. outcomes you have, rationales you do not

basics

~20 s

Reinforcement fine-tuning fits when correctness is machine-checkable and the model already succeeds sometimes. You supply problems and a grader instead of gold answers, and training reinforces the model's own successful attempts — useful when you have verifiable outcomes but no written solutions.

solid answer

~50 s

Supervised fine-tuning needs demonstrations: an input and the exact output you want. Reinforcement fine-tuning needs only a **grader** — a program that scores an attempt — plus a supply of problems. That difference decides the choice. If your task has a deterministic check (a closed case's known outcome, a passing test suite, a schema that validates, a numeric tolerance) and you have thousands of checkable problems but nobody to write gold completions, RFT is the shape that fits. It also needs the base model to succeed at least occasionally, because training amplifies its own successful trajectories; a model that never succeeds gives the optimizer nothing. The failure mode is reward hacking: optimize hard against a grader that only approximates the goal and the model will find the gap. As of mid-2026 RFT is exposed by some providers and absent on others, so availability is a real constraint on the decision.

code

python · 12 lines
python
def grade(attempt: dict, case: dict) -> float:
    """Score one adjudication attempt against a closed case's recorded outcome."""
    if attempt.get("decision") not in {"clear", "escalate"}:
        return 0.0  # malformed output earns nothing
    score = 1.0 if attempt["decision"] == case["final_decision"] else 0.0
    cited = set(attempt.get("cited_rule_ids", []))
    required = set(case["rule_ids_applied"])
    if required:
        score += 0.5 * len(cited & required) / len(required)
    if cited - required:
        score -= 0.25  # penalise padding the citation list
    return max(0.0, score)

go deeper

for a junior

Know that some fine-tuning is driven by a grading program rather than by example answers, and that it only applies when correctness can be checked automatically.

for a middle

Explain the data difference clearly: supervised tuning needs written target outputs, reinforcement tuning needs problems plus a scorer, and preference tuning needs comparisons. Say which you would pick given what data exists.

for a senior

Show the preconditions and the diagnostics: a non-zero base success rate, a grader that checks outcomes rather than surface features, a held-out grading path, and hand-inspection of top-scoring samples to catch exploitation early.

for a principal

Own the objective design. Argue about whether the computable reward is a faithful proxy for the business goal, what it will be worth optimizing hard against, and whether the option is even available on the model the organization has standardized on.

## Three supervision shapes, one decision When a team says "we'll fine-tune", they usually mean supervised fine-tuning. In 2026 the practical menu has three shapes, and choosing between them is a distinct decision from choosing whether to train at all. **Supervised fine-tuning (SFT)** learns from *demonstrations*: pairs of input and the exact target output. It needs someone to have written the right answer. **Preference optimization** learns from *comparisons*: for the same input, this output is better than that one. It needs raters, not authors, which is cheaper — you only have to judge, not produce. **Reinforcement fine-tuning (RFT)** learns from a *grader*: a program that takes an attempt and returns a score. It needs neither authors nor raters, only problems and a way to check answers. ## What makes a task RFT-shaped The defining property is a **verifiable reward** — correctness you can compute rather than judge. Concrete instances: a code change is right if the test suite passes; an extraction is right if it validates against a schema and matches known field values; a routing decision is right if it matches the outcome the closed case eventually recorded; a numeric answer is right within tolerance. The second requirement is a **non-zero success rate**. RFT works by sampling several attempts per problem and pushing the model toward the ones that scored well. If the base model never succeeds, every sample scores zero and there is no gradient of preference to exploit. This is why RFT usually elevates a task the model half-does rather than teaching one it cannot do at all — and why, if the pass rate is zero, the correct move is back down the ladder to prompting or retrieval until it is not. The third is **volume of problems**. You need many checkable instances; you do not need answers for them. That inversion is the whole reason to prefer RFT — a compliance team may have hundreds of thousands of alerts with a recorded final disposition and zero written rationales, which is useless for SFT and ideal for a grader. ## Against prompting Prompting can describe the task and can show a few examples, but it cannot search. RFT lets the model discover a strategy — an ordering of checks, a decision procedure, a reasoning length — that nobody wrote down, by rewarding whatever actually scores. When the task has a crisp objective and the gap is "the model is inconsistent rather than ignorant", that search is what buys the improvement. When the task has no objective check, RFT has nothing to optimize and you are back to prompting or SFT. ## Reward hacking, the characteristic failure A grader is a proxy. Optimize against it long enough and the model will exploit the difference between the proxy and the goal: satisfying an assertion without doing the work, producing output that validates but is empty, exploiting a tolerance, or gaming a length or keyword heuristic. Mitigations are unglamorous — write graders that check outcomes rather than surface features, add penalty terms for degenerate answers, hold out a grading path the model was never trained against, and inspect the highest-scoring samples by hand because that is where the exploit will be visible first. A related trap is a grader that is *correct but sparse*: a single pass/fail on a long task gives very little signal. Partial credit that is genuinely meaningful helps; partial credit invented to be dense usually becomes the next thing to hack. ## Cost, and where it sits on the ladder RFT is more expensive per unit of progress than SFT because it samples multiple attempts per problem and runs the grader on each. It sits after prompting and retrieval, and it is chosen over SFT on the basis of *what data you have*, not on the basis of ambition. If you already possess good demonstrations, SFT is cheaper and simpler. If you possess outcomes but not demonstrations, RFT is the only shape that uses them. ## Availability is part of the decision As of mid-2026 the three supervision shapes are not uniformly available. Some hosted providers expose supervised, preference and reinforcement fine-tuning as separate products; several major closed providers expose no public fine-tuning API at all; open-weight models support all three through open training libraries but require you to own the compute and the infrastructure. A recommendation that ignores whether the option exists on your chosen model is not a real recommendation. ## Answering well Say what RFT needs that SFT does not (a grader, not gold answers), what it needs that prompting cannot supply (search over strategies), the two hard preconditions (machine-checkable correctness, non-zero base success rate), and the characteristic failure (the model optimizing the proxy). Then note that availability constrains the choice.

  • Your base model's pass rate on the grader is zero. What do you do?
    Do not start reinforcement training — with every sample scoring zero there is nothing for the optimizer to prefer. Go back down the ladder: improve the prompt, add retrieval so the model has the inputs it was missing, decompose the task into steps that are individually achievable, or move to a stronger base. If demonstrations exist, a round of supervised tuning first to lift the pass rate above zero is the standard bootstrap.
  • How do you tell reward hacking from genuine improvement?
    Read the highest-scoring samples by hand — exploits show up at the top of the distribution, not the middle. Keep a held-out grading path the model was never optimized against, and watch for score climbing on the training grader while the held-out check and human review flatten or fall. Degenerate patterns are usually visible: empty-but-valid outputs, padded citation lists, answers that satisfy an assertion without doing the work.
  • You have both outcomes and written rationales. Which shape do you pick?
    Usually supervised tuning first, because demonstrations are the cheaper and more stable signal and they set the output shape directly. Reinforcement tuning then becomes an optional second stage that pushes consistency on the cases the supervised model still gets wrong, using the outcomes you already hold. Running both in that order is common; starting with the reinforcement stage when good demonstrations exist wastes sampling budget.

Supervised tuning is learning from worked solutions in the back of the book. Reinforcement fine-tuning is having only an answer key: you attempt, get marked, and keep what scored.

saying these in an interview costs you the question

  • Reinforcement fine-tuning can teach a task the model never gets right
  • Any task can be graded if you write a clever enough judge prompt
  • Reward hacking is a theoretical concern, not a practical one
  • Reinforcement tuning removes the need for an evaluation set
  • Recommending a training shape the chosen model does not support

context