skip to content

How do you obtain step-level labels for reasoning traces without hand-annotating every step?

level: seniorimportance: should knowfreq 28%

answer

  1. labels are the real bottleneck
  2. correct, neutral or incorrect per step
  3. stop labelling at the first error
  4. sample continuations from each prefix
  5. recoverability is not correctness

basics

~20 s

The cheap route samples several continuations from each prefix and labels a step by how often it still reaches the correct final answer, needing no annotator. Human labelling stays for a stratified audit sample used to validate the automated labels.

solid answer

~50 s

Human step labelling is expensive — a published dataset of human-labelled competition-maths steps runs to roughly 800,000 labels, with a protocol that marks each step correct, neutral or incorrect and stops at the first error, since everything after it is conditioned on a broken state. The scalable alternative is rollout-based labelling: from the prefix ending at step j, sample k continuations, and score the step by the fraction that reach the correct final answer. It needs no annotator but requires an automatically checkable answer, and it conflates "this step is wrong" with "this step makes recovery hard" — a valid but awkward step gets a low label. A grader model is the third option, faster than humans and biased in its own way. In practice you combine them: automated labels in bulk, a stratified human-labelled sample to measure agreement, and published labelling conventions so the numbers mean something to the next team.

go deeper

for a junior

Know that step labels do not come for free and that the common categories are correct, neutral and incorrect. Recognising that an automated method exists is enough at this level.

for a middle

Explain the rollout method concretely — truncate at step j, sample k continuations, label by the fraction reaching the right answer — and state its two preconditions: a checkable answer and enough compute for k generations per step.

for a senior

Design the mixture: bulk automated labels, a stratified human audit sample, measured agreement reported with any derived metric, and re-calibration when the model changes. Be explicit that rollouts measure recoverability rather than validity.

for a principal

Own the annotation programme as a budget and a standard. Decide what fraction of labels must be human, what agreement threshold gates use of the labels, which conventions are mandated so results stay comparable over time, and when a smaller trustworthy set beats a large noisy one.

## Why labels are the bottleneck Every step-level metric and every process reward model depends on a labelled step. Final answers are often free to check; steps never are. The design question is not whether to spend on labels but which mixture of humans, models and automated procedures buys usable labels at the volume you need. ## The human protocol The established convention labels each step of a solution as correct, neutral or incorrect. Correct means mathematically sound and progressing; incorrect means false or invalid; neutral covers steps that are not wrong but do not advance, such as restating the problem or an exploratory move that goes nowhere. Annotation stops at the first incorrect step, because subsequent steps are generated conditioned on a broken premise and their labels would mix two different judgments. The operational details are where projects fail. Annotators need a written rubric with worked edge cases, because "neutral" is where disagreement concentrates. You need double annotation on a fraction of items to measure inter-annotator agreement, and an adjudication path for disagreements. Agreement below your threshold usually means the rubric is under-specified, not that annotators are careless. Volume gives a sense of scale: a widely cited human step-labelled maths dataset carries on the order of 800,000 step labels across tens of thousands of sampled solutions, which is a serious annotation programme, not a sprint task. ## Rollout-based automatic labelling The main way to avoid that bill exploits the same property that makes outcome supervision cheap: a checkable final answer. Truncate the solution after step j, sample k independent continuations from that prefix, and check what fraction reach the correct answer. That fraction becomes the step's label — high means the step left the model in a good position, low means it did not. This needs no human at all and scales with compute rather than headcount. Its limits are real and should be stated in an interview. It requires an automatically verifiable answer, so it does not transfer to open-ended reasoning. It measures recoverability, not correctness: a mathematically valid step that leads into a hard subcase gets a low score, and an incorrect step whose error is commonly self-corrected downstream gets a high one. And the estimate is noisy at small k, so cost is k generations per step, which for long chains is substantial. ## Grader models A capable model can be prompted to label steps against the same rubric a human would use. It is far cheaper than people and far more consistent run to run, and it is also biased in ways that correlate across items — toward confident, conventionally formatted steps, and toward agreement with the final answer it can see, which is why the grader should be shown the step without the outcome where possible. Treat a grader model as an instrument that needs calibration against human labels on a sample, not as a substitute for having any human labels. ## The practical mixture Most teams end up with a three-tier arrangement. Bulk labels come from rollouts or a grader model. A stratified sample — stratified by problem difficulty, by step position, and deliberately over-sampling the disagreement region between the automated methods — goes to human annotators. The agreement between the automated labels and the human sample is measured and reported alongside any metric computed from those labels, so a reader knows the error bar on the labelling itself. When agreement drifts, usually after a model change, the sample is re-drawn. ## What to publish with the labels Whatever the source, the conventions must travel with the numbers: how steps were segmented (line, sentence, or model-emitted delimiter), what neutral means, whether labelling stopped at the first error, and for rollouts, the value of k and the sampling temperature. Two step-level scores computed under different conventions are not comparable, and a comparison across teams that ignores this is the most common way step-level evaluation produces confident nonsense. ## Choosing under a budget If the answers are verifiable and you need volume, start with rollouts and buy a human audit sample. If the domain has no checkable answer — legal reasoning, clinical rationale — rollouts are unavailable and you are choosing between a grader model with human calibration and a smaller, entirely human-labelled set. In that case a smaller, well-adjudicated human set usually beats a large noisy one, because the point of step labels is to be a reference standard.

  • Why does rollout-based labelling sometimes mark a mathematically valid step as bad?
    Because it measures recoverability, not validity. A correct step that steers into a hard subcase yields few continuations that reach the right answer, so it scores low, while a flawed step the model routinely self-corrects can score high. If your downstream use is grading derivations rather than steering search, that mismatch matters and needs a human-labelled check.
  • What do you do when inter-annotator agreement on step labels comes back low?
    Assume the rubric is at fault before the annotators. Disagreement usually concentrates on the neutral category and on steps that are valid but unproductive. Pull the disputed items, write explicit worked examples for those cases, re-train the annotators, and re-measure on a fresh double-annotated sample. Publishing the final agreement number alongside any metric derived from the labels is part of the deliverable.
  • How would you get step labels in a domain with no automatically checkable answer, such as clinical rationale?
    Rollouts are unavailable, so the choice is a grader model calibrated against humans, or a smaller fully human-labelled set. For a reference standard, prefer the smaller adjudicated human set — its value is being trustworthy, not large. Use the grader model for triage and routing items to annotators rather than as the label of record.

saying these in an interview costs you the question

  • Assumes an automated step label is as good as a human one
  • Ignores that rollout labelling needs a checkable final answer
  • Labels steps past the first error without noting the confound
  • Compares step scores across differing segmentation conventions
  • Uses a grader model with no human-labelled calibration sample

context