skip to content

Why does GRPO drop the value critic that PPO-style RLHF requires?

level: middleimportance: must knowfreq 62%

answer

  1. advantage needs a baseline
  2. the baseline can be learned or measured
  3. sample many answers to the same prompt
  4. group mean replaces the value network
  5. fewer models, more rollouts

basics

~20 s

PPO-style RLHF trains a second network to predict expected return as a baseline. GRPO samples a group of responses per prompt instead and uses the group's own mean reward as that baseline, removing a whole model from memory and from the failure surface.

solid answer

~50 s

An advantage estimate needs a baseline - some notion of how good a response is *relative to what was expected* - or the gradient is dominated by how hard the prompt happens to be. PPO-style RLHF gets that baseline from a learned value network, a second full-sized model trained alongside the policy. GRPO (group-relative policy optimisation) observes that for a prompt you can just sample a group of responses - eight, sixteen, sixty-four - score them all, and use the group's mean reward as the baseline, typically normalising by the group's spread. A response better than its siblings gets positive advantage; worse gets negative. No critic is trained at all. The payoff is roughly a quarter to a third less memory and one fewer model to destabilise, at the cost of many more rollouts per prompt. As of mid-2026 GRPO and its variants - DAPO, GSPO - are the reference family for reasoning post-training, while PPO-style critics have receded.

code

python · 12 lines
python
def group_relative_advantages(rewards):
    n = len(rewards)
    mean = sum(rewards) / n
    var = sum((r - mean) ** 2 for r in rewards) / n
    std = var ** 0.5
    if std == 0.0:
        return [0.0] * n
    return [(r - mean) / std for r in rewards]


print(group_relative_advantages([1.0, 0.0, 1.0, 0.0, 1.0, 0.0, 1.0, 1.0]))
print(group_relative_advantages([1.0, 1.0, 1.0, 1.0]))

go deeper

for a junior

Know that GRPO scores several answers to the same prompt and rewards the ones better than the group average, so no separate value network is needed.

for a middle

Explain why advantage needs a baseline at all, how the group mean supplies one, and the trade of extra sampling for less memory and fewer unstable components.

for a senior

Discuss what happens at the extremes - zero-variance groups, binary rewards, coarse token-level credit - and how variants like DAPO's dynamic sampling or GSPO's sequence-level ratios address them in real runs.

for a principal

Frame the choice as compute allocation: sampling throughput versus training memory, and whether your reward signal is dense enough that a critic would have bought better credit assignment than more rollouts would.

## The problem a baseline solves Policy-gradient methods push up the probability of actions that did better than expected. The words *than expected* matter: raw reward is a terrible training signal because prompt difficulty dominates it. An easy prompt scores 0.9 for every rollout and a hard one scores 0.2 for every rollout, so learning from raw reward mostly teaches the model which prompts are easy. Subtracting a baseline turns reward into an **advantage**, and advantage is what carries the useful signal. ## How PPO-style RLHF gets its baseline The classic recipe trains a value network - a critic - to predict expected return from a given state. It is typically initialised from the reward model or the policy and is comparable in size, so a training run holds four models in play: policy, reference (for KL), reward model, and critic. That is expensive and fragile. The critic must be learned online while the thing it evaluates keeps changing, so its errors are correlated with the policy's, and a mis-fit critic quietly poisons every advantage estimate. It also adds its own hyperparameters and its own way to diverge. ## The group-relative trick GRPO removes the critic by exploiting a property of language-model RL that classic robotics RL does not have: sampling is cheap and repeatable. For each prompt, sample a group of G responses from the current policy, score each, then compute each response's advantage as its reward minus the group mean, commonly divided by the group's standard deviation. The empirical group mean *is* the baseline, and it is unbiased by construction because the samples come from the current policy. Every token of a response inherits that response-level advantage. The clipped-objective machinery and the KL term stay; only the source of the baseline changes. A useful picture: post-training a refund-policy assistant, you sample eight replies to a customer message, score them, and the four above the group mean get reinforced while the four below get pushed down. If all eight score identically the group has zero spread and contributes no gradient - which is a real and important behaviour, not a bug. ## Why it fits verifiable rewards especially well When the reward is a grader - unit tests passing, a checked numeric answer - rewards are often binary. Group-relative advantage converts a batch of binary outcomes into a graded signal automatically: five of eight correct means the correct ones get moderate positive advantage. This pairing of GRPO with verifiable rewards is why the technique spread so quickly through reasoning-model training. ## The mainstream variants All of these are adjustments to the same core, and naming them signals currency: - **DAPO** decouples the clipping range so the upper bound can be raised (clip-higher), adds dynamic sampling that discards all-correct and all-wrong groups which carry no signal, moves the loss to token level, and shapes rewards for over-long generations. - **GSPO** moves the importance ratio from token level to sequence level, which noticeably stabilises training of mixture-of-experts models where token-level ratios are noisy. ## What you give up The critic was not free but it was not useless either. Without it: - **Sampling cost multiplies.** G rollouts per prompt instead of one. The saving is memory and stability, not total compute. - **Credit assignment is coarse.** A single scalar spread over every token means a mostly-correct chain with one bad step gets uniformly punished. A critic could in principle localise blame; process-level rewards are the alternative route. - **Zero-variance groups waste work.** Prompts that are always solved or never solved produce no gradient, which is why dynamic sampling and difficulty-curated prompt sets matter. - **Normalising by group standard deviation can bias updates** toward low-variance prompts, which is why some variants drop that division. ## Answering it well Say what a baseline is *for* first, then that GRPO replaces a learned baseline with an empirical one drawn from a group of samples of the same prompt. Then be concrete about the trade: less memory and fewer moving parts, more sampling, coarser credit assignment.

  • What happens when every response in a group gets the same reward?
    The advantages are all zero, so that prompt contributes no gradient - the group is wasted compute. This is common at the extremes: prompts the model always solves and prompts it never solves. DAPO's dynamic sampling addresses it by resampling until a group contains a mix of outcomes, and curriculum-style prompt selection aims at problems near the model's current success boundary.
  • How do you choose the group size G?
    It trades gradient quality against sampling cost. Too small and the group mean is a noisy baseline, so advantages swing on luck; too large and you burn rollouts for diminishing variance reduction. Typical runs sit in the eight-to-sixty-four range, pushed higher when rewards are binary and success is rare, since you need several successes per group to get any signal at all.
  • Does removing the critic remove the reference model too?
    No. The reference policy serves the KL term, which keeps the policy anchored near its starting behaviour, and that is orthogonal to how the baseline is computed. GRPO runs still hold a frozen reference alongside the policy, though some implementations drop or shrink the KL term for verifiable-reward reasoning training where drift is less of a concern.

saying these in an interview costs you the question

  • Saying GRPO is cheaper overall, ignoring the extra rollouts
  • Claiming GRPO removes the KL penalty and reference model
  • Thinking the group baseline is a fixed constant rather than per-prompt
  • Describing the critic as the reward model - they are different networks
  • Assuming per-token credit assignment survives without a value network

context