skip to content

How does TRPO's hard KL constraint differ from PPO's clipped objective?

level: seniorimportance: nice to knowfreq 30%

answer

  1. one constrains, the other only discourages
  2. average divergence versus per-sample ratio
  3. Fisher matrix as the local metric
  4. conjugate gradient plus backtracking line search
  5. first-order simplicity bought the popularity

basics

~20 s

TRPO maximises the surrogate subject to an explicit constraint that average KL between old and new policy stays under a threshold, solved by a natural-gradient step plus a line search. PPO approximates that idea first-order, with clipping.

solid answer

~50 s

TRPO poses a constrained problem: maximise the importance-weighted surrogate subject to the mean KL divergence between old and new policy staying below a threshold. It solves that with a natural-gradient direction - the gradient preconditioned by the inverse Fisher information matrix, obtained by conjugate gradient on Fisher-vector products so the matrix is never formed - then backtracks along that direction until a step both satisfies the constraint and improves the surrogate. The trust region is therefore checked on every update. PPO keeps the motivation and drops the machinery: instead of measuring distribution distance it bounds each sample's probability ratio and lets plain minibatch gradient ascent do the rest. That is far simpler to implement, cheap enough to run several epochs per batch, and indifferent to architecture. The price is that nothing is enforced - the clip is a per-sample proxy for a constraint on the whole distribution.

go deeper

for a junior

It is enough to know that an earlier method solved a constrained problem with an explicit limit on how much the policy distribution may change, and that the clipped objective is the cheap, popular successor to that idea.

for a middle

Be able to state the constrained problem in words - maximise the surrogate subject to average divergence below a threshold - and say why a first-order clipped objective replaced it: simpler code, cheaper updates, several epochs per batch.

for a senior

Show that you know what is genuinely lost: the clip constrains per-sample ratios, not an averaged divergence, and it enforces nothing. Talk about what you measure in a run to recover the assurance the constraint used to provide.

for a principal

Own the argument about when a checked guarantee is worth its complexity - high-cost or safety-relevant environments where a destroyed policy is expensive, versus cheap parallel simulation where a heuristic plus checkpoints is the better engineering trade.

## The theory both descend from Both methods come from the same result: the expected return of a new policy can be written in terms of the old policy's advantages plus a term that depends on how much the state distribution shifts. That gives a lower bound of the form *surrogate improvement minus a penalty proportional to the divergence between the policies*. Maximising the bound guarantees the true return does not go down. The awkward part is that the theoretically correct penalty coefficient is very conservative, so faithfully optimising the bound produces steps too small to be practical. ## What TRPO does with it TRPO converts the penalty into a constraint with a knob you can set: maximise `E[ (pi_new(a|s)/pi_old(a|s)) * A ]` subject to `E_s[ KL(pi_old(.|s) || pi_new(.|s)) ] <= delta`. The constraint is on a *divergence between distributions*, averaged over the states in the batch - a statement about the policy as a whole, not about any individual sample. Solving it needs three pieces: 1. **A local quadratic model of the KL.** Near the old parameters, the average KL is well approximated by a quadratic form whose matrix is the Fisher information matrix of the policy distribution. This makes the trust region an ellipse in parameter space whose shape reflects how sensitively the policy's *output distribution* responds to each parameter direction - not a sphere of fixed parameter radius. 2. **A natural-gradient direction.** The step direction is the gradient preconditioned by the inverse Fisher matrix. That matrix is far too large to build for a network, so it is handled with conjugate gradient using Fisher-vector products, which require only extra backward-style passes rather than an explicit matrix. 3. **A backtracking line search.** The quadratic model is only an approximation, so the proposed step is shrunk geometrically until an actual evaluation shows the KL constraint holds and the surrogate genuinely improved. This is what makes the constraint real rather than nominal. The result is a method that will refuse a bad step. The cost is substantial: second-order machinery, several extra passes per update, fragility with parameter sharing between policy and value networks, and an implementation most teams find hard to get right. ## What PPO does instead PPO keeps the goal - do not let one update move the policy far - and abandons every part of the apparatus. It has no constraint, no Fisher matrix, no line search. The clipped surrogate is an ordinary differentiable objective that first-order minibatch gradient ascent optimises directly, which is why several epochs over one batch are affordable and why it composes with shared trunks and any architecture you like. The approximation is doubled, and it is worth being precise about where. First, PPO measures closeness in *ratio space per sample* rather than *KL over the state distribution*: a bound on individual likelihood ratios is related to but not the same as a bound on average divergence, and many samples each moving a permitted amount can still add up to a large distributional shift. Second, it does not enforce anything even in ratio space - flattening the objective removes the incentive to move further, but a ratio already outside the band contributes zero gradient rather than being pulled back, and a large learning rate can step past the boundary before the flat region is felt. ## How to talk about the tradeoff TRPO offers a checked guarantee at the price of complexity and per-update cost. PPO offers a heuristic that is much cheaper, much simpler, tolerant of the tricks practitioners want to use, and empirically strong across a wide range of tasks - which is why it became the default and TRPO largely a reference point. The honest framing is that PPO's clip is a first-order stand-in for a trust region, and its reliability comes from a body of practice rather than from a theorem. The practical consequence follows directly: because nothing is enforced, a PPO run needs measurement where TRPO has proof. Watching the fraction of clipped samples and the realised divergence between the pre-update and post-update policy is how you recover, empirically, the assurance the constraint would have given you. ## A middle option There is also a penalty formulation that sits between the two: keep first-order optimisation but add an explicit divergence penalty to the objective, with a coefficient adapted up or down depending on whether the measured divergence overshot or undershot a target. It is simpler than a hard constraint and more directly aimed at distribution distance than clipping is, and it is the natural thing to reach for when the clip alone proves too loose.

  • Why does TRPO need a line search after computing its step direction?
    The step direction comes from a quadratic approximation of the average KL that is only accurate near the old parameters. The full-length step may violate the real constraint or fail to improve the surrogate, so the step is shrunk geometrically and re-evaluated until both conditions hold. Without that check the constraint would be nominal rather than enforced.
  • Why is the Fisher information matrix the right local metric for a trust region on policies?
    Because the constraint is on distribution distance, not parameter distance. To second order the average KL between the old and new policy equals a quadratic form in the parameter change with the Fisher matrix in the middle, so that matrix converts a step in parameters into the change it causes in the policy's output distribution. A fixed-radius ball in raw parameter space would allow huge distribution changes along sensitive directions.
  • In what sense is the clip a weaker statement than a KL constraint?
    It bounds each sample's likelihood ratio individually, whereas the constraint bounds an average divergence over the whole state distribution. Many samples each moving within their permitted band can still produce a large aggregate shift, and the clip provides no restoring force for a ratio already outside its band. It shapes the objective; it does not certify the result.

saying these in an interview costs you the question

  • Says PPO enforces a KL constraint, just more cheaply
  • Claims TRPO uses second-order optimisation to go faster
  • Thinks the trust region is a fixed radius in parameter space
  • Confuses the Fisher matrix with the loss Hessian
  • Cannot say what TRPO's line search is checking

context