skip to content

Policy Gradient Methods

Optimising a parameterised policy directly with the score-function estimator, as REINFORCE does. Interviewers ask when you would prefer this to learning action values at all.

on this pageshow

questions

4

Why do policy gradient methods suit continuous action spaces better than value-based control?

level: juniorimportance: must knowfreq 66%

answer

  1. think about the shape of the action set
  2. what happens at decision time?
  3. maximising over an infinite set
  4. no argmax over real-valued vectors
  5. parameterise the policy, sample the action

basics

~20 s

Policy gradient methods tune the policy's parameters directly, so an action is sampled from a distribution. Value-based control has to maximise a value estimate over all actions at every step, which is intractable when the action is a real-valued vector.

solid answer

~50 s

A value-based controller stores or estimates a value for each state-action pair and acts by taking the action that maximises it. That maximisation is a search over the action set, which is fine for a handful of discrete moves and hopeless for a robot arm whose action is a vector of seven joint torques: there is no finite set to scan, and discretising each joint multiplies out into an unusable grid. A policy gradient method never needs that maximisation. It parameterises the policy itself, for example a Gaussian per joint whose mean and standard deviation are functions of the state, and moves the parameters by gradient ascent on expected return. Acting means sampling from that distribution, which costs the same whether the action space holds two options or a real-valued vector. Exploration comes for free from the spread of the distribution, and the learned spread can shrink as the controller improves.

go deeper

for a junior

Be ready to state the core contrast in one breath: value-based methods pick the action that maximises an estimated value, policy methods sample from a learned distribution over actions. Know that the maximisation is what breaks on continuous actions.

for a middle

Explain the mechanics: what a Gaussian policy outputs per action dimension, why discretisation costs grow exponentially in the number of dimensions, and how the sampled spread doubles as the exploration mechanism.

for a senior

Show you know the price of the trade. Policy gradients buy usable continuous control at the cost of noisy gradients and experience you can reuse barely at all, and they converge to a local optimum of expected return rather than a guaranteed best policy.

for a principal

Own the representation decision. Argue when the action space genuinely justifies a policy-gradient stack versus when a coarse discretisation, a hand-written controller, or a supervised model on logged decisions gets the same result for a fraction of the engineering and interaction budget.

## The two ways to represent "what to do" An agent has to answer one question at every step: given the state I am in, which action do I take? There are two structurally different ways to encode the answer. A **value-based** method learns a numeric score for state-action pairs, usually written `Q(s, a)`, and derives its behaviour from that score: acting means finding the action with the largest `Q(s, a)` in the current state. The policy is not stored anywhere; it is *computed* from the value estimate at decision time. A **policy gradient** method stores the policy itself, as a parameterised probability distribution `pi_theta(a | s)`. There is no value table to consult and no maximisation to solve: the agent samples an action from the distribution. Learning means nudging the parameter vector `theta` so that expected return goes up. ## Why the action space decides which one is usable The hidden cost in the value-based route is the argmax. `max over a of Q(s, a)` is an optimisation problem solved at every single timestep. When the action set is `{left, right, up, down}`, that is four lookups. When the action is a 7-dimensional vector of joint torques for a robot arm, the set is uncountably infinite. There is no argmax to take. The usual escape is discretisation: chop each joint's torque range into, say, 10 bins. Now the action space has `10^7` = ten million entries, for a coarse resolution that produces jerky control. The cost is exponential in the number of action dimensions, which is why discretisation stops being viable at around two or three continuous dimensions. You also throw away the geometry of the space: two neighbouring torque values are near-identical physically, but the discretised representation treats them as unrelated symbols and has to learn about each separately. ## What a continuous policy looks like The standard parameterisation for a continuous action is a **Gaussian policy**. For each action dimension, the policy maps the state to a mean and a standard deviation, and the action is drawn from that normal distribution. For a 7-joint arm, that is 7 means and 7 spreads. Sampling is O(dimensions) regardless of how fine the control needs to be, and the representation is smooth: shifting the mean slightly shifts behaviour slightly. The spread carries a second job. Because the action is sampled rather than chosen, the policy explores by construction, and the amount of exploration is itself a learned parameter. Early in training a wide spread tries a broad band of torques; as the returns improve, gradient ascent typically narrows the spread and the controller becomes more decisive. Exploration is part of the object being optimised rather than something bolted on outside it. A discrete policy can be parameterised the same way: keep a preference score per action and pass them through a softmax, so the policy is a categorical distribution. Policy gradients work on both; it is the continuous case where they are close to mandatory. ## What you give up This is a trade, not a free win. - **Sample cost.** The gradient estimate is an expectation over trajectories sampled from the current policy, so a batch of experience is valid for roughly one update. Value-based methods can typically replay old experience many times. - **Variance.** A gradient estimated from sampled episode returns is noisy, which is why variance-reduction techniques such as subtracting a baseline are standard rather than optional. - **Local optima.** Gradient ascent on the policy parameters converges to a local optimum of expected return. There is no guarantee of finding the globally best policy, whereas tabular value methods on a small discrete problem come with convergence guarantees to the optimum. ## How to say it in an interview The short version: value-based control needs to *search* the action space at every decision, policy gradient methods need only to *sample* it. When the action is a real-valued vector, searching is the thing you cannot afford, so you learn the policy directly and let its distribution do the choosing and the exploring.

  • How does a Gaussian policy over joint torques explore, if nothing external adds randomness?
    The action is a sample, not a choice: the policy outputs a mean and a standard deviation per dimension and draws from that. The spread is a learned parameter, so exploration is optimised along with everything else and typically narrows as returns improve. Nothing has to be added around the policy to make it try new actions.
  • Could you just discretise the continuous action space and run a value-based method?
    For one or two action dimensions, yes, and it is often the pragmatic choice. Beyond that the grid grows exponentially: ten bins across seven joints is ten million actions. You also cap the control resolution and destroy the ordering between neighbouring torques, so the agent has to learn about physically near-identical actions independently.
  • Do policy gradient methods have any advantage on a discrete action space?
    Yes. They can represent a genuinely stochastic policy, which is what you want when the best behaviour is to mix actions rather than commit, and they change behaviour smoothly as parameters move. A value-based controller's behaviour can flip discontinuously when two action values cross.

Choosing by value is like ordering dinner by reading every line of the menu and picking the highest-rated dish. If the menu is infinite, you cannot read it; instead you learn to describe roughly what you want and let that description produce the order.

saying these in an interview costs you the question

  • Claims a value-based method handles continuous actions unchanged
  • Says policy gradients are just value learning renamed
  • Thinks a policy must output one action, never a distribution
  • Assumes discretising a seven-dimensional action space is cheap
  • Cannot say what quantity the parameters are being tuned against

context

open as a page

How does REINFORCE turn sampled episode returns into a gradient on the policy's parameters?

level: middleimportance: must knowfreq 72%

basics

~20 s

REINFORCE weights the gradient of each action's log-probability by the return that followed it: theta <- theta + lr * G_t * grad log pi(a_t | s_t). Actions from high-return episodes get more probable. The reward itself is never differentiated.

open as a page

In REINFORCE, why does subtracting a baseline from the return cut gradient variance?

level: middleimportance: should knowfreq 57%

basics

~10 s

Weighting by return minus a baseline centres the weights, so better-than-average actions are pushed up and worse-than-average ones down, instead of everything rising. Any baseline independent of the chosen action leaves the gradient unbiased.

open as a page

Why must a REINFORCE run discard its collected trajectories after a single gradient update?

level: seniorimportance: should knowfreq 43%

basics

~20 s

The policy gradient is an expectation over trajectories drawn from the current policy. One update moves the policy, so the batch now comes from a different distribution and estimates the wrong expectation. Every step needs fresh interaction.

open as a page