Why does the pathwise gradient estimator usually beat the score-function estimator on variance?
answer
- both are unbiased, variance is the issue
- one differentiates f, one only reads it
- score times scalar value carries little direction
- expectation of the score is zero
- subtract a baseline, bias unchanged
basics
~20 sBoth estimators are unbiased, but the pathwise one differentiates the objective at the sample, so each draw carries directional information. The score-function estimator only multiplies a scalar value by a score vector, which is far noisier and needs a baseline.
solid answer
~50 sBoth estimate the same gradient of an expectation, and both are unbiased. The pathwise (reparameterized) estimator writes the sample as a differentiable transform of parameter-free noise and returns `df/dz` times `dz/dtheta`, so each draw tells you which way the objective actually slopes at that point. The score-function estimator returns `f(z)` times `grad log q(z)`: it never looks inside `f`, only at its numerical value, and infers a direction from the correlation between high values and the score. That is a weaker signal whose variance typically grows with latent dimension. Concretely, for `f(z) = z` with a Gaussian latent, the pathwise estimate of the gradient with respect to the mean is exactly one on every draw, variance zero; the score-function version is unbiased but has variance `mu^2/sigma^2 + 2`. A baseline near the mean of `f` removes the first term, which is why the raw form is rarely usable.
code
python · 19 linesimport random, statistics
mu, sigma, n = 2.0, 1.0, 200000
random.seed(0)
pathwise, score = [], []
for _ in range(n):
eps = random.gauss(0.0, 1.0)
z = mu + sigma * eps # reparameterized draw
pathwise.append(1.0) # d/dmu (mu + sigma*eps)
score.append(z * (z - mu) / sigma ** 2) # f(z) * d/dmu log q(z)
print(statistics.fmean(pathwise), statistics.pstdev(pathwise)) # 1.0 0.0
print(statistics.fmean(score), statistics.pstdev(score)) # ~1.0 ~2.45
# with a baseline b = mu, the score-function spread drops from
# sqrt(mu**2/sigma**2 + 2) to sqrt(2)
based = [(z - mu) * (z - mu) / sigma ** 2
for z in (mu + sigma * random.gauss(0.0, 1.0) for _ in range(n))]
print(statistics.fmean(based), statistics.pstdev(based)) # ~1.0 ~1.41go deeper
Recall that there are two ways to get a gradient of an expectation: differentiate through a reparameterized sample, or weight the objective's value by the gradient of the log-density. Know that both are unbiased.
Explain the mechanics of each estimator and what each one requires: a reparameterizable distribution plus a differentiable objective for one, only an evaluable objective and a differentiable log-density for the other.
Demonstrate judgment about variance. Be ready to say why the score-function form needs a baseline before it is usable, name the variance-reduction tools you would try, and admit when the pathwise route is closed.
Own the tradeoff across a system: whether to redesign a component so the pathwise route stays open, or accept an estimator whose variance budget becomes a permanent tuning burden for every team touching that model.
## Two unbiased estimators of the same gradient You want `grad_theta` of `E over z ~ q_theta(z) of f(z)`. The parameter sits in the distribution, not in the integrand, so you cannot simply push the gradient inside. Two standard constructions fix that, and both are unbiased. **Pathwise (reparameterized).** Write the sample as `z = g(theta, eps)` where `eps` comes from a fixed, parameter-free base distribution. Then the parameter has moved into the integrand and the estimator is just backpropagation through the transform: ``` estimate = df/dz * dz/dtheta, z = g(theta, eps) ``` It requires two things: `q` must be reparameterizable, and `f` must be differentiable with respect to `z`. **Score-function (likelihood-ratio).** Use the identity `grad q = q * grad log q`: ``` grad_theta E_q[f(z)] = E_q[ f(z) * grad_theta log q_theta(z) ] ``` It requires far less: `f` need only be *evaluated*, never differentiated, and `z` may be discrete. All you need is a differentiable log-density in the parameters. That generality is exactly why it is the fallback whenever the pathwise route is closed. ## Why the variances differ so much The pathwise estimator uses the *shape* of `f`. It asks: at this particular sample, which direction increases the objective, and how does moving the parameter move the sample? Each draw is an informed direction. The score-function estimator never differentiates `f`. It draws a sample, reads off a scalar, and multiplies a score vector — a direction that says "make this sample more likely" — by that scalar. The estimator learns only from the correlation between the objective's value and where the sample landed. Nothing about the local slope of `f` reaches it. The score vector's magnitude also tends to grow with the dimension of the latent, so the estimator typically degrades as the latent gets wider, while the pathwise estimator's variance depends mostly on how smooth `f` is. **A head-to-head you can do on paper.** Take `z` normal with mean `mu` and scale `sigma`, and the trivial objective `f(z) = z`, whose exact gradient with respect to `mu` is one. - Pathwise: `z = mu + sigma * eps`, so the estimate is `d(mu + sigma*eps)/dmu = 1` on *every* draw. Unbiased, variance exactly zero. - Score-function: `grad_mu log q = (z - mu)/sigma^2`, so the estimate is `z * (z - mu)/sigma^2`. Its mean is one — also unbiased — but writing `z = mu + sigma*u` with `u` standard normal turns the estimate into `mu*u/sigma + u^2`, whose variance is `mu^2/sigma^2 + 2`. So on an objective where one estimator is *exact*, the other has variance that blows up as the latent's scale shrinks or its mean moves away from the origin. That is not a pathological example; it is the typical gap. ## Baselines Because `E[grad log q] = 0` (the integral of the density is one, and the gradient of a constant is zero), you may subtract any constant `b` from `f(z)` without introducing bias: ``` estimate = (f(z) - b) * grad_theta log q_theta(z) ``` In the example above, taking `b = mu`, which is the mean of `f`, leaves `u^2` and variance `2`: the `mu^2/sigma^2` term is gone. A baseline that tracks the running mean of the objective, or a constant fitted from recent batches, is the cheapest variance reduction there is, and a score-function estimator used without one is usually unusable at realistic step sizes. Other standard tools are control variates, common random numbers across the compared terms, and simply drawing several samples per example. ## Choosing between them Use pathwise whenever you can: it is exact-ish, cheap, and needs no tuning. Reach for the score function when one of its preconditions fails — the latent is discrete, the objective is a non-differentiable measurement of the decoded output, or the sampler is a black box you cannot differentiate through. Then budget for variance reduction as part of the design, not as an afterthought. One honest caveat: "pathwise has lower variance" is an overwhelming empirical regularity, not a theorem. If `f` has huge or wildly varying derivatives, the pathwise estimator inherits that roughness and can be the noisier of the two. It is also worth remembering that low variance is not the same as low bias: both estimators here are unbiased, and the biased alternatives — continuous relaxations of discrete draws — trade exactly along that axis.
- When is the score-function estimator the only option available?When a precondition of the pathwise route fails: the latent is discrete, so no differentiable transform of continuous noise produces it; the objective is a non-differentiable measurement of the decoded output; or the sampling procedure is a black box you cannot differentiate through. In those cases the score-function form still applies, and you pay for it with variance reduction.
- Why does subtracting a baseline leave the estimator unbiased?Because the expected score is zero: the density integrates to one, and the gradient of that constant is zero, so the expectation of the gradient of the log-density vanishes. Subtracting any constant therefore subtracts zero in expectation while shrinking the spread of the product, provided the baseline is close to the mean of the objective and not correlated with the current sample.
- Can the pathwise estimator ever be the higher-variance choice?Yes. It differentiates the objective at the sample, so it inherits the objective's roughness. If the function has very large or rapidly varying derivatives, or long chains that amplify them, the pathwise estimate can swing more than a well-baselined score-function estimate. The usual ordering is an empirical regularity, not a theorem, so measure it if the training is unstable.
The pathwise estimator feels the slope of the hill under its feet. The score-function estimator only hears a score shouted from wherever it happens to stand and has to infer the uphill direction from many such shouts.
saying these in an interview costs you the question
- Says the score-function estimator is biased
- Thinks a baseline changes what is being optimised
- Claims the pathwise estimator works for discrete latents
- Believes lower variance means lower bias
- Cannot say what the score-function estimator needs from f