For a fixed GAN generator, what is the optimal discriminator and what does the generator then minimize?
answer
- fix the generator, optimize pointwise
- a log y + b log(1 - y)
- a ratio of two densities
- substitute back, get a divergence
- floor at minus log four
basics
~20 sFor a fixed generator, the optimal discriminator is D*(x) = p_data(x) / (p_data(x) + p_g(x)). Substituting it back turns the objective into a constant plus twice the Jensen-Shannon divergence between the two distributions, which is zero only when they are equal.
solid answer
~50 sHold the generator fixed, so `p_g` is a fixed distribution. The objective becomes an integral over x of `p_data(x) log D(x) + p_g(x) log(1 - D(x))`, and since D may take any value in (0, 1) independently at each x, you can maximize pointwise. The scalar map `a log y + b log(1 - y)` peaks at `y = a/(a + b)`, so `D*(x) = p_data(x) / (p_data(x) + p_g(x))` — a density ratio, not a hard classifier. Substituting it back gives `V(D*, G) = -log 4 + 2 * JSD(p_data, p_g)`, where JSD is the symmetric divergence between each distribution and their mixture. So against a perfect judge, the generator is minimizing that divergence; it hits zero only when `p_g = p_data`, where `D*` is 1/2 everywhere and V sits at its floor of about -1.386 nats. The result assumes an unrestricted D and exact expectations, which real training gives you neither of.
code
python · 23 linesimport math
def gauss(x, mu, s=1.0):
return math.exp(-((x - mu) ** 2) / (2 * s * s)) / (s * math.sqrt(2 * math.pi))
def value_at_optimal_d(shift, lo=-16.0, hi=26.0, n=42001):
step = (hi - lo) / (n - 1)
total = 0.0
for i in range(n):
x = lo + i * step
p = gauss(x, 0.0) # data: mean 0, sd 1
q = gauss(x, shift) # generator: mean shift, sd 1
d = p / (p + q + 1e-300) # optimal discriminator at this x
total += (p * math.log(max(d, 1e-300)) + q * math.log(max(1 - d, 1e-300))) * step
return total
for shift in (0.0, 1.0, 4.0, 10.0):
print(shift, round(value_at_optimal_d(shift), 4))
# 0.0 -1.3863 the two distributions coincide: floor of -log 4, D* = 0.5 everywhere
# 1.0 -1.1635
# 4.0 -0.1209
# 10.0 -0.0 near-disjoint: a perfect discriminator exists, objective flat at 0go deeper
You are not expected to reproduce the derivation. Recall the punchline: the best possible discriminator reports a ratio of the two densities, and it sits at 1/2 everywhere once the generator matches the data.
Be able to run the pointwise argument: fix the generator, maximize a log y + b log(1 - y) at each point, and land on p_data/(p_data + p_g). Knowing the substitution yields a constant plus a divergence is the next step up.
Do the derivation and then attack it. Name what the theorem assumes — an unrestricted discriminator, exact expectations, convexity in distribution space — and explain why a near-disjoint generator makes the idealised objective flat.
Frame the decision this result informs: an objective whose gradient dies when the model is far from the data is a structural risk, and choosing to accept it, replace the objective, or use a different generative family is your call to justify.
## What the question is really asking This is the one piece of GAN theory that gets whiteboarded. It has two halves: solve the inner maximization in closed form, then see what the outer minimization becomes once you substitute that solution back in. ## Half one: the optimal discriminator Start from the value function `V(D, G) = E_x~p_data[log D(x)] + E_z~p_z[log(1 - D(G(z)))]` and rewrite the second expectation as an expectation over the generator's induced distribution `p_g` rather than over the noise. Both terms are then integrals over the same sample space: `V(D, G) = INT_x [ p_data(x) log D(x) + p_g(x) log(1 - D(x)) ] dx` Now the key move. Hold `G` fixed and treat `D` as an arbitrary function — not a network with parameters, but a free choice of a number in (0, 1) at every point x. Because the integrand at one point does not constrain the integrand at any other, maximizing the integral means maximizing the bracket separately at each x. The bracket has the form `f(y) = a log y + b log(1 - y)` with `a = p_data(x)`, `b = p_g(x)`, `y = D(x)`. Differentiate: `f'(y) = a/y - b/(1 - y)`. Set it to zero: `a(1 - y) = b y`, so `y = a/(a + b)`. The second derivative `-a/y^2 - b/(1 - y)^2` is negative throughout (0, 1), so this is the maximum. Hence `D*(x) = p_data(x) / (p_data(x) + p_g(x))` Read that as an estimate of a **density ratio**, mapped into (0, 1). It is 1/2 where the two densities agree, near 1 where data dominates, near 0 where the generator dominates. It does not depend on the discriminator's architecture — the architecture only determines how well a trainable network can approximate this function. A useful sanity check: at a point where the generator puts three times the data's density, `D* = p/(p + 3p) = 0.25`. ## Half two: what the generator is then minimizing Substitute `D*` back: `V(D*, G) = INT p_data log[ p_data/(p_data + p_g) ] + INT p_g log[ p_g/(p_data + p_g) ]` Write `p_data + p_g = 2m` where `m = (p_data + p_g)/2` is the equal mixture of the two. Each log splits into `log(p/m) - log 2`, and since both densities integrate to 1 the two `-log 2` terms contribute `-2 log 2 = -log 4`. What is left is the average of the divergence from `p_data` to the mixture and from `p_g` to the mixture — that average is exactly the Jensen-Shannon divergence, so `V(D*, G) = -log 4 + 2 * JSD(p_data, p_g)` JSD is non-negative, symmetric, and equals zero only when the two distributions coincide. Therefore the outer minimization is minimized precisely at `p_g = p_data`, where the value is `-log 4` (about -1.386 in nats) and `D*(x) = 1/2` everywhere. That is the equilibrium statement people quote. ## The awkward corollary JSD is *bounded*: in nats it never exceeds `log 2`. Take the toy where the data is a unit-variance Gaussian at 0 and the generator produces a unit-variance Gaussian at some shift. As the shift grows, overlap vanishes, `D*` approaches a step function that is 1 on one side and 0 on the other, JSD saturates at its ceiling, and `V(D*, G)` flattens out at 0. A shift of 8 and a shift of 20 score essentially the same. The idealised objective can tell you the generator is wrong but not which way to move it — and that is with the *optimal* discriminator, so it is a property of the objective, not a training bug. ## Why the theorem does not describe your training run Four gaps are worth naming, because the follow-up is usually "so why doesn't it just work?": 1. The proof optimizes over *all functions* D; you optimize over a parameterized family. 2. It assumes the inner maximization is solved exactly at every step; you take a handful of gradient steps. 3. The convexity argument lives in the space of *distributions* `p_g`; you move network *parameters*, and the map from parameters to distributions is not convex. 4. Expectations are replaced by minibatch estimates from a finite training set. So treat the derivation as the statement of what the objective *wants*, not as a description of the dynamics you observe. ## What a strong answer sounds like Get to `D* = p_data/(p_data + p_g)` via the pointwise argument in three lines, state the `-log 4 + 2 JSD` substitution and the `D* = 1/2` equilibrium, and then volunteer the boundedness corollary and at least one of the four gaps above. That combination — clean derivation plus honest limits — is what the question is screening for.
- What happens to this picture when the data and generator distributions barely overlap?The optimal discriminator approaches a step function, the Jensen-Shannon term saturates at its ceiling of log 2 nats, and the value function flattens near 0. A generator far from the data and one merely far-ish score almost identically, so the idealised objective supplies no direction of improvement even with a perfect discriminator.
- What is D*(x) at a point where the generator's density is three times the data's?0.25. The optimal discriminator is p_data/(p_data + p_g) = p/(p + 3p) = 1/4. It is worth saying out loud that this is a smooth density-ratio value, not a hard reject — an optimal discriminator only outputs near 0 or 1 where one distribution genuinely dominates.
- Why doesn't this proof guarantee that GAN training converges?The convexity and uniqueness argument lives in the space of distributions and assumes the inner maximization is solved exactly. Real training moves network parameters with a few alternating gradient steps, over a non-convex parameterization, using minibatch estimates. The pair can orbit an equilibrium rather than settle into one.
saying these in an interview costs you the question
- Says the optimal discriminator is a hard 0/1 classifier everywhere
- Claims D* depends on the discriminator's architecture
- States the equilibrium value of the objective is 0 rather than -log 4
- Presents the divergence result as a description of actual training dynamics
- Forgets that the derivation assumes the inner maximization is solved exactly