skip to content

Why do GANs minimize -log D(G(z)) instead of the minimax form log(1 - D(G(z)))?

level: middleimportance: must knowfreq 58%

answer

  1. gradient magnitude, not the optimum
  2. the discriminator wins early
  3. flat exactly where you need signal
  4. sigmoid saturation, then chain rule
  5. one scales as D, the other as 1 - D

basics

~20 s

Early on, the discriminator rejects generated samples confidently, and log(1 - D(G(z))) is flat in that regime, so the generator receives almost no gradient. The non-saturating form -log D(G(z)) is steepest exactly where samples are being rejected.

solid answer

~50 s

The literal minimax game has the generator minimize `log(1 - D(G(z)))`. At the start of training the samples are obviously fake, so `D(G(z))` sits near 0 — and that is precisely where this term is flattest. Differentiating through the discriminator's sigmoid, the gradient of `log(1 - D)` with respect to D's pre-sigmoid logit is `-D(G(z))`, which vanishes as the discriminator grows confident: the generator stalls exactly when it most needs to move. The non-saturating variant keeps the same direction of preference but has the generator maximize `log D(G(z))` — equivalently minimize `-log D(G(z))` — whose gradient with respect to that same logit is `-(1 - D(G(z)))`, near its largest when samples are being rejected. It is no longer literally the minimax objective: both forms want `D(G(z))` driven to 1, but their gradient magnitudes differ, so the clean divergence reading of the game no longer holds exactly.

go deeper

for a junior

Recall that the loss actually used in practice is -log D(G(z)) and that the reason is a vanishing gradient early in training, when the discriminator easily spots fakes. Knowing which form is which is the minimum here.

for a middle

Explain the mechanics: differentiate both losses through the discriminator's sigmoid and show one gradient scales as D(G(z)) and the other as 1 - D(G(z)). Say explicitly which regime training starts in.

for a senior

Show you know the cost of the heuristic: the swap keeps the generator's preferred direction but breaks the exact divergence interpretation of the minimax game, and it leaves the discriminator's objective untouched, so the two players no longer share one value function.

for a principal

Be ready to defend taking an interpretable objective and replacing it with a heuristic that trains. Articulate when a theoretically clean formulation is worth defending and when shipping requires the version whose gradients survive initialization.

## The two candidate generator losses The value function is `V(D, G) = E[log D(x)] + E[log(1 - D(G(z)))]`, and the game says the generator minimizes V. Since only the second term contains the generator, the **saturating** (minimax) generator loss is `L_G_sat = E_z[ log(1 - D(G(z))) ]`, minimized. The **non-saturating** (heuristic) alternative is `L_G_ns = E_z[ -log D(G(z)) ]`, minimized. Both are minimized by driving `D(G(z))` toward 1, so they agree about what the generator wants. They disagree about how hard they push at each level of the discriminator's confidence, and that difference decides whether training starts at all. ## Where the saturation comes from The discriminator ends in a sigmoid: `D = sigmoid(s)`, where `s` is the pre-sigmoid logit the network computes. Everything the generator learns arrives through `s`, so the honest way to compare the two losses is to differentiate each with respect to `s`. For the saturating form: `d/ds log(1 - sigmoid(s)) = -1/(1 - D) * D(1 - D) = -D`. For the non-saturating form: `d/ds [-log sigmoid(s)] = -1/D * D(1 - D) = -(1 - D)`. So the two gradient magnitudes are `D(G(z))` and `1 - D(G(z))` — mirror images. Now think about early training. The generator is random, the discriminator learns within a few hundred steps to reject its output, and `D(G(z))` falls to something like 0.02. The saturating loss then supplies a gradient of magnitude 0.02: essentially nothing. The non-saturating loss supplies 0.98: nearly full strength. Late in training, if the generator ever fools the discriminator completely, the two swap roles — but that regime is not where training gets stuck, so the asymmetry is the one you want. Notice the mechanism carefully, because this is where candidates fumble. It is *not* that `log(1 - D)` has a small derivative with respect to `D` — `d/dD log(1 - D) = -1/(1 - D)`, which is about -1 near D = 0. The flatness comes from the **sigmoid saturating**: a confident discriminator sits far out on the logit axis, where `dD/ds = D(1 - D)` is tiny, and the chain rule kills the signal before it reaches the generator. ## What the swap costs you theoretically The elegance of the minimax form is that, against an optimal discriminator, the generator is minimizing a symmetric divergence between its distribution and the data's. The non-saturating loss gives that up. It is not the same function, not a monotone transform of it, and not a constant offset — only the fixed point (the generator wants to be indistinguishable) survives. A published analysis of the non-saturating gradient under an optimal discriminator shows it corresponds to a combination of a reverse KL term and a *negatively* weighted Jensen-Shannon term, so it is not descending any single well-behaved divergence. In practice this is a trade everyone takes: an objective with a beautiful interpretation and no usable gradient at initialization is worse than a heuristic that actually moves. The discriminator's own objective is untouched. It still maximizes the original value function on real and generated batches; only the generator's loss changes. That asymmetry — different loss for each player, no longer a single shared value function — is why some people prefer to describe modern adversarial training as two coupled objectives rather than as a minimax game. ## When the saturating form is not actually a problem If the discriminator is weak or deliberately handicapped, `D(G(z))` hovers near 0.5, both losses give gradients around 0.5 in magnitude, and the saturating form trains fine. The pathology is specific to a *confident* discriminator, which is the normal state at initialization because distinguishing random noise output from real data is trivial. So the failure is a start-of-training failure above all, and it can recur later whenever the discriminator pulls decisively ahead. ## Interview framing A complete answer has four beats: state both losses; explain that the discriminator wins early so `D(G(z))` is near 0; show that the saturating loss's gradient through the logit scales as `D(G(z))` while the non-saturating one scales as `1 - D(G(z))`; and concede that the swap preserves the generator's preferred direction but breaks the exact divergence interpretation of the game. If you can add that the discriminator's update is unchanged, you have covered everything the question is testing.

  • Do the two generator losses share the same optimum?
    They share the direction and the fixed point: both are minimized by driving D(G(z)) toward 1. What differs is the gradient magnitude as a function of the discriminator's confidence. The consequence is theoretical rather than about where training aims — the non-saturating form no longer corresponds to minimizing the symmetric divergence the minimax form does.
  • Does swapping the generator loss change the discriminator's update?
    No. The discriminator still maximizes the original value function using real samples labelled real and generated samples labelled fake. Only the generator's loss changes, which is why the pair is no longer literally optimizing one shared value function in opposite directions.
  • If the discriminator were weak rather than strong, would the saturating loss still stall?
    No. With D(G(z)) near 0.5 both losses give comparable gradients, so the saturating form trains perfectly well. The stall is specific to a confident discriminator, which is the default at initialization because random generator output is trivially separable from real data.

saying these in an interview costs you the question

  • Says the non-saturating loss moves where the optimum is
  • Claims the two forms are algebraically equivalent
  • Explains the stall without mentioning the discriminator's sigmoid saturating
  • Asserts the non-saturating form still minimizes the same divergence exactly
  • Thinks the swap also changes the discriminator's objective

context