skip to content

Why is a GAN's discriminator never trained to convergence between generator updates?

level: seniorimportance: nice to knowfreq 32%

answer

  1. the inner maximum is never solved
  2. cost, then overfitting
  3. a perfect judge teaches nothing
  4. k steps, then one step
  5. keep it ahead, not finished

basics

~20 s

Solving the inner maximization at every generator step is prohibitively expensive, and on a finite training set a converged discriminator overfits and saturates, leaving the generator with almost no gradient. Training instead alternates a few discriminator steps with one generator step.

solid answer

~50 s

The objective is written as a minimum over the generator of a maximum over the discriminator, which suggests solving the inner problem first. Practice alternates instead: k discriminator steps, then one generator step, with k usually small. Two reasons. Cost — an inner loop run to convergence at every generator update multiplies training time, and on a finite sample the converged discriminator overfits, learning which exact points are in the training set rather than a useful density ratio. Signal — a near-perfect discriminator is exactly the regime where its output saturates and the generator's gradient collapses, so pushing the inner player to optimality can stall the outer one. The working compromise keeps the discriminator slightly ahead but never finished: strong enough that its judgement means something, weak enough that it still transmits gradient. Which way to move k depends on which player is currently outpacing the other.

go deeper

for a junior

Recall the shape of the loop: a few discriminator steps, then one generator step, repeating. Knowing that the discriminator is not trained to completion each round is the point to take away.

for a middle

Explain why the loop departs from the nested notation: the inner solve is unaffordable, and a discriminator pushed to optimality on a finite sample memorizes it. Be able to state what the generator's gradient does as the discriminator's confidence rises.

for a senior

Demonstrate that you have balanced this in a real run. Name the symptoms of a discriminator that is too strong and one that lags, and say which observable you would watch before changing the step ratio.

for a principal

Own the framing that the ratio, the learning-rate asymmetry and the relative capacity are one decision, not three. Be ready to argue how much tuning budget a training scheme with no monotone progress signal deserves before you switch approach.

## The gap between the notation and the algorithm `min_G max_D V(D, G)` reads like a nested optimization: for each candidate generator, solve the inner maximization, then take a step on the outer one. That is not what any GAN training loop does. The standard loop is: 1. Take k gradient steps on the discriminator, each on a fresh batch of real samples and a fresh batch of generated ones, holding the generator's parameters fixed. 2. Take one gradient step on the generator, holding the discriminator's parameters fixed. 3. Repeat. The original formulation used exactly this alternating scheme with a small k, and reported experiments at k = 1. The inner maximization is therefore approximated by a handful of steps that never converge, and the discriminator is always chasing a generator that keeps moving. ## Reason one: cost Each inner solve would need many passes over data before the discriminator settled, and you would need one such solve per generator update. The generator typically needs tens or hundreds of thousands of updates. The nested version is simply not affordable, and since the generator only moves a little between updates, most of that work would be re-derived from scratch to reach an answer barely different from the last one. Warm-starting from the previous discriminator is exactly what the alternating loop already does — the k steps are the continuation of an optimization that never restarts. ## Reason two: a converged discriminator overfits The theory's `D*(x) = p_data(x)/(p_data(x) + p_g(x))` is defined over *distributions*. Your discriminator sees a finite training set. Given enough capacity and enough steps, the best thing it can do on that empirical objective is memorize: output near 1 on the specific training points and near 0 everywhere else, including on regions where the true data density is high but no training sample happens to land. That function is a perfect classifier of the sample and a useless estimate of the density ratio, and the generator that chases it is being pushed toward the training points themselves rather than toward the distribution they came from. ## Reason three: a perfect judge gives no gradient This is the one that bites in practice. The generator learns only through the discriminator, and the amount of signal it receives depends on the discriminator's confidence. As the discriminator approaches a perfect separator, its output saturates on generated samples, and the gradient reaching the generator shrinks toward zero. Push the inner maximization all the way and you can arrive at a state where the discriminator is right about everything and the generator has stopped learning — technically the correct inner solution, practically a dead run. So the discriminator is deliberately kept in a middle band: ahead of the generator, since a discriminator that lags gives noisy or misleading feedback and lets the generator drift toward whatever the current weak judge happens to dislike, but never finished. ## Choosing k - **k = 1** is the common default: cheapest, and it keeps the two players moving on similar timescales. - **Larger k (say 5)** is chosen when the discriminator is visibly underfitting the current generator — its accuracy on freshly generated samples sits near chance — so each generator step is guided by a more accurate estimate of the ratio, at roughly k times the discriminator's cost per generator update. - The symptom of **k too large** is the stall described above: the discriminator separates the batches almost perfectly, and the gradient reaching the generator collapses. - The symptom of **k too small** is a generator that exploits a judge which has not caught up, moving in directions that stop helping as soon as the discriminator does catch up. Step count is not the only knob. Giving the two players different learning rates achieves a similar asymmetry without duplicating batches — a two-timescale schedule where the discriminator moves faster is a published and widely used alternative. Capacity is a third: a discriminator much larger than the generator behaves like a large k even at k = 1. ## Why none of this converges the way an ordinary loss does Alternating gradient steps on two coupled objectives is not gradient descent on any single scalar function. There is no quantity guaranteed to decrease each iteration, and the parameter pair can orbit an equilibrium rather than approach it — an effect familiar from simple two-player games where each player's best response keeps chasing the other's last move. Tuning the update ratio is partly an attempt to damp exactly this circling, which is why the answer is a band to stay inside rather than a single correct number. ## Interview framing Say the notation implies a nested solve, say the algorithm does something else, and give the three reasons: cost, overfitting on a finite sample, and vanishing generator signal from a saturated judge. Then name the knobs — step ratio, learning-rate asymmetry, relative capacity — and describe how you would tell which direction to move.

  • What practically differs between one discriminator step per generator step and five?
    Five keeps the discriminator closer to its optimum for the current generator, so each generator step follows a more accurate density-ratio estimate, at roughly five times the discriminator's cost. It also risks pushing the discriminator into the confident regime where the generator's gradient collapses. One step is cheaper and keeps a lagging but still informative judge.
  • How would you tell the discriminator has become too strong?
    It separates real from freshly generated batches almost perfectly, its outputs on generated samples cluster near zero, and the gradient norms arriving at the generator's parameters shrink by orders of magnitude while the generator's parameters stop moving meaningfully. Those three observed together point at the inner player having run away.
  • Besides the step ratio, what else controls the balance between the two players?
    Relative learning rates — a two-timescale schedule with a faster discriminator produces a similar asymmetry without extra batches — and relative capacity, since a discriminator much larger than the generator behaves like a large step ratio even at one step each. All three knobs pull in the same direction and are usually tuned together.
  • Is alternating gradient descent here the same as descending a single loss?
    No. Two coupled objectives with opposed interests admit no single scalar that decreases every iteration, so the parameter pair can circle an equilibrium instead of settling. That is a property of the game's dynamics, not of a bad learning rate, and it is why balance between the players is something you manage throughout training.

saying these in an interview costs you the question

  • Says the discriminator should be trained to optimality because the notation says max
  • Claims a stronger discriminator is always better for the generator
  • Ignores that a converged discriminator on finite data memorizes the training points
  • Treats the update ratio as a constant that never needs revisiting
  • Describes the pair as jointly descending one loss function

context