In a GAN's minimax objective, what is the discriminator maximizing and the generator minimizing?
answer
- two networks, one shared score
- binary cross-entropy, real versus fake
- outer minimize, inner maximize
- only one term contains the generator
- generator learns through the judge's gradients
basics
~20 sThe discriminator maximizes the log-probability of labelling real data real and generated data fake. The generator minimizes that same quantity, pushing the discriminator toward calling its samples real. One shared value function, optimized in opposite directions.
solid answer
~40 sA GAN trains two networks against one shared value function: `V(D, G) = E_x~p_data[log D(x)] + E_z~p_z[log(1 - D(G(z)))]`. The discriminator D outputs the probability that its input is real and climbs V — it wants `log D(x)` high on real samples and `log(1 - D(G(z)))` high on generated ones, which is exactly binary cross-entropy for a real-versus-fake classifier. The generator G appears only in the second term and descends V, so it wants `D(G(z))` pushed toward 1. G never touches real data; its entire learning signal is gradients passed back through D. Writing the whole thing as `min_G max_D V(D, G)` says the outer player commits to a generator while the inner player is free to respond as well as it can.
go deeper
Be ready to state the two players and their opposite directions in one breath, and to say which term the generator appears in. Knowing that the generator learns only through the discriminator's gradients is the piece most candidates miss.
Explain the value function term by term and connect it to binary cross-entropy for a real-versus-fake classifier. An interviewer expects you to say what expectation each term is taken over and where the generator's gradient physically comes from.
Show you know why this is a game and not a loss: the generator's objective is defined by a discriminator that is itself moving, so nothing decreases monotonically and progress cannot be read off a single number.
Own the framing decision. Be able to argue when an adversarial objective is worth its instability at all versus a likelihood-based generative model, and what it costs a team in tuning time and evaluation infrastructure.
## The setup A GAN is an *implicit* generative model. It never writes down a density for the data. Instead a generator network `G` maps a noise vector `z`, drawn from a fixed simple prior `p_z` (a standard Gaussian, say), to a sample `G(z)`. Pushing the prior through `G` induces some distribution over samples, written `p_g`. Training's goal is to make `p_g` match the data distribution `p_data`, and the whole trick of the adversarial framework is that you can do this using only *samples* from both — no density, no likelihood. The measuring instrument is a second network, the discriminator `D`, which takes a sample and outputs a number in (0, 1) interpreted as "probability this came from the real data". ## The shared value function Both networks are scored by one function: `V(D, G) = E_x~p_data[log D(x)] + E_z~p_z[log(1 - D(G(z)))]` Read each term separately. The first is large when `D` assigns high probability to real samples being real. The second is large when `D` assigns low probability to generated samples being real, since `1 - D(G(z))` is then close to 1. Add them and you have, up to a factor and a sign, the binary cross-entropy of a classifier trained on a balanced mixture of real examples labelled 1 and generated examples labelled 0. That is why the logs are there: they are the Bernoulli log-likelihood of that classification problem. ## Who moves which way - **Discriminator: maximize V.** It is doing ordinary supervised classification against the current generator. Its gradient comes from both terms. - **Generator: minimize V.** `G` appears only in the second term, `log(1 - D(G(z)))`. Minimizing it means making `D(G(z))` large — that is, fooling `D`. The generator's gradient flows backwards through the discriminator into `G`'s parameters, which is why `D` must be differentiable and why `G` never needs to see a real example directly. The compact notation is `min_G max_D V(D, G)`. The ordering matters conceptually: the inner maximization defines, for any given generator, how well the best possible judge can separate its samples from real ones, and the outer minimization then looks for the generator that even the best judge cannot beat. ## What equilibrium means If the generator's distribution ever exactly equalled the data distribution, no discriminator could do better than chance, and it would output 1/2 on every input. That is the fixed point the game aims at. Note what this does *not* say: it makes no claim that a particular sample is good, only that the two distributions are indistinguishable to the judge on offer. ## Why this is not an ordinary loss This is the part interviewers actually probe. In supervised learning you descend one fixed scalar function of your parameters, and a falling loss means progress. Here there are two parameter sets moving against each other on one surface, and neither player's objective is stationary: the generator's loss landscape is defined by the current discriminator and changes the moment the discriminator updates. Practically that means: - Training is **alternating** (or simultaneous) gradient steps on two objectives, not descent on one. - No scalar is guaranteed to decrease monotonically; the pair can circle an equilibrium. - "The loss went down" is not by itself evidence that the generator got better, because it may simply mean the discriminator got worse. ## Common misreadings A weak answer describes `G` as being trained to reconstruct real images, or `D` as scoring image quality on some open-ended scale. Neither is right. `G` has no target image and no reconstruction term; its only teacher is `D`. And `D` is a plain binary classifier — its output is a probability of the label "real", not a quality score. ## What to say in an interview Write the value function, name what each expectation is over, say which player climbs and which descends, point out that only the second term contains `G`, and then add the sentence that separates a memoriser from someone who has trained one: because the two objectives are coupled, this is a game, and its dynamics — not just its optimum — are what you spend your time managing.
- Does the generator ever see real training data during its own update?No. The generator's loss depends on real data only through the discriminator's parameters. It samples noise, produces fakes, and receives gradients backpropagated through the discriminator. That is why the discriminator must be differentiable, and why a discriminator that has stopped learning anything useful leaves the generator with no teacher at all.
- Why is the objective written with logs rather than raw probabilities?Because maximizing it is maximum likelihood for the discriminator viewed as a binary classifier over a balanced mixture of real and generated samples: the log terms are the Bernoulli log-likelihood, and summing logs is the log of a product of per-sample likelihoods. The log also keeps the classifier's gradients well behaved through the usual cross-entropy-with-sigmoid form.
- What does the min-max ordering imply about who is assumed to move first?The generator is the outer player: it must pick a distribution that survives the best response of an inner player who sees that choice. Real training does not honour that ordering — both players take small alternating steps — which is one reason the theoretical picture and the observed dynamics diverge.
saying these in an interview costs you the question
- Says the generator is trained on real samples with a reconstruction loss
- Claims both networks minimize the same loss
- Describes the discriminator as scoring image quality rather than classifying real versus fake
- Thinks the generator compares its output to the nearest real example
- Treats a falling generator loss as proof the samples improved