skip to content

Generator and Discriminator

Generation as a two-player game where a generator learns to fool a discriminator, with mode collapse and unreadable losses as the standard failures. Interviewers use it to test dynamics reasoning.

on this pageshow

explore

questions

17

In a GAN's minimax objective, what is the discriminator maximizing and the generator minimizing?

level: juniorimportance: must knowfreq 72%

answer

  1. two networks, one shared score
  2. binary cross-entropy, real versus fake
  3. outer minimize, inner maximize
  4. only one term contains the generator
  5. generator learns through the judge's gradients

basics

~20 s

The discriminator maximizes the log-probability of labelling real data real and generated data fake. The generator minimizes that same quantity, pushing the discriminator toward calling its samples real. One shared value function, optimized in opposite directions.

solid answer

~40 s

A GAN trains two networks against one shared value function: `V(D, G) = E_x~p_data[log D(x)] + E_z~p_z[log(1 - D(G(z)))]`. The discriminator D outputs the probability that its input is real and climbs V — it wants `log D(x)` high on real samples and `log(1 - D(G(z)))` high on generated ones, which is exactly binary cross-entropy for a real-versus-fake classifier. The generator G appears only in the second term and descends V, so it wants `D(G(z))` pushed toward 1. G never touches real data; its entire learning signal is gradients passed back through D. Writing the whole thing as `min_G max_D V(D, G)` says the outer player commits to a generator while the inner player is free to respond as well as it can.

go deeper

for a junior

Be ready to state the two players and their opposite directions in one breath, and to say which term the generator appears in. Knowing that the generator learns only through the discriminator's gradients is the piece most candidates miss.

for a middle

Explain the value function term by term and connect it to binary cross-entropy for a real-versus-fake classifier. An interviewer expects you to say what expectation each term is taken over and where the generator's gradient physically comes from.

for a senior

Show you know why this is a game and not a loss: the generator's objective is defined by a discriminator that is itself moving, so nothing decreases monotonically and progress cannot be read off a single number.

for a principal

Own the framing decision. Be able to argue when an adversarial objective is worth its instability at all versus a likelihood-based generative model, and what it costs a team in tuning time and evaluation infrastructure.

## The setup A GAN is an *implicit* generative model. It never writes down a density for the data. Instead a generator network `G` maps a noise vector `z`, drawn from a fixed simple prior `p_z` (a standard Gaussian, say), to a sample `G(z)`. Pushing the prior through `G` induces some distribution over samples, written `p_g`. Training's goal is to make `p_g` match the data distribution `p_data`, and the whole trick of the adversarial framework is that you can do this using only *samples* from both — no density, no likelihood. The measuring instrument is a second network, the discriminator `D`, which takes a sample and outputs a number in (0, 1) interpreted as "probability this came from the real data". ## The shared value function Both networks are scored by one function: `V(D, G) = E_x~p_data[log D(x)] + E_z~p_z[log(1 - D(G(z)))]` Read each term separately. The first is large when `D` assigns high probability to real samples being real. The second is large when `D` assigns low probability to generated samples being real, since `1 - D(G(z))` is then close to 1. Add them and you have, up to a factor and a sign, the binary cross-entropy of a classifier trained on a balanced mixture of real examples labelled 1 and generated examples labelled 0. That is why the logs are there: they are the Bernoulli log-likelihood of that classification problem. ## Who moves which way - **Discriminator: maximize V.** It is doing ordinary supervised classification against the current generator. Its gradient comes from both terms. - **Generator: minimize V.** `G` appears only in the second term, `log(1 - D(G(z)))`. Minimizing it means making `D(G(z))` large — that is, fooling `D`. The generator's gradient flows backwards through the discriminator into `G`'s parameters, which is why `D` must be differentiable and why `G` never needs to see a real example directly. The compact notation is `min_G max_D V(D, G)`. The ordering matters conceptually: the inner maximization defines, for any given generator, how well the best possible judge can separate its samples from real ones, and the outer minimization then looks for the generator that even the best judge cannot beat. ## What equilibrium means If the generator's distribution ever exactly equalled the data distribution, no discriminator could do better than chance, and it would output 1/2 on every input. That is the fixed point the game aims at. Note what this does *not* say: it makes no claim that a particular sample is good, only that the two distributions are indistinguishable to the judge on offer. ## Why this is not an ordinary loss This is the part interviewers actually probe. In supervised learning you descend one fixed scalar function of your parameters, and a falling loss means progress. Here there are two parameter sets moving against each other on one surface, and neither player's objective is stationary: the generator's loss landscape is defined by the current discriminator and changes the moment the discriminator updates. Practically that means: - Training is **alternating** (or simultaneous) gradient steps on two objectives, not descent on one. - No scalar is guaranteed to decrease monotonically; the pair can circle an equilibrium. - "The loss went down" is not by itself evidence that the generator got better, because it may simply mean the discriminator got worse. ## Common misreadings A weak answer describes `G` as being trained to reconstruct real images, or `D` as scoring image quality on some open-ended scale. Neither is right. `G` has no target image and no reconstruction term; its only teacher is `D`. And `D` is a plain binary classifier — its output is a probability of the label "real", not a quality score. ## What to say in an interview Write the value function, name what each expectation is over, say which player climbs and which descends, point out that only the second term contains `G`, and then add the sentence that separates a memoriser from someone who has trained one: because the two objectives are coupled, this is a game, and its dynamics — not just its optimum — are what you spend your time managing.

  • Does the generator ever see real training data during its own update?
    No. The generator's loss depends on real data only through the discriminator's parameters. It samples noise, produces fakes, and receives gradients backpropagated through the discriminator. That is why the discriminator must be differentiable, and why a discriminator that has stopped learning anything useful leaves the generator with no teacher at all.
  • Why is the objective written with logs rather than raw probabilities?
    Because maximizing it is maximum likelihood for the discriminator viewed as a binary classifier over a balanced mixture of real and generated samples: the log terms are the Bernoulli log-likelihood, and summing logs is the log of a product of per-sample likelihoods. The log also keeps the classifier's gradients well behaved through the usual cross-entropy-with-sigmoid form.
  • What does the min-max ordering imply about who is assumed to move first?
    The generator is the outer player: it must pick a distribution that survives the best response of an inner player who sees that choice. Real training does not honour that ordering — both players take small alternating steps — which is one reason the theoretical picture and the observed dynamics diverge.

saying these in an interview costs you the question

  • Says the generator is trained on real samples with a reconstruction loss
  • Claims both networks minimize the same loss
  • Describes the discriminator as scoring image quality rather than classifying real versus fake
  • Thinks the generator compares its output to the nearest real example
  • Treats a falling generator loss as proof the samples improved

context

open as a page

What is mode collapse in GAN training, and how do you spot it in generated samples?

level: juniorimportance: must knowfreq 80%

basics

~20 s

Mode collapse is when a GAN generator maps many different noise vectors to a few nearly identical outputs, covering only part of the real data. You spot it by decoding a fixed batch of noise vectors and seeing the samples repeat.

open as a page

In a conditional GAN, why does the discriminator receive the class label too?

level: middleimportance: must knowfreq 62%

basics

~20 s

Only the discriminator can create pressure to obey the label. If it judges images alone, any realistic image passes, so the generator's cheapest strategy is to ignore its label input and reproduce the overall data distribution.

open as a page

Why do GANs minimize -log D(G(z)) instead of the minimax form log(1 - D(G(z)))?

level: middleimportance: must knowfreq 58%

basics

~20 s

Early on, the discriminator rejects generated samples confidently, and log(1 - D(G(z))) is flat in that regime, so the generator receives almost no gradient. The non-saturating form -log D(G(z)) is steepest exactly where samples are being rejected.

open as a page

Why can a GAN's generator and discriminator loss curves not be read as training progress?

level: middleimportance: must knowfreq 62%

basics

~20 s

Each GAN player's loss is measured against an opponent that changes every step, so a falling or rising curve says who is currently ahead, not whether samples improved. Judge progress from samples and diversity, not from the losses.

open as a page

What does the Fréchet Inception Distance measure between real and generated images?

level: middleimportance: must knowfreq 66%

basics

~20 s

FID embeds real and generated images with a fixed pretrained classifier, fits one Gaussian to each set of feature vectors, and reports the Fréchet distance between those two Gaussians. Lower means the feature distributions are closer.

open as a page

In paired image-to-image translation, why add an L1 loss to the adversarial loss?

level: middleimportance: should knowfreq 44%

basics

~20 s

Pairs give a ground-truth target, and the L1 term pins the output to it — correct layout, colours and large-scale structure. The adversarial term then supplies the sharp local detail that a reconstruction loss alone averages into blur.

open as a page

How do precision and recall for generative models separate fidelity from coverage?

level: middleimportance: should knowfreq 36%

basics

~20 s

Precision is the share of generated samples falling inside the real data's feature manifold, which measures fidelity. Recall is the share of real samples falling inside the generated manifold, which measures coverage. Opposite failures cannot cancel.

open as a page

For a fixed GAN generator, what is the optimal discriminator and what does the generator then minimize?

level: seniorimportance: should knowfreq 42%

basics

~20 s

For a fixed generator, the optimal discriminator is D*(x) = p_data(x) / (p_data(x) + p_g(x)). Substituting it back turns the objective into a constant plus twice the Jensen-Shannon divergence between the two distributions, which is zero only when they are equal.

open as a page

What does switching a GAN to a Wasserstein critic with a gradient penalty actually fix?

level: seniorimportance: should knowfreq 45%

basics

~20 s

It fixes the vanishing signal from a discriminator that has already won. A Lipschitz-constrained critic estimates a distance between the real and generated distributions, so its gradients stay useful and its value tracks sample quality.

open as a page

Why can't you compare your FID number against the one reported in a paper?

level: seniorimportance: should knowfreq 44%

basics

~20 s

FID is only meaningful inside one fixed protocol. The number moves with sample count, the feature network's weights, the resizing applied before it, and which real split you score against. Change any of them and the comparison is void.

open as a page

How can a generator that memorises its training images still score an excellent FID?

level: seniorimportance: should knowfreq 28%

basics

~20 s

Because distributional scores only ask whether the generated distribution matches the real one, and a copy of the training set matches it exactly. Catching memorisation needs a separate nearest-neighbour audit against the training data, calibrated against a held-out baseline.

open as a page

With no paired training data for image translation, how do you decide between cycle consistency and buying pairs?

level: principalimportance: should knowfreq 30%

basics

~20 s

Decide on three things: whether one direction of the mapping can be simulated to fabricate pairs cheaply, whether the task needs geometry changed or only appearance, and what a plausible-but-wrong output costs. Cycle consistency constrains the mapping; it does not guarantee it is meaningful.

open as a page

In a class-conditional GAN, why prefer a projection discriminator over an auxiliary classifier?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

An auxiliary-classifier head rewards the generator for samples that are easy to classify, which pulls each class toward a few prototypical examples. A projection discriminator folds the label into the single adversarial score instead, so no such reward exists.

open as a page

Why is a GAN's discriminator never trained to convergence between generator updates?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Solving the inner maximization at every generator step is prohibitively expensive, and on a finite training set a converged discriminator overfits and saturates, leaving the generator with almost no gradient. Training instead alternates a few discriminator steps with one generator step.

open as a page

A spectrally normalized GAN trains stably at batch size 32 but destabilizes at 256 — how do you diagnose it?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Spectral normalization bounds how sharp the discriminator can be, not how the two players are balanced, and batch size changes that balance. Control for generator update count first, then log per-player gradient norms and held-out discriminator accuracy.

open as a page

How do you evaluate a generator of sensor time series when no standard feature extractor exists?

level: principalimportance: nice to knowfreq 18%

basics

~20 s

You build the yardstick yourself: train a domain encoder on real data and measure distribution distance in its features, back that with train-on-synthetic-test-on-real utility and physically meaningful summary statistics, and accept that the numbers are comparable only inside your own project.

open as a page