skip to content

Autoencoders, GANs and Generative Models

You will learn the pre-LLM generative families: what a VAE's reparameterization trick enables, why GAN training is unstable and how mode collapse shows up, and the intuition behind diffusion. Interviewers use GAN failure modes to test whether you can reason about adversarial training dynamics.

on this pageshow

explore

questions

page 1 of 2

In a GAN's minimax objective, what is the discriminator maximizing and the generator minimizing?

level: juniorimportance: must knowfreq 72%

answer

  1. two networks, one shared score
  2. binary cross-entropy, real versus fake
  3. outer minimize, inner maximize
  4. only one term contains the generator
  5. generator learns through the judge's gradients

basics

~20 s

The discriminator maximizes the log-probability of labelling real data real and generated data fake. The generator minimizes that same quantity, pushing the discriminator toward calling its samples real. One shared value function, optimized in opposite directions.

solid answer

~40 s

A GAN trains two networks against one shared value function: `V(D, G) = E_x~p_data[log D(x)] + E_z~p_z[log(1 - D(G(z)))]`. The discriminator D outputs the probability that its input is real and climbs V — it wants `log D(x)` high on real samples and `log(1 - D(G(z)))` high on generated ones, which is exactly binary cross-entropy for a real-versus-fake classifier. The generator G appears only in the second term and descends V, so it wants `D(G(z))` pushed toward 1. G never touches real data; its entire learning signal is gradients passed back through D. Writing the whole thing as `min_G max_D V(D, G)` says the outer player commits to a generator while the inner player is free to respond as well as it can.

go deeper

for a junior

Be ready to state the two players and their opposite directions in one breath, and to say which term the generator appears in. Knowing that the generator learns only through the discriminator's gradients is the piece most candidates miss.

for a middle

Explain the value function term by term and connect it to binary cross-entropy for a real-versus-fake classifier. An interviewer expects you to say what expectation each term is taken over and where the generator's gradient physically comes from.

for a senior

Show you know why this is a game and not a loss: the generator's objective is defined by a discriminator that is itself moving, so nothing decreases monotonically and progress cannot be read off a single number.

for a principal

Own the framing decision. Be able to argue when an adversarial objective is worth its instability at all versus a likelihood-based generative model, and what it costs a team in tuning time and evaluation infrastructure.

## The setup A GAN is an *implicit* generative model. It never writes down a density for the data. Instead a generator network `G` maps a noise vector `z`, drawn from a fixed simple prior `p_z` (a standard Gaussian, say), to a sample `G(z)`. Pushing the prior through `G` induces some distribution over samples, written `p_g`. Training's goal is to make `p_g` match the data distribution `p_data`, and the whole trick of the adversarial framework is that you can do this using only *samples* from both — no density, no likelihood. The measuring instrument is a second network, the discriminator `D`, which takes a sample and outputs a number in (0, 1) interpreted as "probability this came from the real data". ## The shared value function Both networks are scored by one function: `V(D, G) = E_x~p_data[log D(x)] + E_z~p_z[log(1 - D(G(z)))]` Read each term separately. The first is large when `D` assigns high probability to real samples being real. The second is large when `D` assigns low probability to generated samples being real, since `1 - D(G(z))` is then close to 1. Add them and you have, up to a factor and a sign, the binary cross-entropy of a classifier trained on a balanced mixture of real examples labelled 1 and generated examples labelled 0. That is why the logs are there: they are the Bernoulli log-likelihood of that classification problem. ## Who moves which way - **Discriminator: maximize V.** It is doing ordinary supervised classification against the current generator. Its gradient comes from both terms. - **Generator: minimize V.** `G` appears only in the second term, `log(1 - D(G(z)))`. Minimizing it means making `D(G(z))` large — that is, fooling `D`. The generator's gradient flows backwards through the discriminator into `G`'s parameters, which is why `D` must be differentiable and why `G` never needs to see a real example directly. The compact notation is `min_G max_D V(D, G)`. The ordering matters conceptually: the inner maximization defines, for any given generator, how well the best possible judge can separate its samples from real ones, and the outer minimization then looks for the generator that even the best judge cannot beat. ## What equilibrium means If the generator's distribution ever exactly equalled the data distribution, no discriminator could do better than chance, and it would output 1/2 on every input. That is the fixed point the game aims at. Note what this does *not* say: it makes no claim that a particular sample is good, only that the two distributions are indistinguishable to the judge on offer. ## Why this is not an ordinary loss This is the part interviewers actually probe. In supervised learning you descend one fixed scalar function of your parameters, and a falling loss means progress. Here there are two parameter sets moving against each other on one surface, and neither player's objective is stationary: the generator's loss landscape is defined by the current discriminator and changes the moment the discriminator updates. Practically that means: - Training is **alternating** (or simultaneous) gradient steps on two objectives, not descent on one. - No scalar is guaranteed to decrease monotonically; the pair can circle an equilibrium. - "The loss went down" is not by itself evidence that the generator got better, because it may simply mean the discriminator got worse. ## Common misreadings A weak answer describes `G` as being trained to reconstruct real images, or `D` as scoring image quality on some open-ended scale. Neither is right. `G` has no target image and no reconstruction term; its only teacher is `D`. And `D` is a plain binary classifier — its output is a probability of the label "real", not a quality score. ## What to say in an interview Write the value function, name what each expectation is over, say which player climbs and which descends, point out that only the second term contains `G`, and then add the sentence that separates a memoriser from someone who has trained one: because the two objectives are coupled, this is a game, and its dynamics — not just its optimum — are what you spend your time managing.

  • Does the generator ever see real training data during its own update?
    No. The generator's loss depends on real data only through the discriminator's parameters. It samples noise, produces fakes, and receives gradients backpropagated through the discriminator. That is why the discriminator must be differentiable, and why a discriminator that has stopped learning anything useful leaves the generator with no teacher at all.
  • Why is the objective written with logs rather than raw probabilities?
    Because maximizing it is maximum likelihood for the discriminator viewed as a binary classifier over a balanced mixture of real and generated samples: the log terms are the Bernoulli log-likelihood, and summing logs is the log of a product of per-sample likelihoods. The log also keeps the classifier's gradients well behaved through the usual cross-entropy-with-sigmoid form.
  • What does the min-max ordering imply about who is assumed to move first?
    The generator is the outer player: it must pick a distribution that survives the best response of an inner player who sees that choice. Real training does not honour that ordering — both players take small alternating steps — which is one reason the theoretical picture and the observed dynamics diverge.

saying these in an interview costs you the question

  • Says the generator is trained on real samples with a reconstruction loss
  • Claims both networks minimize the same loss
  • Describes the discriminator as scoring image quality rather than classifying real versus fake
  • Thinks the generator compares its output to the nearest real example
  • Treats a falling generator loss as proof the samples improved

context

open as a page

What is mode collapse in GAN training, and how do you spot it in generated samples?

level: juniorimportance: must knowfreq 80%

basics

~20 s

Mode collapse is when a GAN generator maps many different noise vectors to a few nearly identical outputs, covering only part of the real data. You spot it by decoding a fixed batch of noise vectors and seeing the samples repeat.

open as a page

Why does a plain autoencoder need a bottleneck narrower than its input?

level: juniorimportance: must knowfreq 70%

basics

~20 s

An autoencoder is trained to copy its input, so a wide enough code can simply learn the identity map: near-zero loss, nothing learned. A code narrower than the input makes exact copying impossible and forces the encoder to keep only rebuildable structure.

open as a page

Why does a denoising autoencoder corrupt its input but score the reconstruction against the clean original?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Scoring against the clean original means copying the input no longer wins. The network has to use structure shared across the data to remove the corruption, so the code ends up describing that structure instead of the input itself.

open as a page

In a conditional GAN, why does the discriminator receive the class label too?

level: middleimportance: must knowfreq 62%

basics

~20 s

Only the discriminator can create pressure to obey the label. If it judges images alone, any realistic image passes, so the generator's cheapest strategy is to ignore its label input and reproduce the overall data distribution.

open as a page

Why do GANs minimize -log D(G(z)) instead of the minimax form log(1 - D(G(z)))?

level: middleimportance: must knowfreq 58%

basics

~20 s

Early on, the discriminator rejects generated samples confidently, and log(1 - D(G(z))) is flat in that regime, so the generator receives almost no gradient. The non-saturating form -log D(G(z)) is steepest exactly where samples are being rejected.

open as a page

Why can a GAN's generator and discriminator loss curves not be read as training progress?

level: middleimportance: must knowfreq 62%

basics

~20 s

Each GAN player's loss is measured against an opponent that changes every step, so a falling or rising curve says who is currently ahead, not whether samples improved. Judge progress from samples and diversity, not from the losses.

open as a page

What does the Fréchet Inception Distance measure between real and generated images?

level: middleimportance: must knowfreq 66%

basics

~20 s

FID embeds real and generated images with a fixed pretrained classifier, fits one Gaussian to each set of feature vectors, and reports the Fréchet distance between those two Gaussians. Lower means the feature distributions are closer.

open as a page

How is a diffusion model trained on one step without simulating the whole noising chain?

level: middleimportance: must knowfreq 72%

basics

~20 s

The forward corruption is fixed and Gaussian, so any step is one shot: x_t = sqrt(abar_t)*x_0 + sqrt(1-abar_t)*eps, with abar_t the running product of (1-beta). Training draws a random t, builds x_t, and regresses eps.

open as a page

Why can an autoregressive image model train in parallel but not sample in parallel?

level: middleimportance: must knowfreq 55%

basics

~20 s

An autoregressive model factorises the joint into conditionals over an ordering. Training scores all of them in one masked pass because the ground truth is already there; sampling must draw each value before the next, one pass per dimension.

open as a page

Why can DDIM sample a diffusion model in 20 steps when ancestral DDPM sampling needed 1000?

level: middleimportance: must knowfreq 72%

basics

~20 s

DDIM reverses the diffusion with a non-Markovian, noise-free update that matches the same training marginals, so one trained network can be run on any subsequence of timesteps. Because the update is deterministic, dropping steps degrades gracefully instead of breaking.

open as a page

How does the reparameterization trick get a gradient through a random latent sample?

level: middleimportance: must knowfreq 72%

basics

~20 s

The reparameterization trick moves randomness off the parameter path: draw fixed noise eps, then set z = mu + sigma * eps. That makes z a differentiable function of mu and sigma, so gradients reach the encoder.

open as a page

What two terms make up a variational autoencoder's training loss?

level: middleimportance: must knowfreq 76%

basics

~20 s

A VAE's loss adds a reconstruction term, penalising how badly the decoder rebuilds the input from its latent code, to a KL term that pulls each input's encoded Gaussian toward a standard-normal prior. The two pull against each other.

open as a page

Why does decoding a random point in a plain autoencoder's 16-D code space give an off-manifold artefact?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Reconstruction loss constrains the decoder only at the codes the encoder actually produced for training data. Everywhere else the code space is unconstrained, so a randomly drawn point falls in an unallocated region and the decoder extrapolates into nonsense.

open as a page

In paired image-to-image translation, why add an L1 loss to the adversarial loss?

level: middleimportance: should knowfreq 44%

basics

~20 s

Pairs give a ground-truth target, and the L1 term pins the output to it — correct layout, colours and large-scale structure. The adversarial term then supplies the sharp local detail that a reconstruction loss alone averages into blur.

open as a page

How do precision and recall for generative models separate fidelity from coverage?

level: middleimportance: should knowfreq 36%

basics

~20 s

Precision is the share of generated samples falling inside the real data's feature manifold, which measures fidelity. Recall is the share of real samples falling inside the generated manifold, which measures coverage. Opposite failures cannot cancel.

open as a page

Why did diffusion models move from a linear noise schedule to a cosine one?

level: middleimportance: should knowfreq 55%

basics

~20 s

A linear beta schedule destroys nearly all signal two-thirds of the way through the chain, so the final third of the reverse steps teaches the model almost nothing. A cosine schedule lets signal-to-noise fall gradually to the end.

open as a page

For an autoencoder's reconstruction loss, when do you use cross-entropy over squared error?

level: middleimportance: should knowfreq 50%

basics

~20 s

Use Bernoulli cross-entropy when every output lies in [0,1] and the decoder emits it through a sigmoid, such as scaled pixels. Use squared error for unbounded real-valued outputs. The choice declares what output distribution you are assuming.

open as a page

What does an L1 sparsity penalty on an autoencoder's hidden activations buy over a narrow bottleneck?

level: middleimportance: should knowfreq 42%

basics

~20 s

An activation penalty limits how many units may fire per example rather than how many units exist. The layer can be wider than the input without collapsing into a copy, and each unit specialises on one recurring pattern.

open as a page

Why does a VAE with a pixel-wise squared-error decoder produce blurry image samples?

level: middleimportance: should knowfreq 58%

basics

~20 s

Squared error is the log-likelihood of a fixed-variance Gaussian output, and its minimiser is the mean over every image consistent with the code. Averaging misaligned edges and textures cancels fine detail, which is what blur is.

open as a page

For a fixed GAN generator, what is the optimal discriminator and what does the generator then minimize?

level: seniorimportance: should knowfreq 42%

basics

~20 s

For a fixed generator, the optimal discriminator is D*(x) = p_data(x) / (p_data(x) + p_g(x)). Substituting it back turns the objective into a constant plus twice the Jensen-Shannon divergence between the two distributions, which is zero only when they are equal.

open as a page

What does switching a GAN to a Wasserstein critic with a gradient penalty actually fix?

level: seniorimportance: should knowfreq 45%

basics

~20 s

It fixes the vanishing signal from a discriminator that has already won. A Lipschitz-constrained critic estimates a distance between the real and generated distributions, so its gradients stay useful and its value tracks sample quality.

open as a page

Why can't you compare your FID number against the one reported in a paper?

level: seniorimportance: should knowfreq 44%

basics

~20 s

FID is only meaningful inside one fixed protocol. The number moves with sample count, the feature network's weights, the resizing applied before it, and which real split you score against. Change any of them and the comparison is void.

open as a page

How can a generator that memorises its training images still score an excellent FID?

level: seniorimportance: should knowfreq 28%

basics

~20 s

Because distributional scores only ask whether the generated distribution matches the real one, and a copy of the training set matches it exactly. Catching memorisation needs a separate nearest-neighbour audit against the training data, calibrated against a held-out baseline.

open as a page

How does classifier-free guidance steer a diffusion model without training a classifier?

level: seniorimportance: should knowfreq 58%

basics

~20 s

One network learns both conditional and unconditional prediction, because the condition is replaced by a learned null token on roughly ten percent of training examples. Sampling evaluates it twice per step and extrapolates along the difference between the two predictions.

open as a page

In a diffusion model, why can't the reverse go from pure noise to data in one step?

level: seniorimportance: should knowfreq 55%

basics

~20 s

The forward jump is closed-form because it conditions on the clean sample. Backwards, many clean samples explain one noisy one, so the reverse distribution is multimodal; a single pass returns only its blurry average. Small steps keep each reverse move Gaussian.

open as a page

Why must a normalizing flow be invertible with a cheap log-determinant Jacobian?

level: seniorimportance: should knowfreq 35%

basics

~20 s

A flow's exact likelihood comes from the change-of-variables formula, which needs the inverse map and the log-determinant of its Jacobian. A general determinant costs cubic time, so flow layers are built to have a triangular Jacobian.

open as a page

In diffusion training, what does v-prediction fix that epsilon-prediction breaks at high noise?

level: seniorimportance: should knowfreq 35%

basics

~20 s

At the noisiest steps the input is almost pure noise, so predicting the added noise is nearly trivial while small errors explode when converted back to an image. v-prediction blends the noise and image targets, staying informative at both ends.

open as a page

How would you turn an autoencoder's reconstruction error into a machine anomaly score?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Train the autoencoder on normal data only, score each window by its reconstruction error, and threshold at a high quantile of the error measured on held-out normal data. Keep the code narrow enough that the model cannot rebuild behaviour it never saw.

open as a page

How does a Gumbel-Softmax relaxation make a categorical latent trainable by gradients?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Gumbel-Softmax adds independent Gumbel noise to the logits and replaces the argmax with a temperature-scaled softmax. The noise is parameter-free, so gradients flow to the logits; the price is a soft sample and a biased gradient.

open as a page

showing 1–30 of 45