skip to content

Autoencoders, GANs and Generative Models

You will learn the pre-LLM generative families: what a VAE's reparameterization trick enables, why GAN training is unstable and how mode collapse shows up, and the intuition behind diffusion. Interviewers use GAN failure modes to test whether you can reason about adversarial training dynamics.

on this pageshow

explore

questions

page 2 of 2

Why does the pathwise gradient estimator usually beat the score-function estimator on variance?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Both estimators are unbiased, but the pathwise one differentiates the objective at the sample, so each draw carries directional information. The score-function estimator only multiplies a scalar value by a score vector, which is far noisier and needs a baseline.

open as a page

Why does a sentence VAE's KL term fall to zero while its decoder ignores the latent?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Posterior collapse: a decoder strong enough to model the sentence alone gains nothing from the code, so the encoder drives every posterior onto the prior. The KL cost vanishes and the decoder degenerates into an unconditional language model.

open as a page

With no paired training data for image translation, how do you decide between cycle consistency and buying pairs?

level: principalimportance: should knowfreq 30%

basics

~20 s

Decide on three things: whether one direction of the mapping can be simulated to fabricate pairs cheaply, whether the task needs geometry changed or only appearance, and what a plausible-but-wrong output costs. Cycle consistency constrains the mapping; it does not guarantee it is meaningful.

open as a page

How do you set beta in a beta-VAE when you need faithful reconstructions and interpretable factors?

level: principalimportance: should knowfreq 34%

basics

~20 s

Beta scales the KL term, so it picks an operating point on a rate-distortion curve, not a quality level. Fix the acceptance test first — a reconstruction budget plus a probe on known factors — then sweep and take the knee.

open as a page

Is a linear autoencoder trained with squared error equivalent to PCA?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

Not exactly. At its global optimum a linear autoencoder with a k-unit code spans the same subspace as the top k principal components, but its axes inside that subspace are an arbitrary rotation and rescaling — not orthonormal, not ordered by variance, not unique across runs.

open as a page

Why does a VAE's encoder emit log-variance rather than variance or standard deviation?

level: middleimportance: nice to knowfreq 42%

basics

~20 s

Log-variance is unconstrained, so any real number a linear output emits is legal, and exponentiating it always yields a positive variance. It also keeps tiny variances representable and drops straight into the Gaussian KL formula, which already contains a log.

open as a page

In a class-conditional GAN, why prefer a projection discriminator over an auxiliary classifier?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

An auxiliary-classifier head rewards the generator for samples that are easy to classify, which pulls each class toward a few prototypical examples. A projection discriminator folds the label into the single adversarial score instead, so no such reward exists.

open as a page

Why is a GAN's discriminator never trained to convergence between generator updates?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Solving the inner maximization at every generator step is prohibitively expensive, and on a finite training set a converged discriminator overfits and saturates, leaving the generator with almost no gradient. Training instead alternates a few discriminator steps with one generator step.

open as a page

A spectrally normalized GAN trains stably at batch size 32 but destabilizes at 256 — how do you diagnose it?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Spectral normalization bounds how sharp the discriminator can be, not how the two players are balanced, and batch size changes that balance. Control for generator update count first, then log per-player gradient norms and held-out discriminator accuracy.

open as a page

How does a contractive autoencoder's Jacobian penalty differ from training with added input noise?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

A contractive penalty is analytic: it shrinks the encoder's derivatives with respect to the input, in every direction at once. Input corruption chases the same robustness stochastically, over a finite radius, and for the whole encode-decode function.

open as a page

How do you evaluate a generator of sensor time series when no standard feature extractor exists?

level: principalimportance: nice to knowfreq 18%

basics

~20 s

You build the yardstick yourself: train a domain encoder on real data and measure distribution distance in its features, back that with train-on-synthetic-test-on-real utility and physically meaningful summary statistics, and accept that the numbers are comparable only inside your own project.

open as a page

In diffusion modelling, when is a fixed Gaussian forward process the wrong fit for your data?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Gaussian corruption assumes continuous, comparably scaled features. It suits smooth vector data such as robot action trajectories, and suits discrete tokens, hard-constrained quantities and heavy-tailed features badly - and every sample costs many network evaluations.

open as a page

When a flow beats a GAN on exact likelihood but its samples look worse, what do you conclude?

level: principalimportance: nice to knowfreq 28%

basics

~10 s

Likelihood and sample quality measure different things. Maximum likelihood minimises a mode-covering divergence that punishes missing data but not wasted mass, so a better likelihood is evidence of density fit, not of better-looking samples.

open as a page

When is distilling a diffusion sampler to four steps worth it over just cutting sampler steps?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Cut steps and change solver first - both are free and reversible. Distillation earns its cost only when a hard latency floor sits below what any training-free sampler reaches, and you accept narrower diversity plus a retraining stage per model version.

open as a page

Gumbel-Softmax or a vector-quantized codebook for a discrete latent: how do you choose?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Pick a relaxation when the code may be soft during training and the class count is modest; pick a quantized codebook when the forward pass must be discrete and a downstream model consumes the codes. Both are biased.

open as a page

showing 31–45 of 45