Autoencoders, GANs and Generative Models
You will learn the pre-LLM generative families: what a VAE's reparameterization trick enables, why GAN training is unstable and how mode collapse shows up, and the intuition behind diffusion. Interviewers use GAN failure modes to test whether you can reason about adversarial training dynamics.
on this pageshowhide
explore
- Latent-Variable Models17 questions
- Bottleneck and Reconstruction4 questions
- Denoising and Sparse Codes4 questions
- Variational Autoencoders5 questions
- Reparameterization Trick4 questions
- Generator and Discriminator17 questions
- Minimax Objective4 questions
- Mode Collapse and Instability4 questions
- Conditional Generation4 questions
- Sample Fidelity Metrics5 questions
- Diffusion and Likelihood Families11 questions
- Forward and Reverse Diffusion4 questions
- Autoregressive and Flow Models3 questions
- Noise Schedules and Samplers4 questions
questions
page 2 of 2Why does the pathwise gradient estimator usually beat the score-function estimator on variance?
basics
~20 sBoth estimators are unbiased, but the pathwise one differentiates the objective at the sample, so each draw carries directional information. The score-function estimator only multiplies a scalar value by a score vector, which is far noisier and needs a baseline.
Why does a sentence VAE's KL term fall to zero while its decoder ignores the latent?
basics
~20 sPosterior collapse: a decoder strong enough to model the sentence alone gains nothing from the code, so the encoder drives every posterior onto the prior. The KL cost vanishes and the decoder degenerates into an unconditional language model.
With no paired training data for image translation, how do you decide between cycle consistency and buying pairs?
basics
~20 sDecide on three things: whether one direction of the mapping can be simulated to fabricate pairs cheaply, whether the task needs geometry changed or only appearance, and what a plausible-but-wrong output costs. Cycle consistency constrains the mapping; it does not guarantee it is meaningful.
How do you set beta in a beta-VAE when you need faithful reconstructions and interpretable factors?
basics
~20 sBeta scales the KL term, so it picks an operating point on a rate-distortion curve, not a quality level. Fix the acceptance test first — a reconstruction budget plus a probe on known factors — then sweep and take the knee.
Is a linear autoencoder trained with squared error equivalent to PCA?
basics
~20 sNot exactly. At its global optimum a linear autoencoder with a k-unit code spans the same subspace as the top k principal components, but its axes inside that subspace are an arbitrary rotation and rescaling — not orthonormal, not ordered by variance, not unique across runs.
Why does a VAE's encoder emit log-variance rather than variance or standard deviation?
basics
~20 sLog-variance is unconstrained, so any real number a linear output emits is legal, and exponentiating it always yields a positive variance. It also keeps tiny variances representable and drops straight into the Gaussian KL formula, which already contains a log.
In a class-conditional GAN, why prefer a projection discriminator over an auxiliary classifier?
basics
~20 sAn auxiliary-classifier head rewards the generator for samples that are easy to classify, which pulls each class toward a few prototypical examples. A projection discriminator folds the label into the single adversarial score instead, so no such reward exists.
Why is a GAN's discriminator never trained to convergence between generator updates?
basics
~20 sSolving the inner maximization at every generator step is prohibitively expensive, and on a finite training set a converged discriminator overfits and saturates, leaving the generator with almost no gradient. Training instead alternates a few discriminator steps with one generator step.
A spectrally normalized GAN trains stably at batch size 32 but destabilizes at 256 — how do you diagnose it?
basics
~20 sSpectral normalization bounds how sharp the discriminator can be, not how the two players are balanced, and batch size changes that balance. Control for generator update count first, then log per-player gradient norms and held-out discriminator accuracy.
How does a contractive autoencoder's Jacobian penalty differ from training with added input noise?
basics
~20 sA contractive penalty is analytic: it shrinks the encoder's derivatives with respect to the input, in every direction at once. Input corruption chases the same robustness stochastically, over a finite radius, and for the whole encode-decode function.
How do you evaluate a generator of sensor time series when no standard feature extractor exists?
basics
~20 sYou build the yardstick yourself: train a domain encoder on real data and measure distribution distance in its features, back that with train-on-synthetic-test-on-real utility and physically meaningful summary statistics, and accept that the numbers are comparable only inside your own project.
In diffusion modelling, when is a fixed Gaussian forward process the wrong fit for your data?
basics
~20 sGaussian corruption assumes continuous, comparably scaled features. It suits smooth vector data such as robot action trajectories, and suits discrete tokens, hard-constrained quantities and heavy-tailed features badly - and every sample costs many network evaluations.
When a flow beats a GAN on exact likelihood but its samples look worse, what do you conclude?
basics
~10 sLikelihood and sample quality measure different things. Maximum likelihood minimises a mode-covering divergence that punishes missing data but not wasted mass, so a better likelihood is evidence of density fit, not of better-looking samples.
When is distilling a diffusion sampler to four steps worth it over just cutting sampler steps?
basics
~20 sCut steps and change solver first - both are free and reversible. Distillation earns its cost only when a hard latency floor sits below what any training-free sampler reaches, and you accept narrower diversity plus a retraining stage per model version.
Gumbel-Softmax or a vector-quantized codebook for a discrete latent: how do you choose?
basics
~20 sPick a relaxation when the code may be soft during training and the class count is modest; pick a quantized codebook when the forward pass must be discrete and a downstream model consumes the codes. Both are biased.
showing 31–45 of 45