Why must a normalizing flow be invertible with a cheap log-determinant Jacobian?
answer
- density is mass per unit volume
- two terms: base log-density plus a correction
- a general determinant costs cubic time
- copy half, transform the other half
- triangular Jacobian makes log-det a sum
basics
~20 sA flow's exact likelihood comes from the change-of-variables formula, which needs the inverse map and the log-determinant of its Jacobian. A general determinant costs cubic time, so flow layers are built to have a triangular Jacobian.
solid answer
~50 sA flow defines a bijection `z = g(x)` onto a simple base density and evaluates `log p_x(x) = log p_z(g(x)) + log|det J_g(x)|`. Both terms are mandatory: the first says how probable the point is in latent space, the second corrects for volume the map stretched or squeezed, without which the density would not integrate to one. That forces two constraints. The map must be invertible, so no layer may throw information away — no pooling, no bottleneck, no non-injective activation. And the log-determinant must be computable cheaply, since a general `d x d` determinant is O(d^3). Affine coupling solves both: split the vector, copy one half through untouched, scale and shift the other half by functions of the copied half. The Jacobian is triangular, so its log-determinant is just the sum of the log scales, and inversion never has to invert the scale and shift networks themselves.
code
python · 25 linesimport math
# One affine coupling layer on a 4-dim vector.
# First half is copied through; second half is scaled and shifted
# by functions of the first half.
def scale_shift(a):
s = [0.5 * a[0], -0.2 * a[1]] # log-scales
t = [1.0 * a[1], 0.3 * a[0]] # shifts
return s, t
def forward(x):
a, b = x[:2], x[2:]
s, t = scale_shift(a)
y = [b[i] * math.exp(s[i]) + t[i] for i in range(2)]
logdet = sum(s) # triangular Jacobian: just the log-scales
return a + y, logdet
def inverse(y):
a, c = y[:2], y[2:]
s, t = scale_shift(a) # same nets, never inverted
return a + [(c[i] - t[i]) * math.exp(-s[i]) for i in range(2)]
z, logdet = forward([1.0, 2.0, 3.0, 4.0])
print(z, logdet) # transformed point and its exact log|det J|
print(inverse(z)) # recovers [1.0, 2.0, 3.0, 4.0]go deeper
Be ready to write the change-of-variables formula and say in words why a second term is needed: stretching space spreads the same probability mass thinner.
Explain why a general determinant is unaffordable and how coupling layers make it a sum, including which block of the Jacobian is zero and why.
Show the practical consequences you would hit building one: no bottleneck, full-width activations, restricted layer choices, and the resulting memory profile at image resolution.
Frame it as a family-level bet: you are trading architectural freedom for an exact, comparable likelihood, and you should be able to say when that exchange is worth making.
## The change-of-variables formula A flow assumes the data was produced by pushing a simple latent through an invertible map. Let `z ~ p_z` be a standard Gaussian of the same dimensionality as `x`, let `f` be an invertible differentiable map with `x = f(z)`, and let `g = f^{-1}` so `z = g(x)`. Then `log p_x(x) = log p_z(g(x)) + log|det J_g(x)|` where `J_g(x)` is the Jacobian matrix of `g` at `x`. Equivalently, in the generative direction, `log p_x(x) = log p_z(z) - log|det J_f(z)|`. The intuition for the second term: probability density is mass per unit volume. If the map locally expands a region by a factor, the same probability mass is spread over more volume, so the density there must be smaller. The absolute determinant of the Jacobian *is* that local volume-change factor. Drop the term and you have a function that is not a normalised density at all. ## Constraint one: strict invertibility Because the formula evaluates `g(x)` for real data, every layer must be exactly invertible. This is far more restrictive than it sounds: - No max pooling, no strided reduction, no discarding of channels. - No non-injective activation. A rectified-linear unit maps every negative input to zero and is not invertible; flows use monotone or explicitly-inverted transformations instead. - **No bottleneck at all.** The latent has exactly the dimensionality of the data. An autoencoder is free to compress a `3 x 256 x 256` image into a few hundred numbers; a flow on the same image carries a 196,608-dimensional latent through every layer. Every activation at full width must also be kept or recomputed for the backward pass, so the memory bill per layer is set by the data, not by a design choice. This is a first-order practical reason flows lost ground at high resolution. ## Constraint two: a tractable log-determinant The Jacobian of an arbitrary layer over `d` dimensions is a dense `d x d` matrix, and its determinant costs O(d^3). At `d = 196,608` that is not a slow computation, it is an impossible one — and it is needed on every layer, for every example, in every training step. So flow layers are designed backwards from the determinant. **Affine coupling** is the canonical answer. Split the vector into two halves `x_a` and `x_b`. Then - `y_a = x_a` (copied through unchanged) - `y_b = x_b * exp(s(x_a)) + t(x_a)` where `s` and `t` are ordinary neural networks of any complexity. Three consequences fall out: 1. **The Jacobian is triangular.** `y_a` does not depend on `x_b` at all, so the off-diagonal block is zero. The determinant of a triangular matrix is the product of its diagonal, and the diagonal here is `exp(s(x_a))` on the transformed half and ones on the copied half. So `log|det J| = sum_j s_j(x_a)` — a plain sum of the scale network's outputs, effectively free. 2. **Inversion is closed-form and cheap.** Given `y`, recover `x_a = y_a`, evaluate `s(x_a)` and `t(x_a)` again, and set `x_b = (y_b - t(x_a)) * exp(-s(x_a))`. Note what did *not* happen: `s` and `t` were never inverted. They are always evaluated on the half that passes through untouched, so they may be arbitrarily deep and non-invertible. 3. **One layer is weak, so you stack them and permute.** A single coupling leaves half the vector unchanged. Real flows alternate which half is transformed, or interleave learned invertible channel mixing, so that every dimension is eventually transformed conditioned on every other. Other designs trade differently — autoregressive-style flows give a triangular Jacobian by conditioning dimension `i` on dimensions `< i`, which makes one direction of the map fast and the other serial. The design question is always the same: which of density evaluation and sampling do you want to be the cheap direction. ## The exactness payoff and its price What you buy is a density that is exact and comparable, with no bound and no adversarial proxy, and a latent space you can traverse and interpolate in. What you pay is architectural freedom: every layer is chosen from the small set that is invertible with a cheap determinant, dimensionality can never shrink, and expressive power per parameter is therefore lower than an unconstrained network of the same size. A candidate who can articulate that constraint-versus-exactness trade is answering the question; one who only recites the formula is not. ## What interviewers probe Expect follow-ups on why the determinant term exists at all (normalisation), on why a rectified-linear unit disqualifies a layer (non-injective), and on whether the coupling networks themselves need to be invertible (they do not, and knowing why is the tell that the candidate has actually understood the construction).
- Do the scale and shift networks inside a coupling layer need to be invertible?No, and that is the whole trick. They are only ever evaluated on the half of the vector that passes through unchanged, which is identical in the forward and inverse directions. So you recompute their outputs during inversion rather than undoing them, and they can be arbitrarily deep, non-monotone networks.
- What does the dimension-preserving constraint cost you in practice?Memory and expressiveness. A `3 x 256 x 256` image forces a 196,608-dimensional latent carried at full width through every layer, so activation memory scales with data size rather than with a chosen code size. And because every layer must come from the invertible, cheap-determinant family, you get less modelling power per parameter than an unconstrained network.
- Why does a single coupling layer need to be stacked with permutations?One coupling leaves half the dimensions untouched, so alone it can only model a very restricted family. Stacking layers that alternate which half is transformed — or that interleave an invertible channel mixing — lets every dimension eventually be transformed conditioned on every other, while each individual layer keeps its triangular Jacobian.
Repackaging the same amount of dough into a longer, thinner strip does not create more dough; the log-determinant is the bookkeeping that says how much thinner the density got where the map stretched.
saying these in an interview costs you the question
- Omits the log-determinant term from the density
- Says the coupling networks must themselves be invertible
- Claims a flow can compress to a lower-dimensional latent
- Thinks any deep network can be used as a flow layer
- Believes the determinant is computed by brute force each step