skip to content

Why does the density of Y = e^X carry a factor of 1/y when X is normal?

level: middleimportance: should knowfreq 38%

answer

  1. a density is probability per unit length
  2. the transformation stretches the axis
  3. go through the CDF and differentiate
  4. chain rule on ln y
  5. derivative of the inverse, in absolute value

basics

~20 s

A density must be rescaled by the derivative of the inverse map. The inverse of Y = e^X is X = ln y, whose derivative is 1/y, so f_Y(y) = f_X(ln y)/y for y > 0.

solid answer

~50 s

A density is probability per unit length, so when a transformation stretches the axis the density must be diluted by the same factor. For a strictly monotone, differentiable map `Y = g(X)` the rule is `f_Y(y) = f_X(g_inverse(y)) * |d/dy g_inverse(y)|`, where the derivative factor is the Jacobian. With `Y = e^X` the inverse is `X = ln y` with derivative `1/y`, so the density becomes `f_Y(y) = f_X(ln y) / y` on `y > 0` — for a normal `f_X` that is exactly the log-normal density. The quickest way to see it without memorising the rule is the CDF route: `P(Y <= y) = P(e^X <= y) = P(X <= ln y) = F_X(ln y)`, and differentiating in `y` produces the `1/y` by the chain rule. Skipping the factor gives something that does not integrate to 1, which is the usual sign the Jacobian was dropped.

go deeper

for a junior

Recall that transforming a continuous variable is not just substitution: the density picks up the derivative of the inverse map, which for Y = e^X is 1/y.

for a middle

Be ready to derive it on the spot through the CDF — write P(e^X <= y) as P(X <= ln y), differentiate, and point out that the chain rule is where the 1/y comes from.

for a senior

Show the reflexes: check the transformed density integrates to 1, get the support right, use the absolute value for decreasing maps, and split a non-monotone map into branches before applying the rule.

for a principal

Own the wider point when reviewing others' work: quantities reported on a transformed scale do not translate back by substitution alone, and be able to explain which summaries survive a monotone transformation and which do not.

## Densities are per-unit-length, and transformations change length A probability density is not a probability. `f_X(x)` is a rate — probability per unit of `x` — so `f_X(x) * dx` is the probability of a tiny interval. When you apply a transformation `Y = g(X)`, the probability sitting in a tiny interval must be preserved, but the *width* of that interval changes. If `g` stretches a small interval, the same probability is now spread over more room, so the density there must be smaller. The correction factor that accounts for this stretching is the Jacobian. ## The rule for a monotone map Let `g` be strictly monotone and differentiable, with inverse `h = g_inverse`. Then `f_Y(y) = f_X(h(y)) * |h'(y)|`. Two pieces: `f_X(h(y))` asks "which `x` produced this `y`, and how dense was `X` there?", and `|h'(y)|` converts from `x`-length to `y`-length. The absolute value matters because a decreasing `g` has a negative derivative, and a density can never be negative — for a decreasing map the CDF derivation produces a minus sign that the absolute value absorbs. ## The exponential map, step by step Take `Y = e^X` with `X` normal, mean `m`, variance `s^2`. Then `h(y) = ln y`, defined for `y > 0`, and `h'(y) = 1/y`. So `f_Y(y) = f_X(ln y) * (1/y) = (1 / (y * s * sqrt(2*pi))) * exp(-((ln y) - m)^2 / (2 * s^2))` for `y > 0`, and zero otherwise. That is the log-normal density, and the `1/y` is not decoration. The exponential map crushes the far-left tail of `X` into a tiny sliver just above zero and blows up the far-right tail into an enormous stretch of the positive axis. Near `y = 0` the same probability is packed into a much narrower interval, so the density there is *multiplied* by a large `1/y`; far out to the right it is divided down. Note also that `Y` is automatically supported only on the positive numbers, because the exponential of a real number is always positive. ## The CDF route, which needs nothing memorised If you cannot recall the Jacobian rule, derive it. Because the exponential is increasing, `F_Y(y) = P(Y <= y) = P(e^X <= y) = P(X <= ln y) = F_X(ln y)` for `y > 0`. Differentiate both sides with respect to `y`; the chain rule contributes the derivative of `ln y`, which is `1/y`: `f_Y(y) = f_X(ln y) * (1/y)`. The Jacobian is exactly the chain-rule factor. This route also handles decreasing maps gracefully: there the event flips to `P(X >= h(y))`, the derivative comes out negative, and the sign is absorbed by the absolute value. ## The check that catches a dropped Jacobian Any candidate density must integrate to 1. Substituting `ln y` into a normal density without the `1/y` gives a function whose integral over `y > 0` is not 1, so the omission is detectable without any theory. Making that check a reflex is worth more than memorising the formula. ## Non-monotone maps The rule as stated needs invertibility. For a map like `Y = X^2`, two values of `X` produce each positive `y`, so you split the domain into monotone branches and add their contributions: `f_Y(y) = [f_X(sqrt(y)) + f_X(-sqrt(y))] * (1 / (2 * sqrt(y)))` for `y > 0`, where `1 / (2*sqrt(y))` is the Jacobian of each branch. The general recipe is always the same: go back to the CDF, describe the event `{g(X) <= y}` as a union of intervals in `x`, and differentiate. ## Discrete variables need no Jacobian Probability mass functions assign probability to points, not to length, so nothing is stretched. For a discrete `X`, `P(Y = y) = sum of P(X = x)` over every `x` with `g(x) = y`. Applying a derivative factor to a PMF is a category error; the Jacobian exists only because continuous probability is spread over intervals. ## A transformation worth knowing One consequence of the change-of-variables idea deserves separate billing. If `X` is continuous with strictly increasing CDF `F`, then `U = F(X)` is Uniform on `(0,1)` — the probability integral transform. Running it backwards, if `U` is Uniform on `(0,1)` then `F_inverse(U)` has distribution `F`, which is the basis of inverse-transform sampling: turn uniform numbers into draws from any distribution whose quantile function you can evaluate. ## Why interviewers ask It tests whether you know what a density *is*. Candidates who think of `f(x)` as "the probability of x" substitute and stop; candidates who think of it as probability per unit length know a stretching factor must appear, and can produce it from the CDF in two lines even under pressure.

  • What changes for a non-monotone transformation such as Y = X^2?
    The map is no longer invertible, so you split it into monotone branches and add their contributions: f_Y(y) = [f_X(sqrt(y)) + f_X(-sqrt(y))] / (2*sqrt(y)) for y > 0. The safe general recipe is to return to the CDF, write the event that X^2 is at most y as -sqrt(y) <= X <= sqrt(y), and differentiate.
  • Do discrete random variables need a Jacobian when transformed?
    No. A probability mass function assigns probability to individual points rather than per unit length, so nothing is stretched. You simply collect the mass: P(Y = y) is the sum of P(X = x) over all x mapping to y. Applying a derivative factor to a PMF is a category error, and if the map is many-to-one the summing is all the work there is.
  • What is the distribution of U = F(X) when F is the continuous, strictly increasing CDF of X?
    Uniform on (0,1). P(F(X) <= u) = P(X <= F_inverse(u)) = F(F_inverse(u)) = u, which is the uniform CDF. This is the probability integral transform, and running it backwards gives inverse-transform sampling: if U is uniform, F_inverse(U) has distribution F, so uniform numbers can be turned into draws from any distribution with a computable quantile function.
  • How would you catch a dropped Jacobian factor without redoing the derivation?
    Integrate the candidate density and check it comes to 1. Substituting the inverse map into the original density but omitting the derivative factor almost always produces total mass different from 1, and it may also live on the wrong support. That integral check is a cheap reflex that catches the error regardless of which transformation was applied.

Population density on a map redrawn with a stretched scale: the same people now cover more paper, so the people-per-square-inch figure must be divided by the stretch factor.

saying these in an interview costs you the question

  • Substitutes ln y into the normal density and stops
  • Drops the absolute value for a decreasing transformation
  • Treats a density value as a probability
  • Applies a Jacobian factor to a discrete PMF
  • Uses the monotone formula on a map like squaring

context