skip to content

Session durations follow a log-normal distribution: why does the mean sit above the median?

level: seniorimportance: should knowfreq 45%

answer

  1. the log is the normal one
  2. quantiles survive a monotone transform
  3. averaging does not survive it
  4. median exp(mu), mean exp(mu + sigma^2/2)
  5. ratio grows with log-scale spread

basics

~20 s

A log-normal variable is the exponential of a normal one, so its right tail is long. The median is exp(mu) while the mean is exp(mu + sigma squared over 2), which is always larger: rare long sessions lift the average.

solid answer

~40 s

By definition, a variable X is log-normal when `log(X)` is normal with mean `mu` and standard deviation `sigma`. Because the exponential function is strictly increasing, the median passes through unchanged: `median(X) = exp(mu)`. The mean does not, because averaging happens on the original scale where the upper tail is stretched: `E[X] = exp(mu + sigma^2 / 2)`. Their ratio is `exp(sigma^2 / 2)`, which is greater than 1 for any positive sigma and grows fast — with `sigma = 1` the mean is about 1.65 times the median. Watch the parameter trap: mu and sigma describe the log of the variable, not the variable itself. The practical consequence for session durations is that the average session is not a typical session, so I would report the median and a few percentiles alongside any mean.

go deeper

for a junior

Be ready to say that a log-normal variable is one whose logarithm is normal, that it is positive and right-skewed, and that its mean exceeds its median.

for a middle

Expect to explain why a monotone transformation preserves the median but not the mean, and to quote exp(mu) and exp(mu + sigma^2/2) correctly.

for a senior

Show you can separate the summary from the question: mean for totals and capacity, median and percentiles for typical experience, and no symmetric error bars on a skewed metric.

for a principal

Own the reporting standard: dashboards and SLOs that average a multiplicative quantity mislead by construction, so argue for percentile-based targets before the metric definition becomes organisational habit.

## Definition A positive random variable X is **log-normal** when its logarithm is normally distributed: ``` log(X) ~ Normal(mu, sigma^2) equivalently X = exp(Y), Y ~ Normal(mu, sigma^2) ``` The two parameters live on the log scale. `mu` is the mean of `log(X)`, and `sigma` is the standard deviation of `log(X)` — neither is the mean or standard deviation of X itself. That single point is the most frequently botched fact about the family, and getting it right in the first sentence signals you have actually used it. ## Why the median survives the exponential and the mean does not The median is a quantile, and quantiles are preserved by any strictly increasing transformation. Half of Y lies below mu, so half of `X = exp(Y)` lies below `exp(mu)`: ``` median(X) = exp(mu) ``` The mean is not a quantile — it is an average, and averaging does not commute with a nonlinear transformation. Because `exp` is convex, it stretches large values much more than small ones, so the mean lands strictly above `exp(mu)`: ``` E[X] = exp(mu + sigma^2 / 2) Var(X) = (exp(sigma^2) - 1) * exp(2*mu + sigma^2) mode = exp(mu - sigma^2) ``` Ordering the three centres gives the textbook right-skew signature: ``` mode < median < mean exp(mu - sigma^2) < exp(mu) < exp(mu + sigma^2/2) ``` ## How big is the gap? The mean-to-median ratio depends only on sigma: ``` E[X] / median(X) = exp(sigma^2 / 2) ``` - `sigma = 0.5` -> ratio `exp(0.125) ≈ 1.13`, a mild 13% gap. - `sigma = 1.0` -> ratio `exp(0.5) ≈ 1.65`. - `sigma = 2.0` -> ratio `exp(2) ≈ 7.4`, so the mean is over seven times the typical value. Spread on the log scale therefore translates into a multiplicative gap on the original scale. This is why average duration and typical duration can disagree wildly for the same population without anyone making an arithmetic error. ## Why durations and sizes land in this family A normal distribution arises when many small independent effects **add**. A log-normal arises when many small independent effects **multiply**, because multiplication on the original scale is addition on the log scale. Session length is plausibly multiplicative: a session that runs long tends to do so because several factors each stretched it by some percentage, not because each added a fixed number of seconds. The same logic covers file sizes, income and many response-time measurements. Two structural consequences follow: the support is strictly positive, so a log-normal can never produce the negative values a naive normal model would, and the distribution is right-skewed for every positive sigma — skewness is not an accident of the data but a property of the family. ## Symmetry on the log scale Plotted on a log axis, a log-normal is exactly a normal bell curve, symmetric about mu. That gives a useful pair of facts. The **geometric mean** of a log-normal — the exponential of the average log — equals `exp(mu)`, which is also the median. So on the log scale the mean, median and mode all coincide at mu, and it is only the return to the original scale that separates them. It also means multiplicative statements are the natural ones: "sessions vary by a factor of about 3 either side of typical" is a well-formed sentence about a log-normal, while "plus or minus 40 seconds" is not. ## What to do about it when reporting For a strongly right-skewed duration, the arithmetic mean answers a specific question — total time divided by number of sessions — and that is genuinely the right number for capacity and cost, because totals are what infrastructure pays for. It is the wrong number for "what does a user experience", because most users fall below it. The senior habit is to be explicit about which question you are answering: report the median and a couple of upper percentiles for the user-experience story, keep the mean for aggregate volume, and never let a single average stand in for a distribution whose shape you already know is skewed. Confidence statements built on the assumption of a symmetric spread around the mean will also be wrong in a predictable direction: too narrow on the high side and, if taken literally, extending below values the variable can actually take. ## Fast checks A quick way to notice a log-normal in the wild is that the mean sits well above the median, the ratio of high percentiles to the median is large, and the values are strictly positive with no upper wall. If the mean and median are close, you do not have a strongly log-normal quantity — sigma on the log scale must be small.

  • For a log-normal whose log has sigma = 1, how many times the median is the mean?
    About 1.65 times. The ratio is `exp(sigma^2 / 2) = exp(0.5) ≈ 1.65`, independent of mu. It grows quickly with log-scale spread: at `sigma = 2` the ratio is `exp(2) ≈ 7.4`. That single formula is the cleanest way to say how misleading an average will be before you look at any data.
  • How do the mode, median and mean of a log-normal order themselves?
    `mode = exp(mu - sigma^2)` below `median = exp(mu)` below `mean = exp(mu + sigma^2/2)`. That mode-median-mean ordering is the standard signature of right skew, and here it holds exactly for every positive sigma rather than approximately, because all three have closed forms driven by the same two parameters.
  • If a variable is log-normal, what exactly is the distribution of its logarithm?
    Normal with mean mu and standard deviation sigma — that is the definition of the family, not a numerical coincidence. On a log axis the distribution is a symmetric bell curve, so mean, median and mode all coincide at mu there, and the geometric mean of the original variable equals `exp(mu)`, the median.
  • Is the arithmetic mean ever the right summary for log-normal session durations?
    Yes, when the question is about totals. Mean duration times session count gives total time served, which is what capacity and cost depend on, so it is the correct input there. It is the wrong summary for typical user experience, where most sessions sit below it; report the median and upper percentiles for that.

Think of growth compounding rather than accumulating: each factor multiplies the duration by a percentage, so the long side of the distribution stretches out while the short side is squeezed against zero.

saying these in an interview costs you the question

  • Says mu and sigma are the mean and SD of the variable itself
  • Claims the mean of a log-normal equals exp(mu)
  • Treats the average of a skewed duration as typical
  • Assumes any right-skewed positive variable is exactly log-normal
  • Builds symmetric intervals around a log-normal mean

context