Why does the mean of 1/X exceed 1 divided by the mean of X for a random latency X?
answer
- expectation does not pass through curves
- which way the function bends
- small values dominate the reciprocal
- equality only for a constant
basics
~10 sBecause the reciprocal function is strictly convex on positive values, Jensen's inequality gives E[1/X] >= 1/E[X], strict unless X is constant. Averaging per-item rates therefore overstates the rate implied by the average latency.
solid answer
~50 sJensen's inequality says that for a convex function g, `E[g(X)] >= g(E[X])`, and for a concave g the inequality reverses. The reciprocal `g(x) = 1/x` is strictly convex for positive x, so `E[1/X] >= 1/E[X]` with strict inequality whenever X actually varies. Concretely, if latency is 1 second half the time and 3 seconds half the time, the mean latency is 2 seconds and its reciprocal is 0.5 per second, but the mean of the per-request rates is `(1 + 1/3)/2 = 0.667` per second. The gap grows with the spread of X and is driven by the small values of X, whose reciprocals are enormous. The practical rule: an average of ratios is not the ratio of the averages. If you want throughput over a window, divide total work by total time rather than averaging per-request rates, which weights every request equally regardless of how long it took.
go deeper
Recall that an expectation can be moved inside only for affine functions, and that the mean of reciprocals is not the reciprocal of the mean.
State Jensen's inequality with the correct direction for convex and for concave functions, and produce a two-point example that exhibits the gap numerically.
Spot the error live in a dashboard or report: an average of per-unit ratios presented as an aggregate rate, and know the correct aggregation and how large the discrepancy is likely to be.
Own the standard for how ratio metrics are defined and aggregated across the organisation, so that team-level and company-level numbers reconcile instead of quietly disagreeing.
## The inequality A function g is convex when its second derivative is non-negative wherever it is defined, meaning the curve bends upward. Jensen's inequality states that for such a g and any random variable X with finite mean, `E[g(X)] >= g(E[X])` If g is concave, the curve bends downward and the inequality flips to `E[g(X)] <= g(E[X])`. If g is affine, that is `g(x) = ax + b`, both hold and the relation is an equality: this is exactly the linearity case, and it is the only case in which you may freely move an expectation inside a function. For `g(x) = 1/x` on positive x, the second derivative is `2/x^3`, which is positive, so the reciprocal is strictly convex and `E[1/X] >= 1/E[X]` with equality only if X takes a single value with probability 1. ## A concrete gap Suppose a request takes 1 second with probability 0.5 and 3 seconds with probability 0.5. - Mean latency: `E[X] = 2` seconds. Reciprocal of the mean: `1/E[X] = 0.5` requests per second. - Mean of the rates: `E[1/X] = 0.5*(1/1) + 0.5*(1/3) = 0.6667` requests per second. The two differ by a third. Widen the spread and the gap explodes: with latencies of 0.1 and 3.9 seconds the mean is still 2 seconds, but `E[1/X] = 0.5*10 + 0.5*0.2564 = 5.13` per second, ten times the reciprocal of the mean. The asymmetry comes from the shape of `1/x`: near zero the reciprocal blows up, while for large x it flattens toward zero, so a few very fast observations dominate the average of rates while contributing almost nothing to the average of durations. ## Why this matters when reporting The two quantities answer different questions, and quoting one when the audience wants the other is a real and common reporting error. - `1/E[X]` is the throughput of a single sequential worker: how many items you get through per unit time, since total time is what limits you. - `E[1/X]` is the unweighted average of per-item rates, which gives a fast item the same weight as a slow one even though the fast one occupied a sliver of the window. The same trap appears whenever a ratio is averaged across units: per-user conversion rates averaged across users, per-region error rates averaged across regions, per-query cost averaged across queries. In each case the average of ratios is not the ratio of the totals, and the difference is not a rounding artefact but a systematic bias whose direction is fixed by the curvature of the map. The fix is to decide which question you are answering. If you want the aggregate rate over a window, aggregate the numerator and the denominator separately and divide once: total items divided by total time. If you genuinely want each unit weighted equally, an average of ratios is the right object, but then say so, because the number will not reconcile with the aggregate and someone will eventually ask why. ## The general pattern Once you see the shape of the argument you can predict the direction of the bias for any non-affine transformation without doing algebra: - Squaring is convex, so `E[X^2] >= (E[X])^2`. That is the same fact as variance being non-negative. - The logarithm is concave, so `E[log X] <= log E[X]`. Averaging on a log scale and then exponentiating gives something no larger than the plain average. - The exponential is convex, so `E[e^X] >= e^(E[X])`. - The maximum of a variable and a constant is convex, so `E[max(X, c)] >= max(E[X], c)`. This is why plugging a mean into a payoff with a floor understates the payoff. The last one is the general lesson for planning: substituting an average input into a non-linear model does not give you the average output. Feeding mean traffic into a queueing formula understates mean delay, because delay curves upward in load. ## How big is the gap? The size of the Jensen gap grows with both the curvature of g and the spread of X. A second-order approximation around the mean, when X is concentrated enough for it to be meaningful, gives `E[g(X)] approximately g(E[X]) + 0.5 * g''(E[X]) * Var(X)`. For the reciprocal this reads `E[1/X] approximately 1/mu + Var(X)/mu^3`, so the excess scales with the variance and blows up as the mean approaches zero. Two consequences: for a tight, well-behaved distribution the two numbers may agree to within noise, and for a heavy-tailed or near-zero-mass distribution they can differ by orders of magnitude, or `E[1/X]` may not even be finite. Knowing when the gap is negligible is as valuable in an interview as knowing that it exists.
- Which way does the inequality go for a concave function such as the logarithm?It reverses: `E[log X] <= log E[X]`. Concave functions bend downward, so averaging the inputs and then transforming overstates the average of the transformed values. Only affine functions give equality, which is precisely the linearity-of-expectation case.
- When is the gap between E[1/X] and 1/E[X] zero?Only when X is constant with probability 1. Any genuine variability makes the inequality strict for a strictly convex function. A second-order approximation gives the excess as roughly `Var(X)/mu^3`, so the gap shrinks toward zero as the spread shrinks and grows without bound as the mean approaches zero.
- How should you aggregate per-request throughput across a window?Sum the numerators and the denominators separately and divide once: total requests divided by total elapsed time. That weights each request by the time it actually consumed. Averaging per-request rates weights a one-millisecond request the same as a ten-second one, which is a different question and will not reconcile with the aggregate.
- Does the same reasoning explain why E[X^2] is at least (E[X])^2?Yes. Squaring is convex, so Jensen gives `E[X^2] >= (E[X])^2` directly, and the difference is exactly the variance. The two facts are the same statement seen from different sides, which is a useful sanity check that you have the direction of the inequality right.
saying these in an interview costs you the question
- Assumes E[g(X)] equals g(E[X]) for any function g
- Gets the direction of the inequality backwards for a convex function
- Calls the gap a rounding or sampling artefact
- Averages per-unit ratios and calls it the aggregate rate
- Plugs a mean input into a non-linear model to get a mean output