Why can a probability density function take values greater than 1 when a probability cannot?
answer
- probability per unit, not probability
- area, not height
- narrow support forces a tall curve
- uniform on [0, 0.5] has height 2
- change the units, the height changes
basics
~20 sA density is probability per unit of x, not a probability. Probability is the area under the curve, so a tall density over a narrow interval still has total area 1. The uniform density on [0, 0.5] equals 2 everywhere.
solid answer
~40 sBecause a density is a rate, not a probability. `f(x)` has units of probability per unit of x, and only `f(x)` multiplied by a width becomes a probability. The two constraints on a density are `f(x) >= 0` everywhere and total area equal to 1 — there is no constraint that `f(x) <= 1`. The cleanest example is a uniform variable on `[0, 0.5]`: one unit of probability spread over a width of 0.5 forces a constant height of `1 / 0.5 = 2`, and the area is still `2 * 0.5 = 1`. Squeeze the support to `[0, 0.01]` and the height becomes 100. A PMF is different: its values *are* probabilities, each in `[0, 1]` and summing to 1, so a PMF value above 1 really would be a bug.
go deeper
Be ready to state that probability is the area under the density, not its height, and to give one example where the height exceeds 1, such as the uniform density on [0, 0.5] having value 2.
An interviewer at this level expects the mechanics: the two real constraints are non-negativity and total area 1, and density carries units of probability per unit of x. Be able to derive the height 2 from width 0.5.
Demonstrate the units argument — rescaling milliseconds to seconds multiplies density values by 1000 while leaving every probability unchanged — and show you can spot the practical failure of reading peak heights off a plotted curve as if they were shares of the data.
Own how distributions get reported to non-specialists. Decide when to present cumulative or interval probabilities rather than density curves, since heights invite misreading, and set the convention so teams compare areas or thresholds instead of peaks across differently scaled metrics.
## A density is a rate For a continuous random variable X, the probability density function f is defined so that for any interval, ``` P(a <= X <= b) = area under f between a and b ``` Notice what the definition does *not* say: it never says `f(x)` is the probability of anything. `f(x)` is probability **per unit of x**. Its units are the reciprocal of x's units — if X is a latency in milliseconds, f is measured in probability per millisecond. A per-unit quantity has no reason to be capped at 1, in the same way that a speed of 30 metres per second is not "more than all the distance there is". Height becomes probability only after you multiply by a width. ## The worked example Let X be uniform on `[0, 0.5]`. All of the probability, which must total 1, is spread evenly across a width of 0.5. For a rectangle, height times width equals area, so ``` height = 1 / 0.5 = 2 ``` The density is `f(x) = 2` on `[0, 0.5]` and 0 elsewhere. Its value is double the largest probability that exists, and the distribution is entirely legitimate: f is non-negative everywhere, and the total area is `2 * 0.5 = 1`. Keep shrinking the support and the height keeps climbing. Uniform on `[0, 0.01]` has `f(x) = 100`. Uniform on `[0, 0.0001]` has `f(x) = 10000`. There is no ceiling at all: a density may even be unbounded near a point and remain valid, provided the area under it stays finite and equal to 1. ## The unit check that settles it Suppose latency X is measured in milliseconds and has some density `f_ms`. Re-express the same latency in seconds, `Y = X / 1000`. The probability of any physical event has not changed — it is the same latency — but the density has: ``` f_s(y) = 1000 * f_ms(1000 * y) ``` The per-second density is a thousand times larger than the per-millisecond one, because the same probability now sits in intervals that are numerically a thousand times narrower. If density values were probabilities, a change of units would have changed the probabilities. They are not, so it did not: the *areas* are identical. This is the argument to reach for when an interviewer pushes. ## The contrast with a PMF A discrete random variable is described by a probability mass function: ``` p(x) = P(X = x), 0 <= p(x) <= 1, sum over all x of p(x) = 1 ``` Here the values genuinely are probabilities. `p(x) > 1` is impossible, and so is a sum above 1. Almost every candidate who insists that `f(x) <= 1` is silently importing the PMF rules into the continuous case. Saying "PMF values are probabilities, PDF values are rates" out loud is usually enough to show you have the distinction. ## What the actual constraints are A function f is a valid probability density exactly when: 1. `f(x) >= 0` for every x — negative probability is meaningless, and a negative region would let some interval get negative probability. 2. The total area under f over the whole line equals 1 — all the probability is accounted for. That is the complete list. No boundedness, no continuity, no requirement to peak below 1. ## Where the confusion causes real damage - **Misreading a plotted density.** A peak at height 3.8 does not mean "38% of the mass is here". Mass is area; to get a probability you must integrate over a range, or read two CDF values and subtract. - **Comparing peak heights across variables with different units.** One curve looks taller purely because its x-axis is finer. Compare probabilities of comparable intervals instead. - **Panicking at a plot.** A density that spikes above 1 on a narrow, concentrated variable is normal, not evidence that something is broken. The general repair in all three cases is the same: stop reading heights, and read areas — or work with the CDF, whose values *are* probabilities and are always in `[0, 1]`. ## How to say it in an interview "A density is probability per unit of x, so it is only constrained to be non-negative and to integrate to 1 — not to stay below 1. Uniform on `[0, 0.5]` has density 2 everywhere and total area `2 * 0.5 = 1`. A PMF is the one whose values are actual probabilities and therefore capped at 1."
- Is there any upper bound at all on a density value?No. The only requirements are `f(x) >= 0` everywhere and total area 1. Concentrating the same unit of probability on a narrower support pushes the height arbitrarily high — uniform on `[0, 0.01]` has density 100 — and a density may even be unbounded near a point while still integrating to 1.
- Does the same argument apply to a PMF?No, and that is the key contrast. A PMF value is a genuine probability, `p(x) = P(X = x)`, so it lies in `[0, 1]` and the values sum to 1 across the support. A PMF value above 1 is a real error. Most people who insist a density cannot exceed 1 are applying the PMF rule in the wrong place.
- If you re-express a latency density from per millisecond to per second, what happens to its values?They scale up by a factor of 1000, since `f_s(y) = 1000 * f_ms(1000 * y)`. The probabilities are unchanged because the intervals shrink by the same factor and the areas are preserved. That units argument is the cleanest proof that density values are not probabilities.
- How do you turn a density value into an approximate probability?Multiply by a small width: `P(x - w/2 < X <= x + w/2)` is approximately `f(x) * w` when w is small enough that f barely changes over it. For anything larger, integrate properly or take a difference of CDF values, `F(b) - F(a)`.
Density is like speed and probability is like distance. Driving at 200 km/h for one minute is fine; the speed number can be large because you only travel far if you also spend time at it.
saying these in an interview costs you the question
- Says f(x) is the probability that X equals x
- Claims a valid density must be at most 1
- Reads a density peak height as a percentage of the mass
- Applies PMF rules to a continuous distribution
- Calls a plot invalid because the curve rises above 1