skip to content

Uniform, Normal and Exponential

The continuous workhorses: the uniform, the normal with its 68-95-99.7 rule, the memoryless exponential, and the log-normal for skewed positive data. Interviewers ask which one a quantity follows.

on this pageshow

questions

5

Test scores are normal with mean 100 and SD 15: by the 68-95-99.7 rule, what share exceeds 130?

level: juniorimportance: must knowfreq 82%

answer

  1. convert to standard deviations first
  2. z = (x - mu) / sigma
  3. the rule covers both tails at once
  4. half of the outer five percent

basics

~20 s

About 2.5%. A score of 130 is two standard deviations above the mean, so its z-score is 2; the rule puts roughly 95% of values within two SDs, leaving about 5% split evenly between the two tails.

solid answer

~40 s

First standardise: `z = (130 - 100) / 15 = 2`, so 130 sits exactly two standard deviations above the mean. The 68-95-99.7 rule says a normal distribution holds about 68% of its mass within one SD of the mean, about 95% within two, and about 99.7% within three. If 95% lies between z = -2 and z = +2, about 5% lies outside, and because the normal is symmetric that splits into roughly 2.5% below 70 and 2.5% above 130. The rule is a mental-arithmetic approximation — the exact two-SD coverage is 95.45%, so the true tail above 130 is 2.28%. I would quote "about 2.5%, a bit over 1 in 40" in conversation and reach for the actual normal CDF when the number has to be defensible.

go deeper

for a junior

Be ready to compute a z-score in one line and turn it into a tail share with the rule, including the halving step for a one-sided question.

for a middle

Expect to explain why standardising collapses every normal onto one reference curve, and to quote the exact tail areas alongside the rounded rule.

for a senior

Demonstrate that you check the normality assumption before quoting tail percentages, and that you know 1.96 rather than 2 belongs in anything you publish.

for a principal

Own the risk framing: decisions driven by two-SD thresholds on a metric with heavier tails than assumed will under-price rare events, so argue for tail estimates that do not lean on normality.

## The distribution and its two parameters A normal (Gaussian) distribution is fully specified by two numbers: the mean `mu`, which locates the peak, and the standard deviation `sigma`, which sets the width. Here `mu = 100` and `sigma = 15`. Everything else about the curve — its symmetry, its bell shape, the share of mass in any region — follows from those two parameters alone. Note the notation trap: `N(100, 15)` is written by some authors with 15 as the standard deviation and by others with the second slot holding the variance. Always say which you mean; a candidate who silently treats 15 as a variance gets `sigma = 3.87` and every subsequent number wrong. ## Standardising: the z-score A z-score re-expresses a raw value as a number of standard deviations from the mean: ``` z = (x - mu) / sigma ``` For `x = 130`: `z = (130 - 100) / 15 = 30 / 15 = 2`. Standardising maps any normal distribution onto the standard normal, the one with mean 0 and standard deviation 1. That is why a single table or a single function suffices for all normal distributions: the shape never changes, only the location and the scale, so shifting and rescaling reduces every case to one reference curve. A z-score is also unit-free, which is what makes a score from one test comparable to a score from another with a different mean and spread. ## What the 68-95-99.7 rule states For a normal distribution: ``` P(|z| < 1) ≈ 0.68 -> about 68% within one SD P(|z| < 2) ≈ 0.95 -> about 95% within two SDs P(|z| < 3) ≈ 0.997 -> about 99.7% within three SDs ``` Translated to this score scale, roughly 68% of scores lie between 85 and 115, roughly 95% between 70 and 130, and roughly 99.7% between 55 and 145. ## Getting to the answer The question asks for one tail. About 95% of the mass lies inside two SDs, so about 5% lies outside — but that 5% is the two tails together. Symmetry splits it evenly: ``` P(X > 130) = P(z > 2) ≈ (1 - 0.95) / 2 = 0.025 ``` About 2.5%, or one score in forty. Forgetting the halving step and answering 5% is the single most common error on this question. ## Approximate versus exact The three percentages are deliberately rounded for mental arithmetic. The precise values are: ``` P(|z| < 1) = 0.6827 -> one tail 0.1587 P(|z| < 2) = 0.9545 -> one tail 0.0228 P(|z| < 3) = 0.9973 -> one tail 0.00135 ``` So the exact answer above 130 is 2.28%, not 2.50%. The gap matters in one direction especially: the 95% figure people quote for confidence intervals corresponds to `z = 1.96`, not `z = 2`. Using 2 as shorthand is fine for a whiteboard estimate and wrong for a published number. Being explicit about which one you are using is part of a good answer. ## Reading the tails A few landmarks are worth carrying in your head, because interviewers ask them in either direction: - Above one SD: about 16% (half of the 32% outside one SD). Here that is scores above 115. - Above two SDs: about 2.5%. Here, above 130. - Above three SDs: about 0.15%. Here, above 145. The reverse direction is equally fair game: "what score marks the top 16%?" is answered by `mu + 1*sigma = 115` without any table at all. For a z that is not a whole number — say `z = 1.47` for a score of 122 — the rule gives no exact answer; you interpolate roughly (about 7%) or look up the normal CDF. ## The assumption behind the arithmetic Every number above depends on the variable actually being normal. The rule is not a general property of symmetric or bell-ish distributions: a heavy-tailed symmetric distribution can put far more than 5% beyond two SDs, and a skewed distribution has no symmetry to split a tail with. Applying the rule to a bounded or strongly skewed quantity — a count, a duration, a proportion — produces confident nonsense, sometimes including probability mass on impossible negative values. Stating that caveat unprompted is what separates a mechanical answer from a statistical one.

  • What score marks the top 16% when the mean is 100 and the SD is 15?
    About 115, one standard deviation above the mean. The rule puts roughly 68% within one SD, leaving 32% outside, and symmetry splits that into about 16% in each tail. So `mu + sigma = 100 + 15 = 115` is the cutoff for the top 16%, no table needed. The exact upper-tail value at z = 1 is 15.87%.
  • How would you get the share above 122 using the rule?
    You cannot get it cleanly. `z = (122 - 100)/15 ≈ 1.47`, and the rule only pins z = 1, 2 and 3. You can bracket it — less than the 16% above z = 1, more than the 2.5% above z = 2 — and interpolate to roughly 7%, which is close to the true 7.1%. For anything reportable, use the normal CDF instead of the rule.
  • Why does standardising make two different normal distributions comparable?
    Subtracting the mean and dividing by the standard deviation removes both the location and the scale, mapping any normal onto the standard normal with mean 0 and SD 1. A z-score is therefore a unit-free statement of relative position, so z = 2 means the same rarity whether the raw units are points, kilograms or seconds.
  • When does the 68-95-99.7 rule stop being trustworthy?
    As soon as the variable is not close to normal. Heavy tails put far more than 5% beyond two SDs, and skew destroys the symmetry that lets you halve a two-tail figure. On a bounded quantity the rule can even assign mass to impossible values, such as negative durations. Treat it as a normal-only shortcut, not a universal law.

saying these in an interview costs you the question

  • Answers 5%, forgetting the rule spans both tails
  • Treats 15 as the variance rather than the standard deviation
  • Claims 68-95-99.7 are exact rather than rounded
  • Uses z = 2 where a reported 95% needs 1.96
  • Applies the rule to a clearly skewed variable

context

open as a page

If bus waiting time is exponential with mean 10 minutes, why does waiting 10 minutes not shorten the expected remaining wait?

level: middleimportance: must knowfreq 64%

basics

~20 s

The exponential distribution is memoryless: P(T > s + t | T > s) = P(T > t). Time already spent waiting carries no information about what remains, so the expected remaining wait is still the full 10 minutes.

open as a page

A passenger arrives uniformly at random within a 60-minute window: what is the probability the arrival falls in the final 10 minutes?

level: juniorimportance: should knowfreq 50%

basics

~10 s

One sixth, about 16.7%. A continuous uniform distribution has a flat density, so probability is proportional to interval length: the last 10 minutes out of 60 give 10/60 = 1/6.

open as a page

Session durations follow a log-normal distribution: why does the mean sit above the median?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A log-normal variable is the exponential of a normal one, so its right tail is long. The median is exp(mu) while the mean is exp(mu + sigma squared over 2), which is always larger: rare long sessions lift the average.

open as a page

What does a gamma distribution with shape k and rate lambda represent?

level: middleimportance: nice to knowfreq 27%

basics

~20 s

A gamma with shape k and rate lambda is the waiting time until the k-th event when events happen independently at rate lambda. Its mean is k/lambda, and shape k = 1 gives the exponential.

open as a page