skip to content

Distribution Families

The named laws that model counts, waiting times and noise, from Bernoulli and Poisson to normal and log-normal, and how to pick one for a described process. Asked in nearly every data screen.

on this pageshow

questions

16

Test scores are normal with mean 100 and SD 15: by the 68-95-99.7 rule, what share exceeds 130?

level: juniorimportance: must knowfreq 82%

answer

  1. convert to standard deviations first
  2. z = (x - mu) / sigma
  3. the rule covers both tails at once
  4. half of the outer five percent

basics

~20 s

About 2.5%. A score of 130 is two standard deviations above the mean, so its z-score is 2; the rule puts roughly 95% of values within two SDs, leaving about 5% split evenly between the two tails.

solid answer

~40 s

First standardise: `z = (130 - 100) / 15 = 2`, so 130 sits exactly two standard deviations above the mean. The 68-95-99.7 rule says a normal distribution holds about 68% of its mass within one SD of the mean, about 95% within two, and about 99.7% within three. If 95% lies between z = -2 and z = +2, about 5% lies outside, and because the normal is symmetric that splits into roughly 2.5% below 70 and 2.5% above 130. The rule is a mental-arithmetic approximation — the exact two-SD coverage is 95.45%, so the true tail above 130 is 2.28%. I would quote "about 2.5%, a bit over 1 in 40" in conversation and reach for the actual normal CDF when the number has to be defensible.

go deeper

for a junior

Be ready to compute a z-score in one line and turn it into a tail share with the rule, including the halving step for a one-sided question.

for a middle

Expect to explain why standardising collapses every normal onto one reference curve, and to quote the exact tail areas alongside the rounded rule.

for a senior

Demonstrate that you check the normality assumption before quoting tail percentages, and that you know 1.96 rather than 2 belongs in anything you publish.

for a principal

Own the risk framing: decisions driven by two-SD thresholds on a metric with heavier tails than assumed will under-price rare events, so argue for tail estimates that do not lean on normality.

## The distribution and its two parameters A normal (Gaussian) distribution is fully specified by two numbers: the mean `mu`, which locates the peak, and the standard deviation `sigma`, which sets the width. Here `mu = 100` and `sigma = 15`. Everything else about the curve — its symmetry, its bell shape, the share of mass in any region — follows from those two parameters alone. Note the notation trap: `N(100, 15)` is written by some authors with 15 as the standard deviation and by others with the second slot holding the variance. Always say which you mean; a candidate who silently treats 15 as a variance gets `sigma = 3.87` and every subsequent number wrong. ## Standardising: the z-score A z-score re-expresses a raw value as a number of standard deviations from the mean: ``` z = (x - mu) / sigma ``` For `x = 130`: `z = (130 - 100) / 15 = 30 / 15 = 2`. Standardising maps any normal distribution onto the standard normal, the one with mean 0 and standard deviation 1. That is why a single table or a single function suffices for all normal distributions: the shape never changes, only the location and the scale, so shifting and rescaling reduces every case to one reference curve. A z-score is also unit-free, which is what makes a score from one test comparable to a score from another with a different mean and spread. ## What the 68-95-99.7 rule states For a normal distribution: ``` P(|z| < 1) ≈ 0.68 -> about 68% within one SD P(|z| < 2) ≈ 0.95 -> about 95% within two SDs P(|z| < 3) ≈ 0.997 -> about 99.7% within three SDs ``` Translated to this score scale, roughly 68% of scores lie between 85 and 115, roughly 95% between 70 and 130, and roughly 99.7% between 55 and 145. ## Getting to the answer The question asks for one tail. About 95% of the mass lies inside two SDs, so about 5% lies outside — but that 5% is the two tails together. Symmetry splits it evenly: ``` P(X > 130) = P(z > 2) ≈ (1 - 0.95) / 2 = 0.025 ``` About 2.5%, or one score in forty. Forgetting the halving step and answering 5% is the single most common error on this question. ## Approximate versus exact The three percentages are deliberately rounded for mental arithmetic. The precise values are: ``` P(|z| < 1) = 0.6827 -> one tail 0.1587 P(|z| < 2) = 0.9545 -> one tail 0.0228 P(|z| < 3) = 0.9973 -> one tail 0.00135 ``` So the exact answer above 130 is 2.28%, not 2.50%. The gap matters in one direction especially: the 95% figure people quote for confidence intervals corresponds to `z = 1.96`, not `z = 2`. Using 2 as shorthand is fine for a whiteboard estimate and wrong for a published number. Being explicit about which one you are using is part of a good answer. ## Reading the tails A few landmarks are worth carrying in your head, because interviewers ask them in either direction: - Above one SD: about 16% (half of the 32% outside one SD). Here that is scores above 115. - Above two SDs: about 2.5%. Here, above 130. - Above three SDs: about 0.15%. Here, above 145. The reverse direction is equally fair game: "what score marks the top 16%?" is answered by `mu + 1*sigma = 115` without any table at all. For a z that is not a whole number — say `z = 1.47` for a score of 122 — the rule gives no exact answer; you interpolate roughly (about 7%) or look up the normal CDF. ## The assumption behind the arithmetic Every number above depends on the variable actually being normal. The rule is not a general property of symmetric or bell-ish distributions: a heavy-tailed symmetric distribution can put far more than 5% beyond two SDs, and a skewed distribution has no symmetry to split a tail with. Applying the rule to a bounded or strongly skewed quantity — a count, a duration, a proportion — produces confident nonsense, sometimes including probability mass on impossible negative values. Stating that caveat unprompted is what separates a mechanical answer from a statistical one.

  • What score marks the top 16% when the mean is 100 and the SD is 15?
    About 115, one standard deviation above the mean. The rule puts roughly 68% within one SD, leaving 32% outside, and symmetry splits that into about 16% in each tail. So `mu + sigma = 100 + 15 = 115` is the cutoff for the top 16%, no table needed. The exact upper-tail value at z = 1 is 15.87%.
  • How would you get the share above 122 using the rule?
    You cannot get it cleanly. `z = (122 - 100)/15 ≈ 1.47`, and the rule only pins z = 1, 2 and 3. You can bracket it — less than the 16% above z = 1, more than the 2.5% above z = 2 — and interpolate to roughly 7%, which is close to the true 7.1%. For anything reportable, use the normal CDF instead of the rule.
  • Why does standardising make two different normal distributions comparable?
    Subtracting the mean and dividing by the standard deviation removes both the location and the scale, mapping any normal onto the standard normal with mean 0 and SD 1. A z-score is therefore a unit-free statement of relative position, so z = 2 means the same rarity whether the raw units are points, kilograms or seconds.
  • When does the 68-95-99.7 rule stop being trustworthy?
    As soon as the variable is not close to normal. Heavy tails put far more than 5% beyond two SDs, and skew destroys the symmetry that lets you halve a two-tail figure. On a bounded quantity the rule can even assign mass to impossible values, such as negative durations. Treat it as a normal-only shortcut, not a universal law.

saying these in an interview costs you the question

  • Answers 5%, forgetting the rule spans both tails
  • Treats 15 as the variance rather than the standard deviation
  • Claims 68-95-99.7 are exact rather than rounded
  • Uses z = 2 where a reported 95% needs 1.96
  • Applies the rule to a clearly skewed variable

context

open as a page

What are the mean and variance of a binomial random variable with n trials and success probability p?

level: juniorimportance: must knowfreq 78%

basics

~10 s

A binomial count over n independent trials with success probability p has mean np and variance np(1-p). It is the sum of n Bernoulli(p) indicators, each contributing mean p and variance p(1-p).

open as a page

For a Poisson count averaging 3 support tickets per hour, what is the probability of zero tickets in an hour?

level: juniorimportance: must knowfreq 64%

basics

~20 s

The Poisson mass function is P(X = k) = e^(-λ) λ^k / k!. With λ = 3 tickets per hour, P(X = 0) = e^(-3), about 0.0498, so roughly a 5 percent chance of a completely quiet hour.

open as a page

Which distributions model 'outages per week' and 'hours between outages' for the same incident stream?

level: middleimportance: must knowfreq 70%

basics

~10 s

Counts in a fixed window, outages per week, are Poisson. The gaps between consecutive outages, hours between them, are exponential. Both describe one event stream: one view counts events, the other measures waiting times.

open as a page

If bus waiting time is exponential with mean 10 minutes, why does waiting 10 minutes not shorten the expected remaining wait?

level: middleimportance: must knowfreq 64%

basics

~20 s

The exponential distribution is memoryless: P(T > s + t | T > s) = P(T > t). Time already spent waiting carries no information about what remains, so the expected remaining wait is still the full 10 minutes.

open as a page

Which distribution models the number of retries before a flaky request finally succeeds?

level: juniorimportance: should knowfreq 48%

basics

~20 s

The geometric distribution. It counts repeated independent attempts, each succeeding with the same probability p, up to the first success. It is not Poisson, because Poisson counts events inside a fixed window rather than trials.

open as a page

A passenger arrives uniformly at random within a 60-minute window: what is the probability the arrival falls in the final 10 minutes?

level: juniorimportance: should knowfreq 50%

basics

~10 s

One sixth, about 16.7%. A continuous uniform distribution has a flat density, so probability is proportional to interval length: the last 10 minutes out of 60 give 10/60 = 1/6.

open as a page

In the geometric distribution, why is the mean number of trials until the first success 1/p?

level: middleimportance: should knowfreq 44%

basics

~20 s

If independent trials each succeed with probability p, the number of trials up to and including the first success is geometric with mean 1/p. With a 2 percent click rate you expect 50 impressions per first click.

open as a page

When can a binomial with n = 10,000 and p = 0.0003 be approximated by a Poisson distribution?

level: middleimportance: should knowfreq 46%

basics

~20 s

When n is large and p small with λ = np moderate, a binomial is closely approximated by Poisson(np). Here λ = 10,000 × 0.0003 = 3, and the two distributions agree to three decimals.

open as a page

Why does the Poisson distribution have variance equal to its mean?

level: middleimportance: should knowfreq 52%

basics

~20 s

A Poisson count has one parameter λ that fixes both its mean and its variance, so its standard deviation is sqrt(λ). Spread is not free to differ from level, and a variance-to-mean ratio far from 1 signals a wrong model.

open as a page

For disk time-to-failure, when is an exponential model wrong and a Weibull right?

level: seniorimportance: should knowfreq 34%

basics

~20 s

An exponential lifetime assumes the failure rate never changes with age, so an old disk is as likely to fail next month as a new one. When wear-out makes failures rise with age, use a Weibull.

open as a page

Insurance claim sizes have a long right tail; how do you choose between log-normal and Pareto?

level: seniorimportance: should knowfreq 44%

basics

~10 s

Take logs of the claim sizes: a symmetric bell after logging points to log-normal. A tail that traces a straight line on a log-log survival plot points to Pareto. A normal fits neither.

open as a page

Session durations follow a log-normal distribution: why does the mean sit above the median?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A log-normal variable is the exponential of a normal one, so its right tail is long. The median is exp(mu) while the mean is exp(mu + sigma squared over 2), which is always larger: rare long sessions lift the average.

open as a page

Which family models a per-user conversion propensity that must lie between 0 and 1?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

The beta distribution. Its support is exactly the interval from 0 to 1, and two shape parameters let it be flat, bell-shaped or piled at both ends. A normal would put probability on impossible values.

open as a page

What does a gamma distribution with shape k and rate lambda represent?

level: middleimportance: nice to knowfreq 27%

basics

~20 s

A gamma with shape k and rate lambda is the waiting time until the k-th event when events happen independently at rate lambda. Its mean is k/lambda, and shape k = 1 gives the exponential.

open as a page

In a QA lot of 50 units with 5 defectives, why is the count of defectives in 10 draws without replacement not binomial?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Drawing without replacement makes the trials dependent: each draw changes what is left in the lot, so the defect probability is not constant. The exact model is hypergeometric, which shares the binomial's mean but has a smaller variance.

open as a page