skip to content

What bound does Chebyshev's inequality place on how often a value falls far from the mean?

level: middleimportance: should knowfreq 35%

answer

  1. no distributional assumption required
  2. the bound falls as k rises
  3. one over k squared
  4. three standard deviations, at most one ninth
  5. Markov applied to squared deviation

basics

~20 s

Chebyshev's inequality says for any distribution with a finite mean and variance, the probability of landing more than k standard deviations from the mean is at most 1 over k squared — at three standard deviations, at most one ninth.

solid answer

~40 s

For any random variable `X` with finite mean `mu` and finite standard deviation `sigma`, and any `k > 0`: `P(|X - mu| >= k * sigma) <= 1 / k^2` The striking part is what it does *not* assume: no normality, no symmetry, no shape at all — just that the mean and variance exist. So if session lengths have mean `mu` and standard deviation `sigma`, at most 1/9 of sessions can be more than 3 standard deviations from `mu`, whatever the distribution looks like. The price of that generality is looseness: a normal distribution puts about 0.27% beyond 3 standard deviations, roughly forty times under the 11.1% Chebyshev allows. Treat it as a worst-case guarantee, not an estimate. It is also the standard route to the weak law of large numbers.

go deeper

for a junior

Memorise the form and one worked value: at most one over k squared beyond k standard deviations, so at most one quarter beyond two. Say "at most", never "about".

for a middle

Be able to explain why no distributional shape is assumed and why that makes the bound loose. Knowing that Chebyshev comes from Markov applied to the squared deviation is the usual follow-up.

for a senior

Show judgment about when a distribution-free guarantee is the right instrument — service-level promises on data whose shape you refuse to assume — versus when a modelled tail estimate is more useful and defensible.

for a principal

Own the tradeoff between assumption-free guarantees and calibrated estimates in how your organisation sets thresholds and SLOs. Be ready to argue what you are willing to assume about a distribution and what that assumption costs if it is wrong.

## The statement Let `X` be a random variable with finite mean `mu = E[X]` and finite variance `sigma^2 = Var(X)`. For any `k > 0`: ``` P(|X - mu| >= k * sigma) <= 1 / k^2 ``` In words: the probability of being at least `k` standard deviations away from the mean, in either direction, is at most `1/k^2`. A few values are worth memorising: | k | Chebyshev bound | Normal distribution actual | |---|---|---| | 2 | at most 1/4 = 25% | about 4.6% | | 3 | at most 1/9 ≈ 11.1% | about 0.27% | | 4 | at most 1/16 = 6.25% | about 0.006% | Note that the bound is useless for `k <= 1`: at `k = 1` it says the probability is at most 1, which is true of everything. ## Why it is remarkable The inequality makes **no assumption about the shape of the distribution**. Skewed, bimodal, discrete, bounded, long-tailed — it does not matter. If the mean and variance exist, the bound holds. That is why it is the right tool when someone hands you a summary ("mean 400 ms, standard deviation 120 ms") and nothing else, and asks what fraction of the mass can possibly sit beyond some threshold. The honest answer uses Chebyshev, and it is a bound, not a prediction. ## Why it is loose Generality costs precision. The bound must hold for the worst distribution consistent with a given mean and variance, so it is quoted at that worst case. In fact the bound is tight — there are distributions achieving it exactly, built by placing mass at just three points: a large lump at the mean and two small lumps at `mu ± k*sigma`. Real data almost never looks like that, which is why real tail probabilities are far below the bound. The practical rule: use Chebyshev when you refuse to assume a shape and need a guarantee that cannot be wrong. Use a distributional model when you can defend the assumption and need an estimate that is close. ## Markov's inequality, the parent result Chebyshev is a corollary of a simpler statement. **Markov's inequality**: for a random variable `X` that is non-negative, and any `a > 0`, ``` P(X >= a) <= E[X] / a ``` This needs only non-negativity and a finite mean — no variance at all. If average session length is 4 minutes, at most 20% of sessions can last 20 minutes or longer, because 4/20 = 0.2. Chebyshev follows by applying Markov to the non-negative variable `(X - mu)^2` with threshold `a = k^2 * sigma^2`: ``` P(|X - mu| >= k*sigma) = P((X - mu)^2 >= k^2 * sigma^2) <= E[(X - mu)^2] / (k^2 * sigma^2) = sigma^2 / (k^2 * sigma^2) = 1/k^2 ``` So Markov is the weaker, more general tool (mean only, non-negative variable) and Chebyshev is what you get when you also know the variance and centre the variable. ## The route to the weak law of large numbers Chebyshev's real payoff is that it turns "averaging reduces variability" into a proof. Take `n` independent draws from a distribution with mean `mu` and variance `sigma^2`. The sample mean `Xbar_n` has expected value `mu` and variance `sigma^2 / n`, so its standard deviation is `sigma / sqrt(n)`. Apply Chebyshev to `Xbar_n` with any fixed tolerance `eps > 0`: ``` P(|Xbar_n - mu| >= eps) <= sigma^2 / (n * eps^2) ``` For fixed `eps` and `sigma`, the right-hand side goes to zero as `n` grows. That is exactly the weak law of large numbers: for any tolerance you name, the probability that the sample mean misses the true mean by more than that tolerance vanishes as the sample grows. Three lines, no distributional assumption beyond a finite variance. ## Common mistakes - **Quoting it as an equality or an estimate.** It says "at most". Saying "11% of sessions are beyond three standard deviations" is wrong; "no more than 11%" is right. - **Getting `k` and the bound backwards.** The bound is `1/k^2`, so it shrinks as `k` grows. Larger deviations are rarer, and the guarantee gets stronger. - **Applying Markov to a variable that can be negative.** Non-negativity is essential; profit-and-loss figures do not qualify without transformation. - **Assuming it needs a bell shape.** It needs nothing of the sort, and that is the entire point.

  • How does Chebyshev's inequality give a proof of the weak law of large numbers?
    Apply it to the sample mean. With n independent draws the sample mean has expected value mu and variance sigma squared over n, so for any fixed tolerance eps the probability of missing mu by more than eps is at most sigma squared divided by n times eps squared. That goes to zero as n grows, which is exactly the weak law.
  • What does Markov's inequality require that Chebyshev's does not?
    Markov's inequality requires the variable to be non-negative, and in exchange needs only its mean: P(X >= a) <= E[X]/a. Chebyshev drops the non-negativity requirement — it applies to the squared deviation, which is automatically non-negative — but pays for that by needing a finite variance as well.
  • When is Chebyshev's bound too loose to be worth using?
    Whenever you can defend a distributional assumption. For a normal distribution the true mass beyond three standard deviations is about 0.27% against Chebyshev's 11.1% ceiling, so planning capacity from the bound would over-provision enormously. Use it as a guarantee you cannot be wrong about, not as an estimate of what will happen.

Chebyshev is a building code written for every possible design, not for yours. It guarantees a margin so generous that any well-behaved distribution beats it comfortably — which is exactly why it is safe to rely on when you know nothing about the design.

saying these in an interview costs you the question

  • States the bound as an estimate rather than a ceiling
  • Thinks the inequality assumes a bell-shaped distribution
  • Inverts the bound so larger k allows more mass
  • Applies Markov's inequality to a variable that can go negative
  • Cannot say what the bound gives at two or three standard deviations

context