What are the mean and variance of a binomial random variable with n trials and success probability p?
answer
- a sum of independent indicators
- linearity gives the mean
- variances add under independence
- each indicator contributes p(1-p)
basics
~10 sA binomial count over n independent trials with success probability p has mean np and variance np(1-p). It is the sum of n Bernoulli(p) indicators, each contributing mean p and variance p(1-p).
solid answer
~40 sA binomial variable counts successes in n independent trials that each succeed with the same probability p, so it is a sum of n Bernoulli(p) indicators. Each indicator has mean p and variance p(1-p). Expectation is linear, so the mean is np; variances add for independent terms, so the variance is np(1-p) and the standard deviation is sqrt(np(1-p)). For 10 flips of a coin with P(heads) = 0.3, the mean is 3, the variance is 2.1 and the standard deviation is about 1.45. Note that the (1-p) factor is what makes a binomial less variable than a count with the same mean but unbounded support, and that variance is maximised at p = 0.5, where p(1-p) = 0.25.
go deeper
Be ready to state np and np(1-p) instantly and plug in numbers, including the standard deviation as the square root. Know that a single trial is Bernoulli with mean p and variance p(1-p).
Explain where the formulas come from: the count is a sum of independent indicators, linearity gives the mean, and variances add only because the trials are independent. Be able to say where p(1-p) peaks and why.
Show you check the assumptions before the formulas. Point out that fixed n, independence and constant p rarely all hold in real data, and describe how correlated trials leave the mean intact while inflating the true variance.
Own the modelling call: decide when a binomial framing is worth its simplicity versus when dependence or a drifting success rate makes it misleading, and be clear about which downstream decisions the understated variance would corrupt.
## The two distributions involved A **Bernoulli** trial is a single experiment with two outcomes, coded 1 for success and 0 for failure, where P(success) = p. Its indicator variable X has E[X] = 1·p + 0·(1-p) = p. For the variance, note X² = X because 0² = 0 and 1² = 1, so E[X²] = p and Var(X) = E[X²] - (E[X])² = p - p² = p(1-p). A **binomial** variable is the count of successes across n such trials, written X ~ Binomial(n, p). Its probability mass function is `P(X = k) = C(n, k) · p^k · (1-p)^(n-k)` for k = 0, 1, ..., n, where `C(n, k) = n! / (k!(n-k)!)` counts the orderings that give exactly k successes. ## The assumptions that make it binomial Three conditions must hold: the number of trials n is **fixed in advance**, the trials are **independent**, and p is **constant** across trials. Break any one and the formulas below stop applying. If n itself is random, or trials influence each other, or p drifts, you no longer have a binomial count — a point interviewers probe more often than the formulas themselves. ## Deriving the mean Write X = X₁ + X₂ + ... + Xₙ where each Xᵢ is the Bernoulli indicator for trial i. Expectation is linear, and linearity does **not** require independence: `E[X] = E[X₁] + ... + E[Xₙ] = p + p + ... + p = np`. This is why the mean is easy: it survives even correlated trials. ## Deriving the variance Variance adds across a sum only when the terms are uncorrelated. Under the independence assumption the covariance terms vanish, so `Var(X) = Var(X₁) + ... + Var(Xₙ) = n · p(1-p) = np(1-p)`, and `SD(X) = sqrt(np(1-p))`. If the trials were **positively** correlated, the mean would still be np but the covariance terms would be positive and the true variance would exceed np(1-p) — the count would be more dispersed than binomial theory predicts. That is the usual reason a real count looks "too variable" for its binomial model. ## Where the variance peaks The factor p(1-p) is a downward parabola in p, maximised at p = 0.5 with value 0.25. So for fixed n the variance is largest at p = 0.5, giving Var = n/4, and it shrinks toward 0 as p approaches 0 or 1. Intuitively, when p is 0.01 almost every trial fails, so the outcome is nearly deterministic and there is little to vary; when p = 0.5 each trial is maximally uncertain. ## Worked example Take n = 10 flips of a biased coin with p = 0.3. - Mean: np = 10 × 0.3 = **3** heads. - Variance: np(1-p) = 10 × 0.3 × 0.7 = **2.1**. - Standard deviation: sqrt(2.1) ≈ **1.45**. - A specific probability: P(X = 3) = C(10, 3) · 0.3³ · 0.7⁷ = 120 × 0.027 × 0.0823543 ≈ 0.267. So you expect about 3 heads, and outcomes between roughly 2 and 4 are unremarkable while 8 heads would be several standard deviations out. ## Related quantities worth having ready - The **sample proportion** p̂ = X/n has mean p and variance p(1-p)/n, because scaling a variable by 1/n divides its variance by n². Its spread shrinks like 1/sqrt(n). - A **fair coin flipped 100 times** gives mean 50, variance 25, standard deviation 5 — a useful mental anchor, since a result outside roughly 40 to 60 is already far out. - **Bernoulli is the n = 1 case**: mean p, variance p(1-p). Candidates who see the binomial as a sum of indicators rarely misremember either formula. ## Common errors The frequent slip is quoting the variance as np, which is the mean, or as np², which is dimensionally wrong. Another is applying the formulas to sampling without replacement from a small population, where p changes between draws and the trials are not independent. And remember that np is an expected count, not necessarily an achievable one: with n = 3, p = 0.5 the mean is 1.5 even though no single outcome equals 1.5.
- Why is a binomial count's variance largest at p = 0.5?The variance is np(1-p), and p(1-p) is a downward parabola peaking at p = 0.5 with value 0.25, giving Var = n/4. At extreme p almost every trial goes the same way, so the count is nearly deterministic and the variance collapses toward zero.
- If the trials were positively correlated rather than independent, what changes?The mean stays np, because linearity of expectation does not need independence. The variance grows beyond np(1-p), since the positive covariance terms between indicators no longer cancel. The count is then overdispersed relative to the binomial, and binomial-based tail probabilities are too small.
- For 100 flips of a fair coin, what are the mean and standard deviation of the number of heads?Mean is np = 100 × 0.5 = 50. Variance is np(1-p) = 100 × 0.5 × 0.5 = 25, so the standard deviation is 5. That makes results between about 40 and 60 unremarkable, while 70 heads sits four standard deviations out.
Think of n identical loaded dice, each scoring 1 or 0. The total score's average is just n times one die's average, and because they are rolled independently their uncertainties pile up additively rather than cancelling.
saying these in an interview costs you the question
- Quotes the variance as np, which is the mean
- Cannot connect the binomial to a sum of Bernoulli indicators
- Claims variance is largest when p is near 0 or 1
- Applies binomial formulas to draws without replacement
- Says the mean must be a whole number of successes