skip to content

Confidence Intervals

Turning a point estimate into a range built from a critical value times the standard error, and what 95% coverage actually promises. Interviewers probe the interpretation far more than the formula.

on this pageshow

explore

questions

16

What does the 95% in a 95% confidence interval actually refer to?

level: juniorimportance: must knowfreq 88%

answer

  1. the procedure, not this one interval
  2. randomness lives in the endpoints
  3. the parameter is fixed, not moving
  4. picture 100 intervals, about 5 miss

basics

~20 s

The 95% describes the procedure, not one interval. If you repeated the sampling and rebuilt the interval many times, about 95% of those intervals would contain the fixed true parameter. Any single interval either covers it or does not.

solid answer

~40 s

The 95% is a property of the method across repeated sampling, not a probability attached to the numbers you happened to compute. Imagine drawing 100 fresh samples from the same population and building an interval from each one: the true parameter stays put while the intervals jump around it, and about 95 of them cover it while about 5 miss. Once you have your one interval, say `[4.1, 5.3]`, the parameter is a fixed unknown number that is either inside or outside, so "there is a 95% probability the true mean is between 4.1 and 5.3" is the wrong sentence. The defensible wording is "this interval came from a procedure that captures the true mean 95% of the time," and that 95% only holds while the sampling and modelling assumptions behind it hold.

code

python · 14 lines
python
import random, statistics, math

TRUE_MEAN, SIGMA, N, TRIALS = 5.0, 2.0, 30, 20000
random.seed(0)

covered = 0
for _ in range(TRIALS):
    sample = [random.gauss(TRUE_MEAN, SIGMA) for _ in range(N)]
    center = statistics.fmean(sample)
    half = 1.96 * SIGMA / math.sqrt(N)   # sigma treated as known here
    if center - half <= TRUE_MEAN <= center + half:
        covered += 1

print(covered / TRIALS)   # ~0.95: the share of INTERVALS that caught the fixed mean

go deeper

for a junior

Be ready to state the repeated-sampling definition in one clean sentence and to say out loud which quantity is fixed and which one moves. Knowing the wrong phrasing well enough to correct it is half the answer.

for a middle

Explain the mechanics: the endpoints are functions of the sample, so they are random variables, and coverage is the long-run frequency with which that random range traps a constant. Be able to describe the 100-intervals picture without hand-waving.

for a senior

Show that you check whether the nominal level is earned. Interviewers expect you to name the assumptions coverage rests on and to say how you would test them, for instance by simulating the whole pipeline and counting how often the interval covers a known value.

for a principal

Own the framing question: what guarantee does the organisation actually want from a reported range, and is repeated-sampling coverage the right one for that decision? Be prepared to defend the confidence level as a deliberate choice rather than a default.

## The claim being made A confidence interval is a random *range* computed from a sample, offered as a statement about an unknown population quantity: a mean, a proportion, a difference between two groups. The confidence level - 95%, 90%, 99% - is the part almost everyone can quote and almost everyone states incorrectly under interview pressure. The frequentist definition is about the **procedure**. A procedure that produces a 95% confidence interval is one with this property: if you were to repeat the whole exercise many times - draw a new sample from the same population, run the same recipe, get a new interval - then in the long run about 95% of the intervals produced would contain the true parameter value. That is called the **coverage probability** of the procedure, and 95% is the *nominal* level the procedure is designed to hit. ## Where the randomness lives This is the whole question in one sentence: in frequentist statistics the parameter is a fixed unknown constant, and the interval is random. The true population mean does not wobble from experiment to experiment; your sample does, so your estimate does, so the endpoints do. Picture 100 simulated intervals stacked as horizontal bars, with a vertical line drawn at the fixed true value. The line never moves. The bars slide left and right because each came from a different sample. About 95 of the bars cross the line; about 5 sit entirely to one side and miss. Nothing about the picture lets you point at one particular bar and say "this one has a 95% chance of crossing." It crossed or it did not; you simply cannot see which, because you cannot see the line. ## The sentence to avoid, and the sentence to use The classic misreading is: "there is a 95% probability that the true mean is between 4.1 and 5.3." It sounds harmless and it is what people mean colloquially, but it puts a probability distribution on the parameter, which frequentist coverage never does. After the data are in, both numbers in that sentence are fixed and the parameter is fixed, so the only honest probabilities are 0 or 1 - and you do not know which. Acceptable phrasings that keep the meaning straight: - "We are 95% confident the mean lies between 4.1 and 5.3," understood as shorthand for the repeated-sampling property. - "Values between 4.1 and 5.3 are the ones compatible with these data at the 95% level." - "This interval was produced by a method that covers the truth in 95% of repeated samples." ## What coverage does not promise 1. **It says nothing about the specific interval you computed.** Coverage is a long-run frequency of the recipe, not a score for one output. 2. **It is not a range for individual data points.** A 95% interval for a mean is usually far narrower than the spread of the raw observations; a range meant to contain a future observation is a different object entirely. 3. **It is not a ranking of the values inside it.** Frequentist coverage puts no probability on parameter values at all, so "4.5 is more likely than 5.2 because it is nearer the middle" is not a statement the interval supports. 4. **It is conditional on the assumptions.** Coverage is derived under a model: how the sample was drawn, independence between observations, and whatever approximation the recipe rests on. Sample non-randomly, ignore clustering in the data, or lean on a large-sample approximation at a tiny sample size, and the *actual* long-run coverage can sit well below the nominal 95%. The label on the tin does not enforce itself. ## Choosing the level The level is a dial you set before looking. Demanding 99% coverage buys you a stronger guarantee and pays for it with a wider, less informative range; accepting 90% gives a tighter statement that is wrong more often. There is no universally correct setting - 95% is a convention, not a theorem. ## How to answer this in an interview Say what is random and what is fixed, give the repeated-sampling sentence, then explicitly name the misreading and correct it. Interviewers ask this precisely because the wrong version is so fluent: a candidate who volunteers "the parameter is fixed, the interval is what moves" has demonstrated in one line that they understand the framework rather than the formula.

  • Someone writes 'there is a 95% probability the true mean is between 4.1 and 5.3'. What exactly is wrong with it?
    It treats the parameter as random. In this framework the true mean is a fixed unknown constant and 4.1 and 5.3 are fixed once the data are in, so the interval either contains it or it does not - the probability is 0 or 1, you just cannot tell which. The 95% belongs to the method across repeated samples, so the corrected sentence is that this interval came from a procedure covering the true mean 95% of the time.
  • If the modelling assumptions behind the interval are wrong, what happens to the 95% claim?
    It quietly stops being true. Coverage is derived under assumptions about how the sample was drawn and how the observations relate to each other. Break them - a biased sampling frame, correlated observations treated as independent, a large-sample approximation used on very few points - and the actual long-run coverage can fall well below the nominal level. The interval still prints the label 95%; only a simulation or a better model tells you whether it earns it.
  • How is a confidence interval for a mean different from a range meant to contain the next observation?
    They target different things. The confidence interval bounds a population parameter and shrinks as the sample grows, because the estimate gets more precise. A range for a future single observation must also absorb the spread of individuals, so it stays wide no matter how much data you collect. Confusing the two is why people are surprised that a 95% interval for a mean is far narrower than the data itself.

Think of ring toss at a fair. The peg is nailed down and never moves; each throw is a ring that lands somewhere. Saying the game is 95% accurate describes your throwing, not the peg's position - and once a ring has landed, it is either on the peg or it is not.

saying these in an interview costs you the question

  • Says there is a 95% probability the true mean lies in this interval
  • Treats the parameter as random and the computed interval as fixed
  • Claims 95% of the data values fall inside the interval
  • Assumes coverage holds even when the sampling assumptions are violated
  • Says the interval definitely contains the true value

context

open as a page

Why is a 99% confidence interval wider than a 95% one built from the same data?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Higher confidence needs a larger critical value. On a normal reference the two-sided values are 1.645 at 90%, 1.96 at 95% and 2.576 at 99%, so nothing but that multiplier changes and the interval stretches.

open as a page

Why does a sample p95 resist the closed-form confidence interval a sample mean gets?

level: middleimportance: must knowfreq 55%

basics

~20 s

A sample quantile's standard error contains the unknown probability density at that quantile: roughly sqrt(p(1-p)/n) divided by that density. The density has to be estimated from the few points sitting near a p95, so practitioners resample instead.

open as a page

When do you use a t critical value instead of z for a confidence interval for a mean?

level: middleimportance: must knowfreq 76%

basics

~20 s

Use t whenever the population standard deviation is unknown and estimated from the sample, which is almost always. The z shortcut is exact only when that spread is known, and adequate once the sample is large.

open as a page

Your resampled 95% interval for p99 latency from 500 requests is enormous - why?

level: seniorimportance: must knowfreq 45%

basics

~20 s

Only about five of 500 observations exceed the p99, so the estimate rests on two or three sorted values. Resampling only reshuffles those, giving a wide, chunky interval. The width is real information, not a defect.

open as a page

What does quantile regression at tau = 0.9 estimate that an OLS fit does not?

level: middleimportance: should knowfreq 26%

basics

~20 s

Quantile regression at tau = 0.9 estimates the conditional 90th percentile of the outcome; least squares estimates the conditional mean. Its coefficients say how a predictor moves the upper tail, which can differ from its effect on the average.

open as a page

How does a 95% confidence interval relate to a two-sided test at the 5% level?

level: middleimportance: should knowfreq 58%

basics

~20 s

They are two readings of one computation. A 95% confidence interval is the set of null values a two-sided test at the 5% level would not reject, so the interval excludes a value exactly when the test rejects that value.

open as a page

Two 95% confidence intervals overlap - does that prove the two group means do not differ significantly?

level: middleimportance: should knowfreq 52%

basics

~20 s

No. Overlapping intervals can still accompany a significant difference, because the difference has its own smaller standard error: variances add, then you take the square root. Test the difference directly instead of comparing two intervals by eye.

open as a page

How much more data does it take to halve the width of a confidence interval for a mean?

level: middleimportance: should knowfreq 52%

basics

~20 s

About four times as much. The half-width is a critical value times s over sqrt(n), so precision improves with the square root of the sample size; cutting the width in half requires roughly four times the observations.

open as a page

Why does a sample median's precision depend on distribution shape, not just on n?

level: seniorimportance: should knowfreq 32%

basics

~20 s

A sample median's standard error is about 1 divided by (2 times the density at the median times sqrt(n)). A peaked centre pins it down; a flat or hollow centre lets it wander at the same n.

open as a page

A 95% confidence interval for a treatment effect includes zero - does that prove there is no effect?

level: seniorimportance: should knowfreq 48%

basics

~20 s

No. Containing zero only means zero is among the values compatible with the data. Read the endpoints: a wide interval also contains large effects, so nothing is ruled out, while a tight interval around zero genuinely rules out anything big.

open as a page

With 0 successes in 20 trials, why does the Wald interval for a proportion collapse to zero width?

level: seniorimportance: should knowfreq 46%

basics

~20 s

The Wald formula plugs the observed proportion into its own standard error. With zero successes that term becomes 0, so the interval degenerates to [0, 0] - certainty from 20 trials. The Wilson score interval does not.

open as a page

A release moved p95 latency while the median stayed flat, so what changed and what do you do?

level: principalimportance: should knowfreq 28%

basics

~20 s

The distribution changed shape, not location: a pure shift moves every quantile together, so a tail-only move means part of the traffic degraded while typical traffic did not. Check the move against tail uncertainty, then segment.

open as a page

For a median from n = 20, how does an order-statistic interval get its coverage?

level: seniorimportance: nice to knowfreq 16%

basics

~20 s

Count how many observations fall below the population median: for continuous data that count is Binomial(20, 0.5) whatever the distribution. The 6th and 15th sorted values fail only when the count leaves 6 to 14, giving about 95.9% coverage.

open as a page

For the difference between two proportions, should the confidence interval use a pooled standard error?

level: seniorimportance: nice to knowfreq 32%

basics

~10 s

No. Pooling assumes the two proportions are equal, which is the null hypothesis a test assumes but an interval must not. Build the interval with each group's own observed proportion in the standard error.

open as a page

How do you stop stakeholders reading a 95% confidence interval as a probability about the true value?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Fix the sentence, not the audience. Standardise wording that says what the method does and what the range excludes, keep the estimate and its interval in real units, and insist on precision where a decision threshold falls inside the interval.

open as a page