skip to content

Statistical Inference & Hypothesis Testing

You will learn how to go from a sample to a claim about a population: confidence intervals, p-values, the standard test families, error types, and power. Interviewers probe it relentlessly because misreading a p-value is the classic way analysts ship wrong conclusions.

on this pageshow

explore

questions

113 · 6 sections

What does it mean for a sample statistic to be an unbiased estimator of a population parameter?

level: juniorimportance: must knowfreq 70%
basics
~20 s

An estimator is unbiased when its expected value across all possible samples equals the true parameter. Its estimates are centred on the target: too high as often, and by as much, as they are too low.

open as a page

How does a likelihood differ from a probability in maximum likelihood estimation?

level: juniorimportance: must knowfreq 80%
basics
~20 s

A probability treats the parameter as fixed and asks how likely the data are; a likelihood fixes the observed data and reads the same formula as a function of the parameter. Likelihood values do not sum to one over parameters.

open as a page

What does the standard error of the mean measure that the sample standard deviation does not?

level: juniorimportance: must knowfreq 82%
basics
~20 s

The standard error of the mean measures how precisely a sample mean estimates the population mean; the sample standard deviation measures how spread out individual observations are. Only the standard error shrinks as the sample grows.

open as a page

Why is the sum of squared deviations divided by n a biased estimator of the population variance?

level: middleimportance: must knowfreq 74%
basics
~20 s

Deviations are measured from the sample mean, which is pulled toward the data and makes those deviations as small as they can possibly be. Dividing their squares by n therefore underestimates the population variance by the factor (n-1)/n.

open as a page

What is Fisher information, and how does it relate to the score function of a log-likelihood?

level: middleimportance: must knowfreq 58%
basics
~20 s

The score is the derivative of the log-likelihood with respect to the parameter; Fisher information is the variance of the score at the true parameter, equivalently minus the expected second derivative. It measures how sharply data pin down the parameter.

open as a page

What does the 95% in a 95% confidence interval actually refer to?

level: juniorimportance: must knowfreq 88%
basics
~20 s

The 95% describes the procedure, not one interval. If you repeated the sampling and rebuilt the interval many times, about 95% of those intervals would contain the fixed true parameter. Any single interval either covers it or does not.

open as a page

Why is a 99% confidence interval wider than a 95% one built from the same data?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Higher confidence needs a larger critical value. On a normal reference the two-sided values are 1.645 at 90%, 1.96 at 95% and 2.576 at 99%, so nothing but that multiplier changes and the interval stretches.

open as a page

Why does a sample p95 resist the closed-form confidence interval a sample mean gets?

level: middleimportance: must knowfreq 55%
basics
~20 s

A sample quantile's standard error contains the unknown probability density at that quantile: roughly sqrt(p(1-p)/n) divided by that density. The density has to be estimated from the few points sitting near a p95, so practitioners resample instead.

open as a page

When do you use a t critical value instead of z for a confidence interval for a mean?

level: middleimportance: must knowfreq 76%
basics
~20 s

Use t whenever the population standard deviation is unknown and estimated from the sample, which is almost always. The z shortcut is exact only when that spread is known, and adequate once the sample is large.

open as a page

Your resampled 95% interval for p99 latency from 500 requests is enormous - why?

level: seniorimportance: must knowfreq 45%
basics
~20 s

Only about five of 500 observations exceed the p99, so the estimate rests on two or three sorted values. Resampling only reshuffles those, giving a wide, chunky interval. The width is real information, not a defect.

open as a page

What is the difference between the null hypothesis and the alternative hypothesis?

level: juniorimportance: must knowfreq 85%
basics
~20 s

The null hypothesis (H0) is the default claim of no effect or no difference, written as an equality so its sampling distribution can be computed. The alternative (H1) is the competing claim that only evidence in the data can support.

open as a page

What does a p-value of 0.03 from a two-sample t-test actually mean?

level: juniorimportance: must knowfreq 88%
basics
~20 s

A p-value of 0.03 says that if the null hypothesis of no difference were true, data at least this extreme would arise about 3% of the time. It is not the probability that the null hypothesis is true.

open as a page

When should a hypothesis test use a one-tailed alternative rather than a two-tailed one?

level: middleimportance: must knowfreq 66%
basics
~20 s

Use a one-tailed alternative only when a result in the opposite direction would be acted on exactly like no effect, and only when the direction is fixed before the data. Otherwise two-tailed is the honest default.

open as a page

Why can a negligible effect give p = 0.4 at n = 100 but p = 0.001 at n = 1,000,000?

level: middleimportance: must knowfreq 66%
basics
~20 s

A p-value blends how big an effect is with how precisely it was measured. Standard error shrinks roughly as one over the square root of the sample size, so at a million observations even a trivial difference becomes statistically detectable.

open as a page

Why do we say we fail to reject the null hypothesis instead of accepting it?

level: middleimportance: should knowfreq 58%
basics
~20 s

A test controls only the risk of wrongly rejecting a true null; it gives no guarantee that a null which survives is true. A non-significant result is a verdict of not proven, not evidence of no effect.

open as a page

What does a one-way ANOVA test, and what are its null and alternative hypotheses?

level: juniorimportance: must knowfreq 76%
basics
~20 s

One-way ANOVA tests whether the population means of three or more groups defined by a single factor are all equal. The null says every group mean is the same; the alternative says at least one differs.

open as a page

Which assumptions must hold for a two-sample t-test comparing group means to be valid?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A two-sample t-test assumes observations are independent within and between groups, that the sampling distribution of each group mean is approximately normal, and, for the pooled version, that the two populations share a common variance.

open as a page

How does a chi-square goodness-of-fit test decide whether 600 die rolls came from a fair die?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A chi-square goodness-of-fit test compares observed category counts with the counts a hypothesised distribution predicts. A fair die over 600 rolls predicts 100 per face; the statistic sums (observed minus expected) squared, divided by expected, across the six faces.

open as a page

What does a two-sample Kolmogorov-Smirnov test compare between two samples?

level: juniorimportance: must knowfreq 58%
basics
~20 s

A two-sample Kolmogorov-Smirnov test compares the two samples' entire distributions. It builds a step-shaped empirical CDF for each sample and takes the largest vertical gap between them, so a difference in location, spread or shape can all trigger rejection.

open as a page

Why are rank-based tests barely affected by a single mistyped value of 10,000?

level: juniorimportance: must knowfreq 60%
basics
~20 s

Rank-based tests replace each value with its position in sorted order, so a mistyped 10,000 becomes only the largest rank, not a huge number. Its influence is capped at one rank; a mean and a variance have no such cap.

open as a page

What is the difference between a Type I and a Type II error in hypothesis testing?

level: juniorimportance: must knowfreq 88%
basics
~20 s

A Type I error rejects a null hypothesis that is actually true, a false alarm whose long-run rate is alpha. A Type II error fails to reject a null hypothesis that is actually false, a missed effect whose rate is beta.

open as a page

What does it mean to say a hypothesis test has 80% power?

level: juniorimportance: must knowfreq 85%
basics
~20 s

Power is the probability a test rejects the null when a specified real effect exists. 80% power means that if an effect of that size is truly present, the test detects it 80% of the time.

open as a page

How do you compute Cohen's d for a 3-point mean gap when the pooled SD is 15?

level: middleimportance: must knowfreq 74%
basics
~20 s

Divide the difference in means by the pooled standard deviation: 3 / 15 = 0.2. The two groups sit one fifth of a standard deviation apart, which Cohen's conventional benchmarks call a small effect, and their distributions overlap heavily.

open as a page

How does lowering the significance level from 0.05 to 0.01 affect Type II errors at a fixed sample size?

level: middleimportance: must knowfreq 68%
basics
~20 s

Lowering the significance level raises the Type II error rate. Demanding stronger evidence shrinks the rejection region, so with the same data and the same true effect, more real effects fall short of the threshold and go undetected.

open as a page

What four levers determine statistical power in a hypothesis test?

level: middleimportance: must knowfreq 70%
basics
~20 s

Power rises with sample size, with the size of the true effect, with a larger alpha, and with lower outcome variance. Three of those are design choices; the true effect is not yours to set, only to assume honestly.

open as a page

What is bootstrap resampling, and how does it produce a confidence interval for a correlation?

level: juniorimportance: must knowfreq 62%
basics
~20 s

The bootstrap treats your sample as a stand-in for the population: draw many new samples of the same size with replacement, recompute the statistic on each, and read the interval off the middle 95% of those values.

open as a page

If you run 20 independent hypothesis tests at alpha = 0.05, how likely is at least one false positive?

level: juniorimportance: must knowfreq 76%
basics
~20 s

About 64 percent. Each test independently avoids a false positive with probability 0.95, so all 20 stay clean with probability 0.95^20 = 0.36, leaving roughly a 64 percent chance that at least one comes back significant by luck.

open as a page

Model B beats model A by 0.3 accuracy points on a 1,000-example test set - is that real?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Not on that evidence. On 1,000 examples, 0.3 accuracy points is three examples, and a net of three can never reach statistical significance in a paired comparison. The gap sits inside the test set's noise.

open as a page

How does a permutation test build a null distribution, and what does exchangeability require?

level: middleimportance: must knowfreq 55%
basics
~20 s

A permutation test shuffles the group labels thousands of times, recomputing the test statistic on each shuffle to build the null distribution from the data itself. The p-value is the share of shuffles at least as extreme as the observed statistic.

open as a page

Why is the Bonferroni correction called conservative when testing 15 endpoints at once?

level: middleimportance: must knowfreq 68%
basics
~20 s

Bonferroni tests each of the 15 endpoints at 0.05/15 = 0.0033 instead of 0.05. That keeps the family-wise error rate under 5 percent whatever the dependence between endpoints, but the tiny per-test threshold makes real effects far harder to detect.

open as a page