Statistical Inference & Hypothesis Testing
You will learn how to go from a sample to a claim about a population: confidence intervals, p-values, the standard test families, error types, and power. Interviewers probe it relentlessly because misreading a p-value is the classic way analysts ship wrong conclusions.
on this pageshowhide
explore
- Sampling Distributions21 questions
- Standard Error5 questions
- Estimator Properties6 questions
- Maximum Likelihood Estimation5 questions
- Likelihood Asymptotics5 questions
- Confidence Intervals16 questions
- Interval Construction5 questions
- Coverage Interpretation5 questions
- Quantile Uncertainty6 questions
- Test Formulation10 questions
- Null and Alternative5 questions
- P-Value Meaning5 questions
- Test Families32 questions
- T-Tests and Z-Tests5 questions
- Chi-Square Tests5 questions
- Rank-Based Tests6 questions
- Assumption Checks6 questions
- ANOVA and F-Tests5 questions
- Distribution Comparison Tests5 questions
- Errors and Power14 questions
- Type I and II Errors4 questions
- Statistical Power5 questions
- Effect Size5 questions
- Multiplicity and Resampling20 questions
- Multiple Comparisons5 questions
- P-Hacking5 questions
- Bootstrap and Permutation5 questions
- Comparing Two Models5 questions
questions
113 · 6 sectionsWhat does it mean for a sample statistic to be an unbiased estimator of a population parameter?
basics
~20 sAn estimator is unbiased when its expected value across all possible samples equals the true parameter. Its estimates are centred on the target: too high as often, and by as much, as they are too low.
How does a likelihood differ from a probability in maximum likelihood estimation?
basics
~20 sA probability treats the parameter as fixed and asks how likely the data are; a likelihood fixes the observed data and reads the same formula as a function of the parameter. Likelihood values do not sum to one over parameters.
What does the standard error of the mean measure that the sample standard deviation does not?
basics
~20 sThe standard error of the mean measures how precisely a sample mean estimates the population mean; the sample standard deviation measures how spread out individual observations are. Only the standard error shrinks as the sample grows.
Why is the sum of squared deviations divided by n a biased estimator of the population variance?
basics
~20 sDeviations are measured from the sample mean, which is pulled toward the data and makes those deviations as small as they can possibly be. Dividing their squares by n therefore underestimates the population variance by the factor (n-1)/n.
What is Fisher information, and how does it relate to the score function of a log-likelihood?
basics
~20 sThe score is the derivative of the log-likelihood with respect to the parameter; Fisher information is the variance of the score at the true parameter, equivalently minus the expected second derivative. It measures how sharply data pin down the parameter.
What does the 95% in a 95% confidence interval actually refer to?
basics
~20 sThe 95% describes the procedure, not one interval. If you repeated the sampling and rebuilt the interval many times, about 95% of those intervals would contain the fixed true parameter. Any single interval either covers it or does not.
Why is a 99% confidence interval wider than a 95% one built from the same data?
basics
~20 sHigher confidence needs a larger critical value. On a normal reference the two-sided values are 1.645 at 90%, 1.96 at 95% and 2.576 at 99%, so nothing but that multiplier changes and the interval stretches.
Why does a sample p95 resist the closed-form confidence interval a sample mean gets?
basics
~20 sA sample quantile's standard error contains the unknown probability density at that quantile: roughly sqrt(p(1-p)/n) divided by that density. The density has to be estimated from the few points sitting near a p95, so practitioners resample instead.
When do you use a t critical value instead of z for a confidence interval for a mean?
basics
~20 sUse t whenever the population standard deviation is unknown and estimated from the sample, which is almost always. The z shortcut is exact only when that spread is known, and adequate once the sample is large.
Your resampled 95% interval for p99 latency from 500 requests is enormous - why?
basics
~20 sOnly about five of 500 observations exceed the p99, so the estimate rests on two or three sorted values. Resampling only reshuffles those, giving a wide, chunky interval. The width is real information, not a defect.
What is the difference between the null hypothesis and the alternative hypothesis?
basics
~20 sThe null hypothesis (H0) is the default claim of no effect or no difference, written as an equality so its sampling distribution can be computed. The alternative (H1) is the competing claim that only evidence in the data can support.
What does a p-value of 0.03 from a two-sample t-test actually mean?
basics
~20 sA p-value of 0.03 says that if the null hypothesis of no difference were true, data at least this extreme would arise about 3% of the time. It is not the probability that the null hypothesis is true.
When should a hypothesis test use a one-tailed alternative rather than a two-tailed one?
basics
~20 sUse a one-tailed alternative only when a result in the opposite direction would be acted on exactly like no effect, and only when the direction is fixed before the data. Otherwise two-tailed is the honest default.
Why can a negligible effect give p = 0.4 at n = 100 but p = 0.001 at n = 1,000,000?
basics
~20 sA p-value blends how big an effect is with how precisely it was measured. Standard error shrinks roughly as one over the square root of the sample size, so at a million observations even a trivial difference becomes statistically detectable.
Why do we say we fail to reject the null hypothesis instead of accepting it?
basics
~20 sA test controls only the risk of wrongly rejecting a true null; it gives no guarantee that a null which survives is true. A non-significant result is a verdict of not proven, not evidence of no effect.
What does a one-way ANOVA test, and what are its null and alternative hypotheses?
basics
~20 sOne-way ANOVA tests whether the population means of three or more groups defined by a single factor are all equal. The null says every group mean is the same; the alternative says at least one differs.
Which assumptions must hold for a two-sample t-test comparing group means to be valid?
basics
~20 sA two-sample t-test assumes observations are independent within and between groups, that the sampling distribution of each group mean is approximately normal, and, for the pooled version, that the two populations share a common variance.
How does a chi-square goodness-of-fit test decide whether 600 die rolls came from a fair die?
basics
~20 sA chi-square goodness-of-fit test compares observed category counts with the counts a hypothesised distribution predicts. A fair die over 600 rolls predicts 100 per face; the statistic sums (observed minus expected) squared, divided by expected, across the six faces.
What does a two-sample Kolmogorov-Smirnov test compare between two samples?
basics
~20 sA two-sample Kolmogorov-Smirnov test compares the two samples' entire distributions. It builds a step-shaped empirical CDF for each sample and takes the largest vertical gap between them, so a difference in location, spread or shape can all trigger rejection.
Why are rank-based tests barely affected by a single mistyped value of 10,000?
basics
~20 sRank-based tests replace each value with its position in sorted order, so a mistyped 10,000 becomes only the largest rank, not a huge number. Its influence is capped at one rank; a mean and a variance have no such cap.
What is the difference between a Type I and a Type II error in hypothesis testing?
basics
~20 sA Type I error rejects a null hypothesis that is actually true, a false alarm whose long-run rate is alpha. A Type II error fails to reject a null hypothesis that is actually false, a missed effect whose rate is beta.
What does it mean to say a hypothesis test has 80% power?
basics
~20 sPower is the probability a test rejects the null when a specified real effect exists. 80% power means that if an effect of that size is truly present, the test detects it 80% of the time.
How do you compute Cohen's d for a 3-point mean gap when the pooled SD is 15?
basics
~20 sDivide the difference in means by the pooled standard deviation: 3 / 15 = 0.2. The two groups sit one fifth of a standard deviation apart, which Cohen's conventional benchmarks call a small effect, and their distributions overlap heavily.
How does lowering the significance level from 0.05 to 0.01 affect Type II errors at a fixed sample size?
basics
~20 sLowering the significance level raises the Type II error rate. Demanding stronger evidence shrinks the rejection region, so with the same data and the same true effect, more real effects fall short of the threshold and go undetected.
What four levers determine statistical power in a hypothesis test?
basics
~20 sPower rises with sample size, with the size of the true effect, with a larger alpha, and with lower outcome variance. Three of those are design choices; the true effect is not yours to set, only to assume honestly.
What is bootstrap resampling, and how does it produce a confidence interval for a correlation?
basics
~20 sThe bootstrap treats your sample as a stand-in for the population: draw many new samples of the same size with replacement, recompute the statistic on each, and read the interval off the middle 95% of those values.
If you run 20 independent hypothesis tests at alpha = 0.05, how likely is at least one false positive?
basics
~20 sAbout 64 percent. Each test independently avoids a false positive with probability 0.95, so all 20 stay clean with probability 0.95^20 = 0.36, leaving roughly a 64 percent chance that at least one comes back significant by luck.
Model B beats model A by 0.3 accuracy points on a 1,000-example test set - is that real?
basics
~20 sNot on that evidence. On 1,000 examples, 0.3 accuracy points is three examples, and a net of three can never reach statistical significance in a paired comparison. The gap sits inside the test set's noise.
How does a permutation test build a null distribution, and what does exchangeability require?
basics
~20 sA permutation test shuffles the group labels thousands of times, recomputing the test statistic on each shuffle to build the null distribution from the data itself. The p-value is the share of shuffles at least as extreme as the observed statistic.
Why is the Bonferroni correction called conservative when testing 15 endpoints at once?
basics
~20 sBonferroni tests each of the 15 endpoints at 0.05/15 = 0.0033 instead of 0.05. That keeps the family-wise error rate under 5 percent whatever the dependence between endpoints, but the tiny per-test threshold makes real effects far harder to detect.