Math, Statistics & Experimentation
This is the math and statistics layer every data and ML interview leans on: probability, hypothesis testing, regression, and how to design and read an A/B test. Interviewers probe it before any tooling because it separates people who can reason about data from people who can only run libraries.
on this pageshowhide
explore
- Linear Algebra for ML40 questions
- Vectors and Matrix Operations10 questions
- Rank and Linear Systems10 questions
- Eigendecomposition and SVD11 questions
- Norms, Distance and Projections9 questions
- Calculus & Optimization Basics45 questions
- Derivatives & Gradients11 questions
- Chain Rule & Curvature9 questions
- Convexity & Optimality10 questions
- Descent & Constraints15 questions
- Descriptive Statistics & Sampling85 questions
- Center, Spread and Scale14 questions
- Distribution Shape & Outliers15 questions
- Association & Correlation25 questions
- Sampling Design & Bias15 questions
- Diagnosing a Metric Change6 questions
- Back-of-Envelope Estimation5 questions
- Metric Decomposition & Mix Shift5 questions
- Probability & Distributions92 questions
- Combinatorics & Event Rules11 questions
- Conditioning & Bayes' Rule15 questions
- Random Variables & Moments27 questions
- Distribution Families16 questions
- Limit Theorems & Simulation23 questions
- Statistical Inference & Hypothesis Testing113 questions
- Sampling Distributions21 questions
- Confidence Intervals16 questions
- Test Formulation10 questions
- Test Families32 questions
- Errors and Power14 questions
- Multiplicity and Resampling20 questions
- Regression & Statistical Modeling82 questions
- Ordinary Least Squares15 questions
- Coefficient Interpretation15 questions
- Model Diagnostics26 questions
- Goodness of Fit6 questions
- Generalized Linear Models20 questions
- Bayesian Statistics60 questions
- Priors and Likelihoods21 questions
- Posterior Computation15 questions
- Posterior Summaries13 questions
- Posterior Decision Rules11 questions
- A/B Testing & Experiment Design112 questions
- Randomization & Assignment22 questions
- Sample Size Planning10 questions
- Metrics & Guardrails30 questions
- Readout & Decision Rules15 questions
- Validity Threats20 questions
- Sensitivity & Adaptive Designs15 questions
- Causal Inference68 questions
- Potential Outcomes10 questions
- Confounding Structures16 questions
- Adjustment Estimators15 questions
- Quasi-Experimental Designs17 questions
- Heterogeneity and Credibility10 questions
- Time-Series Analysis53 questions
- Series Structure12 questions
- Classical Forecast Models15 questions
- Forecast Evaluation10 questions
- Anomalies and Changepoints5 questions
- Lag Features and Horizons6 questions
- Lead-Lag and Cointegration5 questions
questions
750 · 10 sectionsHow does cosine similarity differ from Euclidean distance between two vectors?
basics
~20 sCosine similarity measures only the angle between two vectors; Euclidean distance also reacts to their magnitudes. Two term-count vectors with the same word mix but different document lengths score cosine 1.0 while sitting far apart in Euclidean distance.
What does it mean for a vector to be an eigenvector of a matrix A?
basics
~20 sA nonzero vector v is an eigenvector of A when Av = lambda v. Multiplying by A leaves v on its own line through the origin, only stretching, shrinking or flipping it; the scalar lambda is that scale factor.
What does a determinant of zero tell you about a square matrix?
basics
~20 sA zero determinant means the matrix is singular: no inverse exists. For a 2x2 matrix [[a, b], [c, d]] the determinant is ad - bc, so ad - bc = 0 is the exact test for non-invertibility.
For the matrix product AB, what shapes must A and B have, and what shape is the result?
basics
~20 sMatrix multiplication needs matching inner dimensions: if A is m x n and B is n x p, then AB is m x p. Entry (i,j) is the dot product of row i of A with column j of B.
When does the linear system Ax = b have no solution, exactly one, or infinitely many?
basics
~20 sRun Gaussian elimination on the augmented matrix. A row that is all zeros on the left but nonzero on the right means no solution. Otherwise, one pivot per unknown means exactly one solution; any free unknown means infinitely many.
What is the formal definition of a convex set?
basics
~20 sA set is convex if the straight segment joining any two of its points lies entirely inside it: for every x and y in the set C and every t in [0, 1], the point t*x + (1-t)*y is also in C.
Why is a zero gradient necessary but not sufficient for a local minimum?
basics
~20 sAt a smooth interior local minimum the gradient must vanish, so a zero gradient is necessary. But maxima, saddle points and flat inflection points also have a zero gradient, so vanishing slope on its own certifies nothing.
Using the limit definition of the derivative, what is the derivative of f(x) = x^2?
basics
~10 sThe derivative of x^2 is 2x. The difference quotient ((x+h)^2 - x^2)/h expands to (2xh + h^2)/h, which simplifies to 2x + h, and letting h shrink to 0 leaves 2x.
In gradient descent, what does the update rule x = x - eta * gradient actually do at each step?
basics
~20 sEach iteration moves every parameter a short distance opposite its own partial derivative: new value = old value minus eta times the gradient. The learning rate eta scales how far you move; the loop repeats until the gradient is near zero.
For f(x, y) = x^2 * y, what are the two partial derivatives and the gradient at (2, 3)?
basics
~20 sDifferentiate one variable at a time, holding the other fixed: df/dx = 2xy and df/dy = x^2. At (2, 3) those evaluate to 12 and 4, so the gradient there is the vector (12, 4).
How would you estimate the number of piano tuners working in Chicago from scratch?
basics
~20 sBreak the target number into a chain of estimable factors: city population, people per household, share of households owning a piano, tunings per piano per year, and tunings one tuner performs per year. Multiply through, divide, and land near 50.
In a cross-tab of device type by plan, what is the difference between joint, marginal and conditional proportions?
basics
~20 sA joint proportion divides a cell count by the grand total. A marginal proportion divides a row or column total by the grand total. A conditional proportion divides a cell by its own row or column total.
Ice-cream sales correlate with drowning deaths — what can and cannot be concluded from that?
basics
~20 sA correlation only says the two series move up and down together in the observed data. It cannot say ice cream causes drownings: something else, such as hot weather, can drive both, and correlation carries no direction while causation does.
What is the difference between sample covariance and Pearson's correlation coefficient r?
basics
~20 sSample covariance measures whether two variables move together, but it carries the product of their units, so its size means little alone. Pearson's r divides covariance by both standard deviations, giving a unitless number between -1 and 1.
For a right-skewed column like household income, why does the mean exceed the median?
basics
~20 sThe mean sums every value, so a long right tail of very high incomes pulls it upward. The median depends only on rank, so extreme values barely move it. Mean above median is the usual signature of right skew.
In the Monty Hall problem, why does switching doors win two-thirds of the time?
basics
~20 sYour first pick wins only one time in three, so two times in three the car is behind another door. The host, who knows where it is, opens a losing door and concentrates that 2/3 onto the single door left.
What does the Central Limit Theorem say about the average of many independent samples?
basics
~10 sThe Central Limit Theorem says that averaging many independent draws from almost any distribution with finite variance produces an average whose distribution is approximately normal, even when the individual observations are not remotely normal.
Test scores are normal with mean 100 and SD 15: by the 68-95-99.7 rule, what share exceeds 130?
basics
~20 sAbout 2.5%. A score of 130 is two standard deviations above the mean, so its z-score is 2; the rule puts roughly 95% of values within two SDs, leaving about 5% split evenly between the two tails.
What are the mean and variance of a binomial random variable with n trials and success probability p?
basics
~10 sA binomial count over n independent trials with success probability p has mean np and variance np(1-p). It is the sum of n Bernoulli(p) indicators, each contributing mean p and variance p(1-p).
For a Poisson count averaging 3 support tickets per hour, what is the probability of zero tickets in an hour?
basics
~20 sThe Poisson mass function is P(X = k) = e^(-λ) λ^k / k!. With λ = 3 tickets per hour, P(X = 0) = e^(-3), about 0.0498, so roughly a 5 percent chance of a completely quiet hour.
Statistical Inference & Hypothesis Testing
all 113 Statistical Inference & Hypothesis Testing questions →What does a one-way ANOVA test, and what are its null and alternative hypotheses?
basics
~20 sOne-way ANOVA tests whether the population means of three or more groups defined by a single factor are all equal. The null says every group mean is the same; the alternative says at least one differs.
Which assumptions must hold for a two-sample t-test comparing group means to be valid?
basics
~20 sA two-sample t-test assumes observations are independent within and between groups, that the sampling distribution of each group mean is approximately normal, and, for the pooled version, that the two populations share a common variance.
What is bootstrap resampling, and how does it produce a confidence interval for a correlation?
basics
~20 sThe bootstrap treats your sample as a stand-in for the population: draw many new samples of the same size with replacement, recompute the statistic on each, and read the interval off the middle 95% of those values.
How does a chi-square goodness-of-fit test decide whether 600 die rolls came from a fair die?
basics
~20 sA chi-square goodness-of-fit test compares observed category counts with the counts a hypothesised distribution predicts. A fair die over 600 rolls predicts 100 per face; the statistic sums (observed minus expected) squared, divided by expected, across the six faces.
What does the 95% in a 95% confidence interval actually refer to?
basics
~20 sThe 95% describes the procedure, not one interval. If you repeated the sampling and rebuilt the interval many times, about 95% of those intervals would contain the fixed true parameter. Any single interval either covers it or does not.
Why is treating three repeated blood-pressure readings per patient as three independent observations wrong?
basics
~20 sReadings from one patient are correlated, so three readings carry far less information than three different patients. Counting them as independent inflates the sample size, shrinks the standard errors and p-values, and manufactures significance that is not there.
In a regression output, what does the standard error of a coefficient tell you?
basics
~20 sThe standard error of a regression coefficient measures how much that estimate would move around if you refit the model on other samples from the same process. A smaller standard error means a more precisely estimated effect.
With day-of-week dummies and Monday as the reference level, what does the Saturday coefficient mean?
basics
~20 sThe Saturday coefficient is the estimated difference in the outcome between Saturday and Monday, with the model's other predictors held fixed. It is a contrast against the omitted baseline day, not Saturday's own average level.
In a subscription churn analysis, how should you record a user who signed up 21 days ago and is still subscribed?
basics
~20 sRecord the user as right-censored at 21 days: the subscription lasted at least 21 days, and the eventual cancellation time is unknown. Deleting the row biases survival downward; coding it as a cancellation invents an event that never happened.
What does the R-squared of a fitted multiple regression tell you about the model?
basics
~20 sR-squared is the share of the outcome's total variation that the fitted model accounts for, computed as 1 minus the residual sum of squares over the total sum of squares. It measures fit on the data used, not correctness.
In a Bayesian A/B test, what does 'probability B beats A is 96%' actually mean?
basics
~20 sIt is the posterior probability that variant B's true rate is higher than variant A's, given the data and the prior. It is a claim about the unknown rates, and says nothing about how large the difference is.
What is the difference between probability as long-run frequency and probability as degree of belief?
basics
~20 sThe frequency reading defines probability as the proportion of times an outcome occurs in repeatable trials. The belief reading defines it as a numeric degree of confidence, so it can also score one-off events such as a single rocket launch.
Starting from a Beta(1,1) prior, what posterior follows from 8 clicks in 100 impressions?
basics
~10 sBeta(1 + 8, 1 + 92), that is Beta(9, 93). The prior contributes one pseudo-success and one pseudo-failure, so the posterior mean is 9/102, about 0.088, slightly above the raw rate of 0.08.
What does a 95% Bayesian credible interval of [1.9%, 2.5%] say about a conversion rate?
basics
~20 sIt says 95% of the posterior probability for the conversion rate falls between 1.9% and 2.5%. Given the model and prior you used, there is a 95% probability the rate lies in that range. The probability attaches to the parameter.
What is the difference between probability and likelihood after seeing 7 heads in 10 coin flips?
basics
~20 sProbability fixes the coin's bias and varies the data; likelihood fixes the observed data and varies the bias. After 7 heads in 10 flips the likelihood is a curve over p that is not a probability distribution.
What is a guardrail metric in an A/B test, and how does it differ from a driver metric?
basics
~20 sA guardrail is a metric the experiment must not damage — latency, crash-free sessions, unsubscribes — and a big enough degradation can veto the launch. A driver metric only explains why the primary metric moved; it never vetoes.
Why does revenue per user make an A/B test hard to call when a few users spend far more than the rest?
basics
~10 sAlmost all the variance in revenue per user comes from a handful of very large spenders, so the interval around the treatment-control difference stays wide. Ordinary experiment sizes then cannot resolve realistic revenue effects.
What is an Overall Evaluation Criterion (OEC) in an A/B test?
basics
~20 sThe Overall Evaluation Criterion is the single metric, agreed before launch, that an experiment's ship-or-not decision is defined against. It expresses what the team means by the product getting better, so the result cannot be reinterpreted afterwards.
What does a DAU/MAU stickiness ratio of 0.2 tell you about how a product is used?
basics
~10 sA DAU/MAU of 0.2 means the typical monthly active user opens the product on about 6 days out of 30. It measures return frequency, not audience size, and averages over very different user types.
What does a difference-in-differences estimate compute from a two-group, two-period panel?
basics
~20 sIt subtracts the comparison group's before-to-after change from the treated group's before-to-after change. That double difference cancels the fixed level gap between the groups and any shock that moved both of them over the same window.
Why can a treatment with an average effect of zero still help some users and harm others?
basics
~20 sAn average pools opposite effects. A drug that raises recovery for patients under 50 and lowers it by a similar amount for patients over 70 posts an overall average near zero while both real effects remain untouched underneath it.
What is a propensity score, and why match on it rather than on the raw covariates?
basics
~20 sA propensity score is a unit's probability of receiving the treatment given its observed covariates. Conditioning on that single number balances the covariates that went into it, so matching happens in one dimension instead of many.
What does adding a confounder to a regression do to the treatment coefficient?
basics
~20 sAdding a measured confounder turns the treatment coefficient from a raw comparison into a within-strata one: it now compares treated and untreated units that share the same covariate value, removing the part of the gap that covariate explained.
What does a sensitivity analysis for unmeasured confounding tell you about a causal estimate?
basics
~10 sA sensitivity analysis says how strong an unmeasured confounder would have to be to overturn the estimate. It never shows confounding is absent; it prices how much hidden bias the finding can tolerate.
Why does MAPE break down on intermittent demand series with many zero-sales days?
basics
~20 sMAPE divides each absolute error by the actual value, so a zero-sales day makes that term undefined and a near-zero day makes it explode. On intermittent demand the average is dominated by a handful of tiny denominators.
In a time series, what is the difference between a point outlier and a level shift?
basics
~20 sA point outlier is one observation far from expectation, after which the series returns to its old level. A level shift moves the series to a new baseline that persists. The difference decides whether you drop a point or re-baseline.
Why is a shuffled random train/test split invalid for evaluating a daily demand forecast?
basics
~20 sA shuffled split trains on future days and tests on past ones, so the model interpolates between neighbouring dates instead of forecasting. Because adjacent days are highly correlated, the score looks excellent and says nothing about future performance.
In a classical time-series decomposition, what do the trend, seasonal and remainder components each represent?
basics
~20 sDecomposition splits a series into three parts: trend-cycle, the slow movement of the level; seasonal, the pattern that repeats at a fixed known period such as 12 months; and remainder, the variation left after removing both.
In simple exponential smoothing, what does the smoothing parameter alpha control?
basics
~20 sAlpha sets how much weight the update puts on the newest observation versus the accumulated past. Alpha near 1 tracks recent data and reacts fast; alpha near 0 averages over a long history and smooths noise.