skip to content

Math, Statistics & Experimentation

This is the math and statistics layer every data and ML interview leans on: probability, hypothesis testing, regression, and how to design and read an A/B test. Interviewers probe it before any tooling because it separates people who can reason about data from people who can only run libraries.

on this pageshow

explore

questions

750 · 10 sections

How does cosine similarity differ from Euclidean distance between two vectors?

level: juniorimportance: must knowfreq 80%
basics
~20 s

Cosine similarity measures only the angle between two vectors; Euclidean distance also reacts to their magnitudes. Two term-count vectors with the same word mix but different document lengths score cosine 1.0 while sitting far apart in Euclidean distance.

open as a page

What does it mean for a vector to be an eigenvector of a matrix A?

level: juniorimportance: must knowfreq 84%
basics
~20 s

A nonzero vector v is an eigenvector of A when Av = lambda v. Multiplying by A leaves v on its own line through the origin, only stretching, shrinking or flipping it; the scalar lambda is that scale factor.

open as a page

What does a determinant of zero tell you about a square matrix?

level: juniorimportance: must knowfreq 76%
basics
~20 s

A zero determinant means the matrix is singular: no inverse exists. For a 2x2 matrix [[a, b], [c, d]] the determinant is ad - bc, so ad - bc = 0 is the exact test for non-invertibility.

open as a page

For the matrix product AB, what shapes must A and B have, and what shape is the result?

level: juniorimportance: must knowfreq 86%
basics
~20 s

Matrix multiplication needs matching inner dimensions: if A is m x n and B is n x p, then AB is m x p. Entry (i,j) is the dot product of row i of A with column j of B.

open as a page

When does the linear system Ax = b have no solution, exactly one, or infinitely many?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Run Gaussian elimination on the augmented matrix. A row that is all zeros on the left but nonzero on the right means no solution. Otherwise, one pivot per unknown means exactly one solution; any free unknown means infinitely many.

open as a page

What is the formal definition of a convex set?

level: juniorimportance: must knowfreq 66%
basics
~20 s

A set is convex if the straight segment joining any two of its points lies entirely inside it: for every x and y in the set C and every t in [0, 1], the point t*x + (1-t)*y is also in C.

open as a page

Why is a zero gradient necessary but not sufficient for a local minimum?

level: juniorimportance: must knowfreq 82%
basics
~20 s

At a smooth interior local minimum the gradient must vanish, so a zero gradient is necessary. But maxima, saddle points and flat inflection points also have a zero gradient, so vanishing slope on its own certifies nothing.

open as a page

Using the limit definition of the derivative, what is the derivative of f(x) = x^2?

level: juniorimportance: must knowfreq 72%
basics
~10 s

The derivative of x^2 is 2x. The difference quotient ((x+h)^2 - x^2)/h expands to (2xh + h^2)/h, which simplifies to 2x + h, and letting h shrink to 0 leaves 2x.

open as a page

In gradient descent, what does the update rule x = x - eta * gradient actually do at each step?

level: juniorimportance: must knowfreq 88%
basics
~20 s

Each iteration moves every parameter a short distance opposite its own partial derivative: new value = old value minus eta times the gradient. The learning rate eta scales how far you move; the loop repeats until the gradient is near zero.

open as a page

For f(x, y) = x^2 * y, what are the two partial derivatives and the gradient at (2, 3)?

level: juniorimportance: must knowfreq 82%
basics
~20 s

Differentiate one variable at a time, holding the other fixed: df/dx = 2xy and df/dy = x^2. At (2, 3) those evaluate to 12 and 4, so the gradient there is the vector (12, 4).

open as a page

How would you estimate the number of piano tuners working in Chicago from scratch?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Break the target number into a chain of estimable factors: city population, people per household, share of households owning a piano, tunings per piano per year, and tunings one tuner performs per year. Multiply through, divide, and land near 50.

open as a page

In a cross-tab of device type by plan, what is the difference between joint, marginal and conditional proportions?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A joint proportion divides a cell count by the grand total. A marginal proportion divides a row or column total by the grand total. A conditional proportion divides a cell by its own row or column total.

open as a page

Ice-cream sales correlate with drowning deaths — what can and cannot be concluded from that?

level: juniorimportance: must knowfreq 85%
basics
~20 s

A correlation only says the two series move up and down together in the observed data. It cannot say ice cream causes drownings: something else, such as hot weather, can drive both, and correlation carries no direction while causation does.

open as a page

What is the difference between sample covariance and Pearson's correlation coefficient r?

level: juniorimportance: must knowfreq 86%
basics
~20 s

Sample covariance measures whether two variables move together, but it carries the product of their units, so its size means little alone. Pearson's r divides covariance by both standard deviations, giving a unitless number between -1 and 1.

open as a page

For a right-skewed column like household income, why does the mean exceed the median?

level: juniorimportance: must knowfreq 84%
basics
~20 s

The mean sums every value, so a long right tail of very high incomes pulls it upward. The median depends only on rank, so extreme values barely move it. Mean above median is the usual signature of right skew.

open as a page

In the Monty Hall problem, why does switching doors win two-thirds of the time?

level: juniorimportance: must knowfreq 64%
basics
~20 s

Your first pick wins only one time in three, so two times in three the car is behind another door. The host, who knows where it is, opens a losing door and concentrates that 2/3 onto the single door left.

open as a page

What does the Central Limit Theorem say about the average of many independent samples?

level: juniorimportance: must knowfreq 85%
basics
~10 s

The Central Limit Theorem says that averaging many independent draws from almost any distribution with finite variance produces an average whose distribution is approximately normal, even when the individual observations are not remotely normal.

open as a page

Test scores are normal with mean 100 and SD 15: by the 68-95-99.7 rule, what share exceeds 130?

level: juniorimportance: must knowfreq 82%
basics
~20 s

About 2.5%. A score of 130 is two standard deviations above the mean, so its z-score is 2; the rule puts roughly 95% of values within two SDs, leaving about 5% split evenly between the two tails.

open as a page

What are the mean and variance of a binomial random variable with n trials and success probability p?

level: juniorimportance: must knowfreq 78%
basics
~10 s

A binomial count over n independent trials with success probability p has mean np and variance np(1-p). It is the sum of n Bernoulli(p) indicators, each contributing mean p and variance p(1-p).

open as a page

For a Poisson count averaging 3 support tickets per hour, what is the probability of zero tickets in an hour?

level: juniorimportance: must knowfreq 64%
basics
~20 s

The Poisson mass function is P(X = k) = e^(-λ) λ^k / k!. With λ = 3 tickets per hour, P(X = 0) = e^(-3), about 0.0498, so roughly a 5 percent chance of a completely quiet hour.

open as a page

What does a one-way ANOVA test, and what are its null and alternative hypotheses?

level: juniorimportance: must knowfreq 76%
basics
~20 s

One-way ANOVA tests whether the population means of three or more groups defined by a single factor are all equal. The null says every group mean is the same; the alternative says at least one differs.

open as a page

Which assumptions must hold for a two-sample t-test comparing group means to be valid?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A two-sample t-test assumes observations are independent within and between groups, that the sampling distribution of each group mean is approximately normal, and, for the pooled version, that the two populations share a common variance.

open as a page

What is bootstrap resampling, and how does it produce a confidence interval for a correlation?

level: juniorimportance: must knowfreq 62%
basics
~20 s

The bootstrap treats your sample as a stand-in for the population: draw many new samples of the same size with replacement, recompute the statistic on each, and read the interval off the middle 95% of those values.

open as a page

How does a chi-square goodness-of-fit test decide whether 600 die rolls came from a fair die?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A chi-square goodness-of-fit test compares observed category counts with the counts a hypothesised distribution predicts. A fair die over 600 rolls predicts 100 per face; the statistic sums (observed minus expected) squared, divided by expected, across the six faces.

open as a page

What does the 95% in a 95% confidence interval actually refer to?

level: juniorimportance: must knowfreq 88%
basics
~20 s

The 95% describes the procedure, not one interval. If you repeated the sampling and rebuilt the interval many times, about 95% of those intervals would contain the fixed true parameter. Any single interval either covers it or does not.

open as a page

Why is treating three repeated blood-pressure readings per patient as three independent observations wrong?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Readings from one patient are correlated, so three readings carry far less information than three different patients. Counting them as independent inflates the sample size, shrinks the standard errors and p-values, and manufactures significance that is not there.

open as a page

In a regression output, what does the standard error of a coefficient tell you?

level: juniorimportance: must knowfreq 78%
basics
~20 s

The standard error of a regression coefficient measures how much that estimate would move around if you refit the model on other samples from the same process. A smaller standard error means a more precisely estimated effect.

open as a page

With day-of-week dummies and Monday as the reference level, what does the Saturday coefficient mean?

level: juniorimportance: must knowfreq 80%
basics
~20 s

The Saturday coefficient is the estimated difference in the outcome between Saturday and Monday, with the model's other predictors held fixed. It is a contrast against the omitted baseline day, not Saturday's own average level.

open as a page

In a subscription churn analysis, how should you record a user who signed up 21 days ago and is still subscribed?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Record the user as right-censored at 21 days: the subscription lasted at least 21 days, and the eventual cancellation time is unknown. Deleting the row biases survival downward; coding it as a cancellation invents an event that never happened.

open as a page

What does the R-squared of a fitted multiple regression tell you about the model?

level: juniorimportance: must knowfreq 78%
basics
~20 s

R-squared is the share of the outcome's total variation that the fitted model accounts for, computed as 1 minus the residual sum of squares over the total sum of squares. It measures fit on the data used, not correctness.

open as a page

In a Bayesian A/B test, what does 'probability B beats A is 96%' actually mean?

level: juniorimportance: must knowfreq 76%
basics
~20 s

It is the posterior probability that variant B's true rate is higher than variant A's, given the data and the prior. It is a claim about the unknown rates, and says nothing about how large the difference is.

open as a page

What is the difference between probability as long-run frequency and probability as degree of belief?

level: juniorimportance: must knowfreq 78%
basics
~20 s

The frequency reading defines probability as the proportion of times an outcome occurs in repeatable trials. The belief reading defines it as a numeric degree of confidence, so it can also score one-off events such as a single rocket launch.

open as a page

Starting from a Beta(1,1) prior, what posterior follows from 8 clicks in 100 impressions?

level: juniorimportance: must knowfreq 58%
basics
~10 s

Beta(1 + 8, 1 + 92), that is Beta(9, 93). The prior contributes one pseudo-success and one pseudo-failure, so the posterior mean is 9/102, about 0.088, slightly above the raw rate of 0.08.

open as a page

What does a 95% Bayesian credible interval of [1.9%, 2.5%] say about a conversion rate?

level: juniorimportance: must knowfreq 82%
basics
~20 s

It says 95% of the posterior probability for the conversion rate falls between 1.9% and 2.5%. Given the model and prior you used, there is a 95% probability the rate lies in that range. The probability attaches to the parameter.

open as a page

What is the difference between probability and likelihood after seeing 7 heads in 10 coin flips?

level: juniorimportance: must knowfreq 75%
basics
~20 s

Probability fixes the coin's bias and varies the data; likelihood fixes the observed data and varies the bias. After 7 heads in 10 flips the likelihood is a curve over p that is not a probability distribution.

open as a page

In an A/B test, what is the difference between a count, a rate, and a share metric?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A count metric sums events, such as total orders. A rate metric divides a count by the units that were exposed, such as orders per visitor. A share metric divides one count by a related total, such as the percentage of orders placed on mobile.

open as a page

What is a guardrail metric in an A/B test, and how does it differ from a driver metric?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A guardrail is a metric the experiment must not damage — latency, crash-free sessions, unsubscribes — and a big enough degradation can veto the launch. A driver metric only explains why the primary metric moved; it never vetoes.

open as a page

Why does revenue per user make an A/B test hard to call when a few users spend far more than the rest?

level: juniorimportance: must knowfreq 74%
basics
~10 s

Almost all the variance in revenue per user comes from a handful of very large spenders, so the interval around the treatment-control difference stays wide. Ordinary experiment sizes then cannot resolve realistic revenue effects.

open as a page

What is an Overall Evaluation Criterion (OEC) in an A/B test?

level: juniorimportance: must knowfreq 74%
basics
~20 s

The Overall Evaluation Criterion is the single metric, agreed before launch, that an experiment's ship-or-not decision is defined against. It expresses what the team means by the product getting better, so the result cannot be reinterpreted afterwards.

open as a page

What does a DAU/MAU stickiness ratio of 0.2 tell you about how a product is used?

level: juniorimportance: must knowfreq 68%
basics
~10 s

A DAU/MAU of 0.2 means the typical monthly active user opens the product on about 6 days out of 30. It measures return frequency, not audience size, and averages over very different user types.

open as a page

What does a difference-in-differences estimate compute from a two-group, two-period panel?

level: juniorimportance: must knowfreq 70%
basics
~20 s

It subtracts the comparison group's before-to-after change from the treated group's before-to-after change. That double difference cancels the fixed level gap between the groups and any shock that moved both of them over the same window.

open as a page

Why can a treatment with an average effect of zero still help some users and harm others?

level: juniorimportance: must knowfreq 62%
basics
~20 s

An average pools opposite effects. A drug that raises recovery for patients under 50 and lowers it by a similar amount for patients over 70 posts an overall average near zero while both real effects remain untouched underneath it.

open as a page

What is a propensity score, and why match on it rather than on the raw covariates?

level: juniorimportance: must knowfreq 76%
basics
~20 s

A propensity score is a unit's probability of receiving the treatment given its observed covariates. Conditioning on that single number balances the covariates that went into it, so matching happens in one dimension instead of many.

open as a page

What does adding a confounder to a regression do to the treatment coefficient?

level: juniorimportance: must knowfreq 76%
basics
~20 s

Adding a measured confounder turns the treatment coefficient from a raw comparison into a within-strata one: it now compares treated and untreated units that share the same covariate value, removing the part of the gap that covariate explained.

open as a page

What does a sensitivity analysis for unmeasured confounding tell you about a causal estimate?

level: juniorimportance: must knowfreq 55%
basics
~10 s

A sensitivity analysis says how strong an unmeasured confounder would have to be to overturn the estimate. It never shows confounding is absent; it prices how much hidden bias the finding can tolerate.

open as a page

Why does MAPE break down on intermittent demand series with many zero-sales days?

level: juniorimportance: must knowfreq 76%
basics
~20 s

MAPE divides each absolute error by the actual value, so a zero-sales day makes that term undefined and a near-zero day makes it explode. On intermittent demand the average is dominated by a handful of tiny denominators.

open as a page

In a time series, what is the difference between a point outlier and a level shift?

level: juniorimportance: must knowfreq 70%
basics
~20 s

A point outlier is one observation far from expectation, after which the series returns to its old level. A level shift moves the series to a new baseline that persists. The difference decides whether you drop a point or re-baseline.

open as a page

Why is a shuffled random train/test split invalid for evaluating a daily demand forecast?

level: juniorimportance: must knowfreq 78%
basics
~20 s

A shuffled split trains on future days and tests on past ones, so the model interpolates between neighbouring dates instead of forecasting. Because adjacent days are highly correlated, the score looks excellent and says nothing about future performance.

open as a page

In a classical time-series decomposition, what do the trend, seasonal and remainder components each represent?

level: juniorimportance: must knowfreq 76%
basics
~20 s

Decomposition splits a series into three parts: trend-cycle, the slow movement of the level; seasonal, the pattern that repeats at a fixed known period such as 12 months; and remainder, the variation left after removing both.

open as a page

In simple exponential smoothing, what does the smoothing parameter alpha control?

level: juniorimportance: must knowfreq 76%
basics
~20 s

Alpha sets how much weight the update puts on the newest observation versus the accumulated past. Alpha near 1 tracks recent data and reacts fast; alpha near 0 averages over a long history and smooths noise.

open as a page