Descriptive Statistics & Sampling
Mean, median, variance, quantiles and correlation for summarising a dataset, plus the sampling designs and biases behind the numbers. Interviewers start here to see whether you read data honestly.
on this pageshowhide
explore
- Center, Spread and Scale14 questions
- Mean, Median, Mode5 questions
- Variance and Standard Deviation5 questions
- Levels of Measurement4 questions
- Distribution Shape & Outliers15 questions
- Quantiles and Histograms5 questions
- Skewness and Kurtosis5 questions
- Outliers and Robust Statistics5 questions
- Association & Correlation25 questions
- Covariance and Pearson r5 questions
- Spearman and Kendall5 questions
- Correlation Traps5 questions
- Cramer's V and Point-Biserial5 questions
- Regression to the Mean5 questions
- Sampling Design & Bias15 questions
- Random and Stratified Sampling5 questions
- Selection and Survivorship Bias5 questions
- Missing Data Mechanisms5 questions
- Diagnosing a Metric Change6 questions
- Back-of-Envelope Estimation5 questions
- Metric Decomposition & Mix Shift5 questions
questions
85 · 7 sectionsFor a right-skewed column like household income, why does the mean exceed the median?
basics
~20 sThe mean sums every value, so a long right tail of very high incomes pulls it upward. The median depends only on rank, so extreme values barely move it. Mean above median is the usual signature of right skew.
What are the four levels of measurement — nominal, ordinal, interval and ratio?
basics
~20 sNominal values are labels with no order. Ordinal values are ranked but without known equal gaps. Interval values have equal gaps and an arbitrary zero. Ratio values have equal gaps and a true zero, so ratios are meaningful.
What is the difference between variance and standard deviation?
basics
~20 sVariance is the average squared deviation from the mean, so it is measured in squared units. Standard deviation is the square root of the variance, which puts the number back into the data's original units and makes it readable.
When computing a variance, when do you divide by n and when by n-1?
basics
~20 sDivide by N when your numbers are the whole population you describe. Divide by n-1 when they are a sample standing in for a larger population's spread. The n-1 version is always the larger of the two.
A dashboard reports overall conversion as the unweighted mean of three teams' rates — why is that wrong?
basics
~10 sAveraging rates weights every team equally regardless of size, so a tiny team counts as much as a huge one. The correct overall rate is total conversions over total users — a size-weighted mean.
Why does one extreme salary shift a dataset's mean but leave its median almost unchanged?
basics
~20 sThe mean adds every value in full, so one enormous salary drags the average up. The median depends only on which value sits in the middle position, so an extreme point changes the ordering but barely moves the centre.
What is the five-number summary of a dataset, and what does it tell you that the mean hides?
basics
~20 sThe five-number summary is the minimum, first quartile, median, third quartile and maximum. Together they show where the bulk of the data sits and how far each tail reaches, which a single average cannot express.
What does the sign of sample skewness tell you about the shape of a distribution?
basics
~20 sPositive sample skewness means the long tail stretches to the right, toward large values. Negative means the long tail stretches to the left. A value near zero indicates a roughly symmetric shape, like a normal bell.
How does a boxplot's 1.5 x IQR rule decide that a data point is an outlier?
basics
~20 sIt takes the first and third quartiles Q1 and Q3, sets IQR = Q3 - Q1, and marks any point below Q1 - 1.5 x IQR or above Q3 + 1.5 x IQR. Those cutoffs are called the fences.
Why do latency SLOs report p95 and p99 instead of the mean response time?
basics
~20 sMean latency averages away the slow requests that users actually notice. The 95th and 99th percentiles report the experience of the worst 5 percent and 1 percent of requests, which is where dissatisfaction, timeouts and retries live.
In a cross-tab of device type by plan, what is the difference between joint, marginal and conditional proportions?
basics
~20 sA joint proportion divides a cell count by the grand total. A marginal proportion divides a row or column total by the grand total. A conditional proportion divides a cell by its own row or column total.
Ice-cream sales correlate with drowning deaths — what can and cannot be concluded from that?
basics
~20 sA correlation only says the two series move up and down together in the observed data. It cannot say ice cream causes drownings: something else, such as hot weather, can drive both, and correlation carries no direction while causation does.
What is the difference between sample covariance and Pearson's correlation coefficient r?
basics
~20 sSample covariance measures whether two variables move together, but it carries the product of their units, so its size means little alone. Pearson's r divides covariance by both standard deviations, giving a unitless number between -1 and 1.
What is regression to the mean, and when should you expect to see it?
basics
~20 sRegression to the mean is the tendency for an extreme measurement to be followed by a less extreme one on remeasurement. Expect it whenever two measurements are correlated but not perfectly, because chance helped produce the extreme.
What does Spearman's rank correlation coefficient measure between two variables?
basics
~20 sSpearman's rho is a correlation computed on the ranks of the data instead of the raw values. It measures how well two variables move together in a consistently increasing or decreasing pattern, not just along a straight line.
What is the difference between simple random sampling and stratified sampling?
basics
~20 sSimple random sampling gives every unit in the frame an equal chance of selection. Stratified sampling first splits the population into non-overlapping groups, then samples inside each one, guaranteeing every group appears and usually giving a more precise estimate.
What is the difference between MCAR, MAR and MNAR missing data?
basics
~20 sMCAR means missingness is unrelated to any variable. MAR means it depends only on variables you observed. MNAR means it depends on the missing value itself, as when the highest earners are the ones who leave salary blank.
How does survivorship bias inflate the average return in a table of funds that still exist today?
basics
~20 sBadly performing funds get closed or merged away, so a table of funds still open today lists mostly winners. Averaging it measures the survivors, not the return an investor could have expected when picking a fund years ago.
Why does filling missing numeric values with the column mean understate the standard deviation?
basics
~20 sEvery filled value sits exactly at the mean, so it adds nothing to the sum of squared deviations while still adding to the row count. The numerator is unchanged and the denominator grows, so variance and standard deviation shrink.
Why did the 1936 Literary Digest poll get the election wrong despite millions of returned ballots?
basics
~20 sTwo selection failures compounded. The mailing list was built from telephone directories, car registrations and subscriber rolls, which excluded poorer households; and only about a quarter of recipients mailed a ballot back, and returners differed from non-returners.
Daily active users are down 8% versus last Tuesday — what do you check before calling it a real drop?
basics
~20 sCheck whether 8% is inside the normal spread of Tuesday-over-Tuesday changes for this metric, then rule out calendar effects and a tracking break. Only a delta outside normal variation, measured on clean data, is worth investigating as a product problem.
How do you tell a broken analytics pipeline from a genuine drop in signups?
basics
~20 sReconcile the metric against an independent source of truth, such as the records the product itself writes. A tracking break usually starts abruptly at a release, hits only one platform or app version, and leaves the downstream business records unchanged.
How would you set a control band on weekly revenue to decide which weeks to investigate?
basics
~20 sPick a baseline of stable weeks, take a centre line from them, and estimate the spread from consecutive-week differences. Set limits at the centre plus and minus about three spreads; weeks outside get investigated. Exclude known-abnormal weeks from the baseline.
You slice a 3% drop across 40 country-by-platform cells and one is down 30% — what now?
basics
~20 sAsk how much of the aggregate drop that cell can explain, and whether an extreme cell was expected anyway. Scanning 40 slices guarantees a few look alarming, and small cells swing hardest, so size the contribution first.
Total revenue fell 6% while paying-customer count is flat — how does account concentration change your read?
basics
~20 sRevenue is usually concentrated in a few large accounts, so check whether one or two explain the whole 6%. Recompute the delta excluding the top accounts: if it vanishes, this is an account event, not a broad product change.
How would you estimate the number of piano tuners working in Chicago from scratch?
basics
~20 sBreak the target number into a chain of estimable factors: city population, people per household, share of households owning a piano, tunings per piano per year, and tunings one tuner performs per year. Multiply through, divide, and land near 50.
How do you keep units consistent when estimating a school district's annual lunch spend?
basics
~20 sWrite the unit beside every factor and check that they cancel to leave dollars per year: students times participation gives meals per day, times school days per year gives meals per year, times dollars per meal gives the answer.
How would you size a national coffee chain's annual revenue both top-down and bottom-up?
basics
~20 sBottom-up builds from unit economics: stores times cups per store per day times price per cup times days per year. Top-down starts from the whole out-of-home coffee market and takes the chain's share. Run both and reconcile the gap.
In a Fermi estimate, how do you find the assumption that dominates the answer's uncertainty?
basics
~20 sRank factors by the width of their plausible range as a ratio, not by their size. In a product relative errors add, so a factor uncertain by 3x dominates one known to 10 percent. Halve and double it.
When is a back-of-envelope estimate good enough to decide on, and when do you insist on measurement?
basics
~10 sAn estimate is good enough when the decision does not flip anywhere across its plausible range. Write down the threshold that would change the call first, then check whether the low-to-high band straddles it.
ARPU fell 5% last month — how do you tell whether revenue or the active-user count drove it?
basics
~20 sARPU is revenue over active users, so pull each part's percent change separately: the ARPU growth factor equals the revenue growth factor divided by the user growth factor. A drop while revenue still rises means the user base grew faster than revenue.
Overall conversion fell while every acquisition channel's own conversion rate rose — how?
basics
~20 sThe overall rate is a traffic-weighted average of the channel rates. If traffic shifts toward a low-converting channel, that weighted average can fall even when every channel improves. This is a mix shift, not a conversion regression.
What does direct standardisation of a mortality rate to a fixed age mix let you compare?
basics
~20 sIt applies both populations' age-specific rates to one shared reference age distribution, so the comparison no longer reflects the fact that one population is older. The output is a hypothetical rate built for comparison, never either population's real death count.
Revenue fell 3%: how do you build a contribution waterfall attributing that delta to country segments?
basics
~20 sPartition the total into mutually exclusive, exhaustive segments, compute each segment's absolute delta, and reconcile the bars to the headline. Report gross gains and gross losses separately, not just the net, and give entrants and exits their own bars.
When a mix shift explains a metric drop, do you report the raw or the mix-adjusted number to leadership?
basics
~20 sReport the raw number as the headline — it is what the business actually experienced — with the mix-adjusted figure beside it as the explanation. Never substitute one for the other, and never choose the standard mix after seeing the results.