skip to content

Descriptive Statistics & Sampling

Mean, median, variance, quantiles and correlation for summarising a dataset, plus the sampling designs and biases behind the numbers. Interviewers start here to see whether you read data honestly.

on this pageshow

explore

questions

85 · 7 sections

For a right-skewed column like household income, why does the mean exceed the median?

level: juniorimportance: must knowfreq 84%
basics
~20 s

The mean sums every value, so a long right tail of very high incomes pulls it upward. The median depends only on rank, so extreme values barely move it. Mean above median is the usual signature of right skew.

open as a page

What are the four levels of measurement — nominal, ordinal, interval and ratio?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Nominal values are labels with no order. Ordinal values are ranked but without known equal gaps. Interval values have equal gaps and an arbitrary zero. Ratio values have equal gaps and a true zero, so ratios are meaningful.

open as a page

What is the difference between variance and standard deviation?

level: juniorimportance: must knowfreq 85%
basics
~20 s

Variance is the average squared deviation from the mean, so it is measured in squared units. Standard deviation is the square root of the variance, which puts the number back into the data's original units and makes it readable.

open as a page

When computing a variance, when do you divide by n and when by n-1?

level: middleimportance: must knowfreq 68%
basics
~20 s

Divide by N when your numbers are the whole population you describe. Divide by n-1 when they are a sample standing in for a larger population's spread. The n-1 version is always the larger of the two.

open as a page

A dashboard reports overall conversion as the unweighted mean of three teams' rates — why is that wrong?

level: seniorimportance: must knowfreq 60%
basics
~10 s

Averaging rates weights every team equally regardless of size, so a tiny team counts as much as a huge one. The correct overall rate is total conversions over total users — a size-weighted mean.

open as a page

Why does one extreme salary shift a dataset's mean but leave its median almost unchanged?

level: juniorimportance: must knowfreq 82%
basics
~20 s

The mean adds every value in full, so one enormous salary drags the average up. The median depends only on which value sits in the middle position, so an extreme point changes the ordering but barely moves the centre.

open as a page

What is the five-number summary of a dataset, and what does it tell you that the mean hides?

level: juniorimportance: must knowfreq 78%
basics
~20 s

The five-number summary is the minimum, first quartile, median, third quartile and maximum. Together they show where the bulk of the data sits and how far each tail reaches, which a single average cannot express.

open as a page

What does the sign of sample skewness tell you about the shape of a distribution?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Positive sample skewness means the long tail stretches to the right, toward large values. Negative means the long tail stretches to the left. A value near zero indicates a roughly symmetric shape, like a normal bell.

open as a page

How does a boxplot's 1.5 x IQR rule decide that a data point is an outlier?

level: middleimportance: must knowfreq 66%
basics
~20 s

It takes the first and third quartiles Q1 and Q3, sets IQR = Q3 - Q1, and marks any point below Q1 - 1.5 x IQR or above Q3 + 1.5 x IQR. Those cutoffs are called the fences.

open as a page

Why do latency SLOs report p95 and p99 instead of the mean response time?

level: middleimportance: must knowfreq 70%
basics
~20 s

Mean latency averages away the slow requests that users actually notice. The 95th and 99th percentiles report the experience of the worst 5 percent and 1 percent of requests, which is where dissatisfaction, timeouts and retries live.

open as a page

In a cross-tab of device type by plan, what is the difference between joint, marginal and conditional proportions?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A joint proportion divides a cell count by the grand total. A marginal proportion divides a row or column total by the grand total. A conditional proportion divides a cell by its own row or column total.

open as a page

Ice-cream sales correlate with drowning deaths — what can and cannot be concluded from that?

level: juniorimportance: must knowfreq 85%
basics
~20 s

A correlation only says the two series move up and down together in the observed data. It cannot say ice cream causes drownings: something else, such as hot weather, can drive both, and correlation carries no direction while causation does.

open as a page

What is the difference between sample covariance and Pearson's correlation coefficient r?

level: juniorimportance: must knowfreq 86%
basics
~20 s

Sample covariance measures whether two variables move together, but it carries the product of their units, so its size means little alone. Pearson's r divides covariance by both standard deviations, giving a unitless number between -1 and 1.

open as a page

What is regression to the mean, and when should you expect to see it?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Regression to the mean is the tendency for an extreme measurement to be followed by a less extreme one on remeasurement. Expect it whenever two measurements are correlated but not perfectly, because chance helped produce the extreme.

open as a page

What does Spearman's rank correlation coefficient measure between two variables?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Spearman's rho is a correlation computed on the ranks of the data instead of the raw values. It measures how well two variables move together in a consistently increasing or decreasing pattern, not just along a straight line.

open as a page

What is the difference between simple random sampling and stratified sampling?

level: juniorimportance: must knowfreq 76%
basics
~20 s

Simple random sampling gives every unit in the frame an equal chance of selection. Stratified sampling first splits the population into non-overlapping groups, then samples inside each one, guaranteeing every group appears and usually giving a more precise estimate.

open as a page

What is the difference between MCAR, MAR and MNAR missing data?

level: juniorimportance: must knowfreq 74%
basics
~20 s

MCAR means missingness is unrelated to any variable. MAR means it depends only on variables you observed. MNAR means it depends on the missing value itself, as when the highest earners are the ones who leave salary blank.

open as a page

How does survivorship bias inflate the average return in a table of funds that still exist today?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Badly performing funds get closed or merged away, so a table of funds still open today lists mostly winners. Averaging it measures the survivors, not the return an investor could have expected when picking a fund years ago.

open as a page

Why does filling missing numeric values with the column mean understate the standard deviation?

level: middleimportance: must knowfreq 58%
basics
~20 s

Every filled value sits exactly at the mean, so it adds nothing to the sum of squared deviations while still adding to the row count. The numerator is unchanged and the denominator grows, so variance and standard deviation shrink.

open as a page

Why did the 1936 Literary Digest poll get the election wrong despite millions of returned ballots?

level: middleimportance: must knowfreq 58%
basics
~20 s

Two selection failures compounded. The mailing list was built from telephone directories, car registrations and subscriber rolls, which excluded poorer households; and only about a quarter of recipients mailed a ballot back, and returners differed from non-returners.

open as a page

Daily active users are down 8% versus last Tuesday — what do you check before calling it a real drop?

level: juniorimportance: must knowfreq 82%
basics
~20 s

Check whether 8% is inside the normal spread of Tuesday-over-Tuesday changes for this metric, then rule out calendar effects and a tracking break. Only a delta outside normal variation, measured on clean data, is worth investigating as a product problem.

open as a page

How do you tell a broken analytics pipeline from a genuine drop in signups?

level: middleimportance: must knowfreq 70%
basics
~20 s

Reconcile the metric against an independent source of truth, such as the records the product itself writes. A tracking break usually starts abruptly at a release, hits only one platform or app version, and leaves the downstream business records unchanged.

open as a page

How would you set a control band on weekly revenue to decide which weeks to investigate?

level: middleimportance: should knowfreq 45%
basics
~20 s

Pick a baseline of stable weeks, take a centre line from them, and estimate the spread from consecutive-week differences. Set limits at the centre plus and minus about three spreads; weeks outside get investigated. Exclude known-abnormal weeks from the baseline.

open as a page

You slice a 3% drop across 40 country-by-platform cells and one is down 30% — what now?

level: seniorimportance: should knowfreq 48%
basics
~20 s

Ask how much of the aggregate drop that cell can explain, and whether an extreme cell was expected anyway. Scanning 40 slices guarantees a few look alarming, and small cells swing hardest, so size the contribution first.

open as a page

Total revenue fell 6% while paying-customer count is flat — how does account concentration change your read?

level: seniorimportance: should knowfreq 40%
basics
~20 s

Revenue is usually concentrated in a few large accounts, so check whether one or two explain the whole 6%. Recompute the delta excluding the top accounts: if it vanishes, this is an account event, not a broad product change.

open as a page

How would you estimate the number of piano tuners working in Chicago from scratch?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Break the target number into a chain of estimable factors: city population, people per household, share of households owning a piano, tunings per piano per year, and tunings one tuner performs per year. Multiply through, divide, and land near 50.

open as a page

How do you keep units consistent when estimating a school district's annual lunch spend?

level: middleimportance: must knowfreq 51%
basics
~20 s

Write the unit beside every factor and check that they cancel to leave dollars per year: students times participation gives meals per day, times school days per year gives meals per year, times dollars per meal gives the answer.

open as a page

How would you size a national coffee chain's annual revenue both top-down and bottom-up?

level: middleimportance: should knowfreq 58%
basics
~20 s

Bottom-up builds from unit economics: stores times cups per store per day times price per cup times days per year. Top-down starts from the whole out-of-home coffee market and takes the chain's share. Run both and reconcile the gap.

open as a page

In a Fermi estimate, how do you find the assumption that dominates the answer's uncertainty?

level: seniorimportance: should knowfreq 43%
basics
~20 s

Rank factors by the width of their plausible range as a ratio, not by their size. In a product relative errors add, so a factor uncertain by 3x dominates one known to 10 percent. Halve and double it.

open as a page

When is a back-of-envelope estimate good enough to decide on, and when do you insist on measurement?

level: principalimportance: nice to knowfreq 27%
basics
~10 s

An estimate is good enough when the decision does not flip anywhere across its plausible range. Write down the threshold that would change the call first, then check whether the low-to-high band straddles it.

open as a page

ARPU fell 5% last month — how do you tell whether revenue or the active-user count drove it?

level: juniorimportance: must knowfreq 72%
basics
~20 s

ARPU is revenue over active users, so pull each part's percent change separately: the ARPU growth factor equals the revenue growth factor divided by the user growth factor. A drop while revenue still rises means the user base grew faster than revenue.

open as a page

Overall conversion fell while every acquisition channel's own conversion rate rose — how?

level: middleimportance: must knowfreq 78%
basics
~20 s

The overall rate is a traffic-weighted average of the channel rates. If traffic shifts toward a low-converting channel, that weighted average can fall even when every channel improves. This is a mix shift, not a conversion regression.

open as a page

What does direct standardisation of a mortality rate to a fixed age mix let you compare?

level: middleimportance: should knowfreq 40%
basics
~20 s

It applies both populations' age-specific rates to one shared reference age distribution, so the comparison no longer reflects the fact that one population is older. The output is a hypothetical rate built for comparison, never either population's real death count.

open as a page

Revenue fell 3%: how do you build a contribution waterfall attributing that delta to country segments?

level: seniorimportance: should knowfreq 52%
basics
~20 s

Partition the total into mutually exclusive, exhaustive segments, compute each segment's absolute delta, and reconcile the bars to the headline. Report gross gains and gross losses separately, not just the net, and give entrants and exits their own bars.

open as a page

When a mix shift explains a metric drop, do you report the raw or the mix-adjusted number to leadership?

level: principalimportance: should knowfreq 33%
basics
~20 s

Report the raw number as the headline — it is what the business actually experienced — with the mix-adjusted figure beside it as the explanation. Never substitute one for the other, and never choose the standard mix after seeing the results.

open as a page