skip to content

Test Families

Picking the right test for the data in front of you: t and z for means, chi-square for counts, ANOVA for several groups, ranks when normality fails. Interviewers ask you to justify the choice.

on this pageshow

explore

questions

page 1 of 2

What does a one-way ANOVA test, and what are its null and alternative hypotheses?

level: juniorimportance: must knowfreq 76%

answer

  1. compares more than two group means
  2. one categorical factor, several levels
  3. omnibus, not pairwise
  4. alternative is 'at least one'

basics

~20 s

One-way ANOVA tests whether the population means of three or more groups defined by a single factor are all equal. The null says every group mean is the same; the alternative says at least one differs.

solid answer

~40 s

One-way ANOVA compares the means of several groups formed by one categorical factor — for example mean crop yield under four different fertilizers. The null hypothesis is that all four population means are equal; the alternative is that **at least one** mean differs from the rest, which is deliberately vague about which one. That makes it an *omnibus* test: a small p-value tells you the factor matters somewhere, not which fertilizer beat which. It does this by comparing how far the group means spread apart against how much the observations scatter inside each group, and it assumes the observations are independent, roughly normal within each group, and have similar spread across groups. If you need to name the winning group you follow the significant result with a post-hoc pairwise procedure.

go deeper

for a junior

Be ready to say in one breath what ANOVA compares and to state both hypotheses correctly, especially the 'at least one differs' alternative. Know that one factor with four levels is still a one-way design.

for a middle

Explain why a test about means is built out of variances: between-group spread against within-group spread, comparable in size when the null holds. Also state the independence, normality and equal-variance assumptions.

for a senior

Show you know what a significant omnibus result does and does not license. Interviewers listen for whether you jump straight to naming the winning group, or correctly stop and reach for a post-hoc pairwise procedure.

for a principal

Own the framing decision: whether the question in front of the business is really an omnibus one at all. Often a small set of pre-planned contrasts against a control answers the real question with more power than an all-groups sweep.

## The setup You have one **categorical factor** with `k` levels — four fertilizers applied to plots of land, three onboarding flows, five machine settings — and one **continuous outcome** measured on each unit, such as crop yield in tonnes per hectare. You want to know whether the factor moves the outcome at all. One-way ANOVA ("analysis of variance", one-way because there is a single factor) is the standard test for exactly this shape of data. "One-way" refers to the number of factors, not the number of groups. Four fertilizers is still a one-way design, because fertilizer is one factor with four levels. ## The hypotheses, stated precisely - **Null (H0):** mu_1 = mu_2 = ... = mu_k. Every group's *population* mean is the same. Note these are population parameters, not the sample means you computed — the sample means will always differ a little by chance. - **Alternative (H1):** at least one mu_i differs from at least one other. Equivalently, "not all the mu_i are equal." The alternative is the part candidates most often state wrongly. It is **not** "all group means differ from each other." A single rogue group among five is enough to make the null false. The alternative is also not directional — there is no "one-tailed ANOVA" in the sense of a t-test, because with three or more means there is no single direction to point in. ## Why it is called an omnibus test An omnibus test answers one broad question with one p-value. It has no machinery for saying *which* means differ. If your four-fertilizer ANOVA returns p = 0.004, the honest sentence is "fertilizer affects yield," not "fertilizer C is the best." Ranking the sample means and declaring the top one the winner is a classic mistake: sampling noise reorders sample means all the time, and the omnibus test never evaluated that specific comparison. To name groups you run a post-hoc pairwise procedure that accounts for the fact that you are now making many comparisons at once. ## How the name makes sense The test is about means, but it is called *analysis of variance* because it works by splitting variability. The total scatter of every observation around the grand mean is decomposed into: - **Between-group** variability — how far each group mean sits from the grand mean, weighted by group size. This grows when the factor really does shift the outcome. - **Within-group** variability — how far individual observations sit from their own group mean. This is the background noise that exists regardless of the factor. If the null is true, the group means are only wandering around the grand mean by chance, so the between-group piece is just another view of the same background noise, and the two pieces come out comparable in size. If some group really is different, the between-group piece inflates. The test statistic is the ratio of the two, and its sampling distribution under the null is an F distribution. ## What it assumes One-way ANOVA assumes the observations are **independent** (one measurement per plot, no plot counted twice, no hidden clustering), the outcome is approximately **normally distributed within each group**, and the groups have roughly **equal variance**. The test is reasonably robust to mild departures from normality, especially with decent and balanced group sizes, and much less forgiving of dependence between observations — dependence quietly shrinks the effective sample size and inflates false positives. ## The two-group special case Nothing breaks if you run a one-way ANOVA on exactly two groups. It is simply equivalent to the pooled two-sample test of equal means: the F statistic comes out as the square of the t statistic, and the two p-values are identical. ANOVA earns its keep from three groups upward, where a single test replaces a pile of pairwise ones. ## What a good answer sounds like "One-way ANOVA asks whether one categorical factor shifts the mean of a continuous outcome. The null is that all group means are equal; the alternative is that at least one differs. It compares between-group spread to within-group spread, and a significant result tells you the factor matters without telling you which groups are responsible — that needs a post-hoc comparison."

  • A four-group ANOVA comes back significant — what can you say about which group is best?
    Nothing yet. The omnibus F only says at least one population mean differs from the others. The group with the highest sample mean may simply have drawn a lucky sample, and the test never evaluated that specific pair. Naming a winner requires a post-hoc pairwise procedure that adjusts for making all the comparisons at once.
  • Can you run a one-way ANOVA on exactly two groups, and what do you get?
    Yes, and it reduces to the pooled two-sample test of equal means. The F statistic equals the square of the corresponding t statistic and the p-values are identical, since F with 1 numerator degree of freedom is a squared t. ANOVA only becomes useful from three groups upward, where one test replaces many pairwise ones.
  • Why is the ANOVA F test evaluated only in the upper tail?
    Because only a large F is evidence against equal means. F is between-group variability over within-group variability; the alternative pushes the numerator up, so evidence lives in the right tail. A very small F means the group means are closer together than chance would usually produce — odd, sometimes a sign of a data problem, but not evidence that the means differ.

It is a smoke alarm for group differences: it tells you something is burning somewhere in the house, but not which room.

saying these in an interview costs you the question

  • States the alternative as 'all group means differ'
  • Claims a significant ANOVA identifies which group is best
  • Thinks one-way ANOVA requires exactly three groups
  • Says ANOVA tests whether the group variances are equal
  • Confuses sample means with the population means being tested

context

open as a page

Which assumptions must hold for a two-sample t-test comparing group means to be valid?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A two-sample t-test assumes observations are independent within and between groups, that the sampling distribution of each group mean is approximately normal, and, for the pooled version, that the two populations share a common variance.

open as a page

How does a chi-square goodness-of-fit test decide whether 600 die rolls came from a fair die?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A chi-square goodness-of-fit test compares observed category counts with the counts a hypothesised distribution predicts. A fair die over 600 rolls predicts 100 per face; the statistic sums (observed minus expected) squared, divided by expected, across the six faces.

open as a page

What does a two-sample Kolmogorov-Smirnov test compare between two samples?

level: juniorimportance: must knowfreq 58%

basics

~20 s

A two-sample Kolmogorov-Smirnov test compares the two samples' entire distributions. It builds a step-shaped empirical CDF for each sample and takes the largest vertical gap between them, so a difference in location, spread or shape can all trigger rejection.

open as a page

Why are rank-based tests barely affected by a single mistyped value of 10,000?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Rank-based tests replace each value with its position in sorted order, so a mistyped 10,000 becomes only the largest rank, not a huge number. Its influence is capped at one rank; a mean and a variance have no such cap.

open as a page

When should you use a paired t-test instead of a two-sample t-test?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Use a paired t-test when each value in one group is naturally matched to one in the other, such as the same patient measured before and after treatment. It tests the mean of the within-pair differences.

open as a page

How is the F statistic in a one-way ANOVA built, and what are its two degrees of freedom?

level: middleimportance: must knowfreq 68%

basics

~20 s

F is the between-group mean square divided by the within-group mean square. With k groups and N observations the numerator has k-1 degrees of freedom and the denominator N-k. F sits near 1 when all group means are equal.

open as a page

In a chi-square test of independence on a 3x4 contingency table, how do you get expected counts and degrees of freedom?

level: middleimportance: must knowfreq 72%

basics

~20 s

Each expected count is that cell's row total times its column total, divided by the grand total. Degrees of freedom are (rows minus one) times (columns minus one), so a 3x4 table has (3-1)(4-1) = 6.

open as a page

What does the Mann-Whitney U test actually compare between two independent samples?

level: middleimportance: must knowfreq 72%

basics

~20 s

Pool both samples, rank them, and U counts how often a value from one group beats a value from the other, testing whether P(X > Y) = 1/2. It compares medians only when the two distributions share a shape.

open as a page

When is a z-test valid for testing a mean instead of a t-test?

level: middleimportance: must knowfreq 68%

basics

~20 s

A z-test for a mean is valid only when the population standard deviation is known rather than estimated. If you plug in the sample standard deviation, the extra uncertainty makes the statistic follow a t distribution instead.

open as a page

Twenty patients are each measured ten times and analysed as 200 independent observations - what breaks?

level: seniorimportance: must knowfreq 46%

basics

~20 s

The independence assumption breaks. Measurements from one patient are correlated, so 200 rows carry far less information than 200 independent ones: standard errors come out too small, confidence intervals too narrow and p-values far too optimistic.

open as a page

Why run one ANOVA across five groups instead of all 10 pairwise t-tests?

level: middleimportance: should knowfreq 61%

basics

~10 s

Each pairwise test spends its own 5% false-positive budget, so 10 of them carry roughly a 40% chance of at least one spurious significant result. A single ANOVA asks one question at one alpha.

open as a page

Why is Levene's test usually preferred over Bartlett's test for checking equal variances?

level: middleimportance: should knowfreq 38%

basics

~20 s

Bartlett's test assumes normality and mistakes heavy tails or skew for unequal variance, so it rejects far too often on real data. Levene's test compares absolute deviations from each group's centre and stays reliable when the data are not normal.

open as a page

On a normal Q-Q plot of a raw sample, what do an S-shape and a single bent tail mean?

level: middleimportance: should knowfreq 48%

basics

~20 s

An S-shape means both tails are heavier than a normal's: the lowest points sit below the reference line and the highest above it. A curve where both ends bend the same way, one tail far more than the other, means skew.

open as a page

When are expected cell counts too small for a chi-square test, and what do you use instead?

level: middleimportance: should knowfreq 52%

basics

~20 s

The common rule wants every expected count at least 5, or more leniently no expected count below 1 with at most a fifth below 5. Below that, collapse categories, gather more data, or run Fisher's exact test.

open as a page

Why are standard Kolmogorov-Smirnov critical values wrong when the reference normal's mean and SD come from the same sample?

level: middleimportance: should knowfreq 34%

basics

~20 s

Fitting the reference curve to the same data pulls it toward the sample, shrinking the maximum gap. Standard tables assume a reference fixed in advance, so p-values come out too large and the test under-rejects.

open as a page

What does a significant Kruskal-Wallis test tell you about three or more groups?

level: middleimportance: should knowfreq 38%

basics

~20 s

Only that the groups are not all alike — at least one tends to produce larger values than another. It is an omnibus test on pooled ranks with k-1 degrees of freedom and never says which pair differs.

open as a page

How does the Wilcoxon signed-rank test use the differences within matched pairs?

level: middleimportance: should knowfreq 46%

basics

~20 s

It takes each pair's difference, drops zeros, ranks the absolute differences from smallest to largest, and sums the ranks belonging to positive differences. Large or small sums indicate the differences are not centred at zero.

open as a page

Why is Welch's t-test usually a safer default than Student's pooled t-test?

level: middleimportance: should knowfreq 58%

basics

~20 s

Student's two-sample t-test pools the two groups into one variance estimate, which is only valid when their variances are equal. Welch's test keeps them separate and adjusts the degrees of freedom, so it stays accurate under unequal variances.

open as a page

Your one-way ANOVA across four groups is significant — how do you find which groups differ?

level: seniorimportance: should knowfreq 49%

basics

~20 s

A significant omnibus F only says at least one mean differs. Follow it with a post-hoc procedure such as Tukey's HSD, which compares all pairs at once while holding the error rate across the whole family of comparisons.

open as a page

Why can Shapiro-Wilk reject normality at n = 5,000 yet pass a clearly skewed n = 8 sample?

level: seniorimportance: should knowfreq 54%

basics

~20 s

A normality test measures evidence against normality, not the size of the departure. At n = 5,000 it detects deviations far too small to matter; at n = 8 it has almost no power, so passing proves nothing about the population.

open as a page

Why does a 2x2 chi-square test give the same p-value as a two-proportion z-test?

level: seniorimportance: should knowfreq 36%

basics

~20 s

On a 2x2 table the Pearson chi-square statistic is algebraically identical to the square of the pooled-variance two-proportion z statistic. A chi-square variable on one degree of freedom is a squared standard normal, so the two-sided tail areas coincide exactly.

open as a page

How does a Kolmogorov-Smirnov test behave on integer star ratings with heavy ties?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A Kolmogorov-Smirnov test is unreliable on heavily tied data: its null distribution assumes continuous measurements with no exact repeats. With five rating values the empirical CDFs jump in blocks and the classical p-value comes out too conservative.

open as a page

What does a team lose by defaulting to rank-based tests whenever the data look skewed?

level: seniorimportance: should knowfreq 44%

basics

~20 s

They quietly change the question. A rank test asks how often one group is higher, not by how much, so it can be significant while the total the business banks is unchanged — and it costs about five percent efficiency.

open as a page

In a two-proportion z-test comparing 42/300 with 61/300, when is the standard error pooled?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Pool for the hypothesis test, not for the confidence interval. The test assumes the null that both rates are equal, so one combined proportion estimates their shared variance; an interval must allow them to differ.

open as a page

A KS test on 1,000,000 sessions returns p below 1e-16 for a shift nobody would notice - how do you decide whether to act?

level: principalimportance: should knowfreq 36%

basics

~20 s

At a million observations the KS rejection threshold shrinks toward zero, so any real difference becomes significant. Judge the magnitude instead: read the statistic D as an effect size and compare it against a materiality threshold agreed beforehand.

open as a page

What information does the sign test use from each matched pair?

level: middleimportance: nice to knowfreq 24%

basics

~10 s

Only the direction: which member of the pair was larger. Magnitudes are discarded, ties are dropped, and the count of positive pairs is compared with a binomial distribution with success probability one half.

open as a page

How does the t distribution with 4 degrees of freedom differ from the standard normal?

level: middleimportance: nice to knowfreq 33%

basics

~20 s

Both are symmetric bell curves centred at zero, but t with 4 degrees of freedom has heavier tails and variance 2 rather than 1. Its two-sided 95% cutoff is about 2.78 against the normal's 1.96.

open as a page

In a two-way drug-by-dose ANOVA, what does a significant interaction term mean?

level: seniorimportance: nice to knowfreq 36%

basics

~20 s

A significant interaction means the effect of one factor depends on the level of the other: the drug's effect changes across doses. The main effects then describe averages that may not hold at any particular drug-dose combination.

open as a page

You log-transform right-skewed income before a t-test - which hypothesis are you now testing?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

One about geometric means, not arithmetic ones. A t-test on logged values compares mean log income between groups, and exponentiating that difference gives a ratio of geometric means - not the difference in average income the business usually asked about.

open as a page

showing 1–30 of 32