skip to content

Model Diagnostics

Checking whether a fit deserves trust: residual patterns, non-constant variance, single points that drag the line, predictors that duplicate each other. Interviewers show a plot and ask what broke.

on this pageshow

explore

questions

26

Why is treating three repeated blood-pressure readings per patient as three independent observations wrong?

level: juniorimportance: must knowfreq 70%

answer

  1. same patient, not new information
  2. n counts units, not rows
  3. errors correlated within patient
  4. standard errors too small, p-values too small
  5. pseudo-replication

basics

~20 s

Readings from one patient are correlated, so three readings carry far less information than three different patients. Counting them as independent inflates the sample size, shrinks the standard errors and p-values, and manufactures significance that is not there.

solid answer

~50 s

Repeated measurements on one patient are positively correlated: knowing the first reading tells you a lot about the second, so three readings sit closer to one independent data point than to three. Ordinary regression assumes independent errors, so 300 readings from 100 patients make it behave as if it had 300 independent units and divide the variance by far too large an `n`. The coefficients themselves usually stay unbiased when the mean model is right, but the standard errors are too small, the confidence intervals too narrow and the p-values too small, so the real Type I error rate runs well above the nominal 5%. That is pseudo-replication: replicating the measurement is not replicating the unit. The honest fixes are to make the patient the unit — aggregate per patient, cluster the standard errors by patient, or fit a random intercept per patient.

go deeper

for a junior

Be ready to say plainly that rows are not units: several readings from one patient do not count as several patients. Naming the within-patient correlation and its effect on the p-value is enough at this level.

for a middle

Expect to explain the mechanics: the errors share a patient-specific offset, the variance formula divides by far too large an n, so the standard error is too small while the coefficient itself stays unbiased.

for a senior

Show that you would catch this in the data, not in the write-up. Spot the repeating id column before modelling, decide the unit deliberately, and justify your choice among aggregation, clustered standard errors and a random intercept.

for a principal

Own the standard: the analysis unit is a design decision, not a post-hoc repair, and it belongs in every review checklist your team uses. Be ready to explain to a stakeholder why a result with a tiny p-value had to be withdrawn once repeated readings were counted honestly.

## Rows are not units Every standard-error formula in ordinary least squares rests on one assumption that has nothing to do with normality or straight lines: the errors are independent of each other. Independence is what licenses the arithmetic `SE = s / sqrt(n)` and, more generally, the whole `(X'X)^-1 * sigma^2` variance formula. When a dataset has several rows per patient, per user, per store or per classroom, that assumption is usually false, and the sample size the software believes in is not the sample size you actually have. Call the readings on patient `j` at visits 1, 2 and 3 `y_j1`, `y_j2`, `y_j3`. Patients differ from each other in baseline blood pressure for reasons the model does not contain — age, weight, medication history, the anxiety of being in a clinic. Write that persistent patient-specific offset as `u_j` and the visit-to-visit noise as `e_jt`, so the error on any row is `u_j + e_jt`. Two rows from the same patient share `u_j`, so their errors are positively correlated; two rows from different patients share nothing and stay independent. The strength of the sharing is the intraclass correlation, `ICC = var(u) / (var(u) + var(e))`. ## What actually breaks Three things are worth separating, because interviewers probe exactly this distinction. **The coefficients are usually fine.** Under within-patient correlation, ordinary least squares remains unbiased for the coefficients as long as the mean model is correctly specified and the clustering is not itself confounded with the predictor. It is no longer the minimum-variance estimator — it is inefficient — but it is not systematically wrong. **The standard errors are not fine.** The reported variance assumes `n` independent errors. With positive within-patient correlation there is less independent information than `n`, so the reported standard error is too small. Confidence intervals are too narrow, t-statistics too large, p-values too small. A test you believe rejects 5% of the time under the null can easily reject 20% or 40% of the time. This is a false-positive machine, and it is silent: nothing in the output looks wrong. **The degrees of freedom are not fine either.** Software counts residual degrees of freedom from rows. With 300 rows from 100 patients, the reference distribution is far too generous. The severity is not uniform across predictors. A predictor that is constant within a patient — a patient-level treatment assignment, sex, a chronic diagnosis — suffers the worst inflation, because each patient really contributes just one observation of that predictor while the software counts three. A predictor that varies freely from visit to visit and is uncorrelated within the patient suffers much less. The damage grows with both the correlation of the errors within a cluster and the correlation of the predictor within a cluster. ## Recognising it before you model The diagnostic is structural, not graphical: look at the identifier columns. If a patient id, user id, device id or session id repeats down the rows, you have clustered data, and the burden of proof is on the analyst to show it does not matter. Counting distinct ids against total rows takes one line and tells you the worst case immediately: 300 rows and 100 patients means your true sample size is somewhere between 100 and 300, never above 300, and closer to 100 the more the patient dominates the variation. ## The three honest analyses **Aggregate to the unit.** Average each patient's readings into one number and analyse 100 rows. Simple, transparent, hard to argue with when cluster sizes are equal and you only care about patient-level predictors. It costs you any within-patient effect and some precision when cluster sizes are unequal. **Keep the rows, fix the inference.** Fit on all 300 rows but compute standard errors clustered by patient. This allows arbitrary correlation inside a patient while assuming independence across patients, and it changes only the uncertainty, never the estimates. **Model the structure.** Fit a random intercept per patient, which estimates `var(u)` explicitly, weights patients sensibly and recovers efficiency. It asks more of you: the random effect must be uncorrelated with the predictors, or the coefficients themselves become biased. ## The interview point The sentence that lands is that a p-value is only as trustworthy as the independence claim underneath it, and adding visits to the same patients buys measurement precision, not statistical power over patients. If you want more power over patients, recruit more patients.

  • Does pseudo-replication bias the coefficients themselves, or only the standard errors?
    Usually only the inference. With a correctly specified mean model, ordinary least squares stays unbiased under within-cluster correlation; it is merely inefficient. What breaks is the variance estimate, so the standard errors, intervals and p-values are all too small. Bias in the coefficients appears only when the clustering is confounded with the predictor, for example when patients with more visits also differ systematically in treatment.
  • For which kind of predictor is the understatement of the standard error worst?
    For a predictor that is constant within a cluster, such as a patient-level treatment or a demographic. Each patient then contributes essentially one observation of that predictor while the software counts every row, so the naive standard error can be several times too small. A predictor that moves from visit to visit and is uncorrelated within the patient suffers far less inflation.
  • Would averaging each patient's readings into a single value be an acceptable analysis?
    Often yes, and it is the simplest defensible answer when every patient has a similar number of readings and the question is about patient-level predictors. You lose the ability to estimate anything that varies within a patient, and you throw away precision when cluster sizes are very unequal, since a mean of ten readings deserves more weight than a mean of two.

Asking one person the same question three times is not a three-person survey. You get a more precise reading of that one opinion, not three opinions.

saying these in an interview costs you the question

  • More rows always mean more statistical power
  • It is fine because every reading is a genuine measurement
  • Independence only matters for the normality assumption
  • Add more visits per patient to gain power over patients
  • The p-value came out tiny, so the effect must be real

context

open as a page

What does heteroscedasticity mean for the error terms in a linear regression?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Heteroscedasticity means the variance of the regression errors is not the same for every observation: it changes systematically, usually with a predictor. Homoscedasticity, the classical assumption, is the opposite - one common error variance for every row.

open as a page

In a regression fit, what is the difference between an outlier and an influential observation?

level: juniorimportance: must knowfreq 58%

basics

~10 s

An outlier has a large residual: its response sits far from the fitted line. An influential observation is one whose removal visibly changes the fitted coefficients. A point can be either, both, or neither.

open as a page

What is multicollinearity in a linear regression, and what does it damage?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Multicollinearity means the predictors are strongly linearly related to each other, so the fit cannot tell their separate effects apart. It inflates coefficient standard errors and makes individual coefficients unstable, while overall fit and predictions stay largely intact.

open as a page

What should a residual-vs-fitted plot from a linear regression look like when the model's assumptions hold?

level: juniorimportance: must knowfreq 74%

basics

~10 s

A structureless horizontal band: residuals scattered randomly around zero across the whole fitted range, with roughly constant vertical spread and no curve, funnel or clustering. Any visible shape means the model is missing something.

open as a page

What do cluster-robust standard errors change in an OLS fit with many observations per user?

level: middleimportance: must knowfreq 58%

basics

~20 s

Cluster-robust standard errors leave the coefficients unchanged and only rewiden the uncertainty around them. They permit arbitrary correlation inside each user while assuming independence across users, so reported precision reflects the number of users rather than the number of rows.

open as a page

Does heteroscedasticity bias OLS coefficient estimates, or only their standard errors?

level: middleimportance: must knowfreq 76%

basics

~20 s

Heteroscedasticity leaves OLS coefficient estimates unbiased and consistent; it breaks the conventional standard-error formula. Every t-statistic, p-value and confidence interval built from those standard errors is therefore untrustworthy, and OLS is no longer the minimum-variance linear estimator.

open as a page

What is the leverage (hat value) of an observation in linear regression, and what determines it?

level: middleimportance: must knowfreq 62%

basics

~20 s

Leverage (the hat value h_ii) measures how unusual a row's predictor values are compared with the rest of the data. It depends only on X, never on the response y, and bounds how far that row can pull its own fitted value.

open as a page

How is the variance inflation factor computed, and what does a VIF of 12 mean?

level: middleimportance: must knowfreq 68%

basics

~20 s

The variance inflation factor for a predictor is 1 divided by (1 minus its R squared from regressing it on the other predictors). A VIF of 12 inflates that coefficient's variance twelvefold and its standard error about 3.5 times.

open as a page

In a residual-vs-fitted plot from an OLS fit, what does a clean U-shaped curve indicate?

level: middleimportance: must knowfreq 64%

basics

~20 s

It indicates the straight-line mean function is wrong. The true relationship curves, so the fit over-predicts in the middle of the fitted range and under-predicts at both ends. The remedy is to change the model, not to delete points.

open as a page

With an intraclass correlation of 0.3 and 20 pupils per classroom, what is the effective sample size?

level: middleimportance: should knowfreq 45%

basics

~10 s

The design effect is 1 + (20 - 1) * 0.3 = 6.7, so divide the pupil count by 6.7. An 800-pupil study in 40 classrooms carries the information of about 119 independent pupils.

open as a page

How does the Breusch-Pagan test check for non-constant error variance in a regression?

level: middleimportance: should knowfreq 54%

basics

~20 s

Breusch-Pagan regresses the squared OLS residuals on the model's predictors. Its statistic, n times the auxiliary R-squared, follows a chi-square distribution with one degree of freedom per auxiliary predictor; a small p-value rejects the null of constant error variance.

open as a page

How does Cook's distance measure the influence of a single observation on a regression fit?

level: middleimportance: should knowfreq 50%

basics

~20 s

Cook's distance summarises how far all the fitted values move when one observation is dropped and the model is refitted. It combines that row's rescaled residual with its leverage, so a point needs both an odd response and an odd predictor position to score high.

open as a page

In a normal Q-Q plot of regression residuals, which pattern signals heavy tails rather than right skew?

level: middleimportance: should knowfreq 55%

basics

~20 s

Heavy tails bend both ends away from the straight line in opposite directions: the lowest points fall below it and the highest rise above it. Right skew bends the whole plot upward, with both ends above the line and the largest gap at the top.

open as a page

What does a rising smoother in a regression's scale-location plot indicate about the errors?

level: middleimportance: should knowfreq 47%

basics

~20 s

It indicates non-constant error variance: the typical size of a residual grows with the fitted value. The same fact shows up in the residual-vs-fitted plot as a band that widens into a funnel toward the right.

open as a page

In a random-intercept model of test scores nested in schools, what does the intercept variance tell you?

level: seniorimportance: should knowfreq 42%

basics

~20 s

It estimates how widely school mean scores spread around the overall mean, as one variance on the outcome scale. Divided by the total variance it gives the intraclass correlation, and it treats the sampled schools as draws from a wider population.

open as a page

For heteroscedastic data, when do you prefer weighted least squares over robust standard errors?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Prefer weighted least squares when the variance model is genuinely known - most cleanly with aggregated rows whose group sizes you have. Otherwise keep ordinary least squares and report heteroscedasticity-robust standard errors, which require no variance model at all.

open as a page

A row with huge Cook's distance dominates your regression - how do you decide whether to drop it?

level: seniorimportance: should knowfreq 46%

basics

~20 s

First establish what the row actually is. Delete it only if it is a data error or falls outside the population you are modelling. If it is genuine, keep it and report the fit both with and without it.

open as a page

A regression's overall F-test is highly significant but no coefficient's t-statistic clears 2 — why?

level: seniorimportance: should knowfreq 52%

basics

~20 s

That pattern is the classic signature of severe multicollinearity. The predictors jointly explain the outcome well, which the F-test detects, but they overlap so heavily that no single coefficient can be pinned down, so every individual t-statistic stays small.

open as a page

With 5,000 rows, a regression's residuals are clearly non-normal in the Q-Q plot. Are the coefficient p-values still usable?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Usually yes. With thousands of observations the sampling distribution of a least-squares coefficient is close to normal whatever the error distribution looks like, so its p-values and confidence intervals are approximately right. Prediction intervals for individual observations get no such protection.

open as a page

Two near-duplicate engagement predictors flip between +40 and -38 across bootstrap refits — how do you fix the model?

level: principalimportance: should knowfreq 41%

basics

~20 s

First decide what the model is for. If it only forecasts, the swings are harmless. If someone will act on a coefficient, the data cannot separate the two effects, so drop one, combine them into a single measure, or report the well-estimated joint effect instead.

open as a page

How does White's test for heteroscedasticity differ from the Breusch-Pagan test?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

White's test uses the same squared-residual auxiliary regression as Breusch-Pagan but adds the squares and all cross-products of the predictors, so it catches variance patterns of any shape. The price is many more degrees of freedom and lower power.

open as a page

Why prefer studentised deleted residuals over standardised residuals when flagging regression outliers?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

A standardised residual divides by an error scale that the suspect observation itself inflates, which can mask it. The studentised deleted version re-estimates that scale with the observation left out, so a genuine outlier cannot hide behind its own effect on the fit.

open as a page

Why are regression residuals plotted against fitted values rather than against the observed outcome?

level: middleimportance: nice to knowfreq 24%

basics

~20 s

Because least-squares residuals are uncorrelated with the fitted values by construction, so any pattern there is a real signal. Residuals are correlated with the observed outcome, so that plot slopes upward even when the model is perfectly specified.

open as a page

Your standard errors are clustered by store but you only have 8 stores — what goes wrong?

level: seniorimportance: nice to knowfreq 26%

basics

~10 s

Cluster-robust standard errors are justified by the number of clusters growing, not the number of rows. With 8 stores they are downward-biased and noisy, so tests over-reject however many transactions each store contributes.

open as a page

When is high multicollinearity safe to ignore in a regression model?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Ignore it when the inflated coefficients are ones you never interpret: pure forecasting, collinearity confined to control variables, or a predictor of interest that is itself uncorrelated with the tangled ones. It matters only when a decision rests on a separated effect.

open as a page