Regression & Statistical Modeling
Linear and logistic regression as statistical models: OLS assumptions, reading coefficients and odds ratios, and diagnosing a broken fit. Interviewers probe interpretation far more than fitting.
on this pageshowhide
explore
- Ordinary Least Squares15 questions
- Normal Equations and Gauss-Markov5 questions
- Coefficient Standard Errors5 questions
- Prediction Intervals5 questions
- Coefficient Interpretation15 questions
- Partial Effects and Units5 questions
- Dummy Variables5 questions
- Interactions and Log Terms5 questions
- Model Diagnostics26 questions
- Residual and Q-Q Plots6 questions
- Heteroscedasticity5 questions
- Leverage and Influence5 questions
- Multicollinearity and VIF5 questions
- Clustered and Repeated Data5 questions
- Goodness of Fit6 questions
- Generalized Linear Models20 questions
- Odds Ratios and Logit5 questions
- Poisson and Count Models5 questions
- Censoring and Kaplan-Meier5 questions
- Cox Proportional Hazards5 questions
questions
82 · 5 sectionsIn a regression output, what does the standard error of a coefficient tell you?
basics
~20 sThe standard error of a regression coefficient measures how much that estimate would move around if you refit the model on other samples from the same process. A smaller standard error means a more precisely estimated effect.
What quantity does ordinary least squares minimise when fitting a linear model?
basics
~20 sOrdinary least squares picks the coefficients that make the sum of squared residuals as small as possible. A residual is an observed value minus the value the fitted line predicts for it, so the misses are vertical.
Why is a regression prediction unreliable at a predictor value far outside the observed data range?
basics
~20 sNothing in the data supports the model's shape out there. The arithmetic still prints a number and a finite interval, but that interval covers only sampling error under an assumed straight line, not the risk that the true relationship bends.
How do you judge whether an OLS coefficient is significant from its estimate and standard error?
basics
~20 sDivide the coefficient by its standard error to get the t-statistic, then compare it against a t distribution with n minus the number of estimated coefficients. A magnitude near 2 or more clears the usual 5% bar.
How do you derive the OLS slope and intercept from the normal equations?
basics
~20 sSet both partial derivatives of the squared-residual sum to zero to get the two normal equations. Solving them gives slope = sum of (x - xbar)(y - ybar) over sum of (x - xbar) squared, and intercept = ybar minus slope times xbar.
With day-of-week dummies and Monday as the reference level, what does the Saturday coefficient mean?
basics
~20 sThe Saturday coefficient is the estimated difference in the outcome between Saturday and Monday, with the model's other predictors held fixed. It is a contrast against the omitted baseline day, not Saturday's own average level.
In a log(wage) regression, how do you interpret a coefficient of 0.07 on years of schooling?
basics
~20 sOne more year of schooling is associated with roughly a 7% higher wage, holding the other predictors fixed. Because the outcome is logged, the coefficient reads as an approximate percent change; the exact figure is exp(0.07) - 1, about 7.25%.
Why does adding all four region dummies plus an intercept break an OLS regression?
basics
~20 sThe four region dummies add to 1 in every row, reproducing the intercept's column of ones exactly. The design matrix loses full rank, so no unique least-squares coefficients exist. This is the dummy variable trap.
In a sales model with a promotion x weekend interaction, what does the promotion coefficient mean?
basics
~20 sIt is the promotion effect when the weekend indicator equals zero, that is on weekdays only. The interaction coefficient is the extra promotion effect on weekends, so the weekend promotion effect is the sum of the two.
In a log-log demand model, what does a price coefficient of -1.4 tell you?
basics
~20 sIt is a price elasticity: a 1% price increase is associated with roughly a 1.4% fall in quantity, holding the other predictors fixed. Because the size exceeds 1, demand is elastic, so raising price lowers revenue.
Why is treating three repeated blood-pressure readings per patient as three independent observations wrong?
basics
~20 sReadings from one patient are correlated, so three readings carry far less information than three different patients. Counting them as independent inflates the sample size, shrinks the standard errors and p-values, and manufactures significance that is not there.
What does heteroscedasticity mean for the error terms in a linear regression?
basics
~20 sHeteroscedasticity means the variance of the regression errors is not the same for every observation: it changes systematically, usually with a predictor. Homoscedasticity, the classical assumption, is the opposite - one common error variance for every row.
In a regression fit, what is the difference between an outlier and an influential observation?
basics
~10 sAn outlier has a large residual: its response sits far from the fitted line. An influential observation is one whose removal visibly changes the fitted coefficients. A point can be either, both, or neither.
What is multicollinearity in a linear regression, and what does it damage?
basics
~20 sMulticollinearity means the predictors are strongly linearly related to each other, so the fit cannot tell their separate effects apart. It inflates coefficient standard errors and makes individual coefficients unstable, while overall fit and predictions stay largely intact.
What should a residual-vs-fitted plot from a linear regression look like when the model's assumptions hold?
basics
~10 sA structureless horizontal band: residuals scattered randomly around zero across the whole fitted range, with roughly constant vertical spread and no curve, funnel or clustering. Any visible shape means the model is missing something.
What does the R-squared of a fitted multiple regression tell you about the model?
basics
~20 sR-squared is the share of the outcome's total variation that the fitted model accounts for, computed as 1 minus the residual sum of squares over the total sum of squares. It measures fit on the data used, not correctness.
Why does adding any predictor to an OLS regression never lower its R-squared?
basics
~20 sLeast squares can always set the new coefficient to zero and reproduce the previous fit, so the minimised residual sum of squares can only tie or shrink. Since R-squared is 1 minus that sum over a fixed total, it never falls.
What null hypothesis does the overall F-test in a linear regression output test?
basics
~20 sIt tests whether every slope coefficient is zero at once, meaning the model does no better than an intercept-only model that predicts the outcome's mean. A small p-value says at least one predictor carries signal, without saying which.
How do you test whether a block of three interaction terms improves a regression fit?
basics
~20 sFit the model with and without the three terms on identical rows and run a nested F-test: the drop in residual sum of squares per added parameter, divided by the full model's residual variance, on 3 and n-k-1 degrees of freedom.
Is an R-squared of 0.04 ever good enough to ship a regression model?
basics
~20 sYes, when the model's job is to identify and size drivers in a noisy outcome rather than to predict individuals. The acceptance bar comes from the decision the model supports, not from any fixed R-squared threshold.
In a subscription churn analysis, how should you record a user who signed up 21 days ago and is still subscribed?
basics
~20 sRecord the user as right-censored at 21 days: the subscription lasted at least 21 days, and the eventual cancellation time is unknown. Deleting the row biases survival downward; coding it as a cancellation invents an event that never happened.
In logistic regression, what is the difference between odds and probability?
basics
~10 sProbability is p, a number between 0 and 1. Odds is p/(1-p), the event weighed against its complement, and runs from 0 to infinity. A probability of 0.8 is odds of 4, or 4-to-1.
Why model event counts with Poisson regression instead of ordinary linear regression?
basics
~20 sCounts are non-negative integers, usually skewed, and their spread grows with their level. Poisson regression models the log of the expected count, so fitted values stay positive and each predictor acts multiplicatively rather than adding a fixed number of events.
How do you compute a Kaplan-Meier survival estimate by hand from a table of event and censoring times?
basics
~20 sAt each event time, divide the events by the number still at risk just before it and multiply the surviving fractions: S(t) = product of (1 - d_i / n_i). Censored subjects shrink the risk set without creating a step.
In a Cox proportional hazards model of churn, what does a hazard ratio of 1.8 on an annual-plan flag mean?
basics
~10 sAnnual-plan customers cancel at 1.8 times the instantaneous rate of the reference group, at every time point the model considers. A hazard ratio compares rates among those still subscribed; it is not a probability.