In a regression output, what does the standard error of a coefficient tell you?
answer
- answers how precise, not how big
- imagine refitting on a fresh sample
- same units as the coefficient itself
- residual noise over predictor spread
- the denominator of the t-ratio
basics
~20 sThe standard error of a regression coefficient measures how much that estimate would move around if you refit the model on other samples from the same process. A smaller standard error means a more precisely estimated effect.
solid answer
~50 sA fitted coefficient is an estimate, not the truth: refit the same model on a fresh sample and the number changes. The standard error is the estimated typical size of that sample-to-sample movement, reported in the same units as the coefficient itself. For a simple regression of `y` on one predictor `x`, it is `SE(b) = s / sqrt(sum of (x_i - xbar)^2)`, where `s` is the residual standard error — the typical vertical gap between the points and the fitted line. So three things drive it: more residual noise pushes it up, more observations pull it down, and more spread in the predictor pulls it down. It is also the ruler everything else is measured against: the t-statistic is the estimate divided by it, and the confidence interval is the estimate plus or minus a critical value times it.
go deeper
Be ready to say in one sentence that it measures how much the estimated coefficient would vary across samples, and that it is reported in the same units as the coefficient.
Expect to write the simple-regression formula and read the three drivers off it: residual noise on top, predictor spread and sample size underneath.
Show that you check the standard error before quoting any coefficient, and that you know the printed value depends on the model's error assumptions holding.
Own the reporting norm: coefficients travel with their uncertainty in your team's outputs, so no decision memo ever carries a bare point estimate.
## Estimate versus true value A fitted regression prints a number for each predictor — a slope of `0.8`, say. That number is an **estimate** computed from one particular sample. The quantity it aims at is the unknown coefficient of the process that generated the data. If you could draw a fresh sample from the same process and refit, you would get a different number: `0.6`, `1.1`, `0.9`. The standard error is the estimated typical size of that movement. It answers *how precisely have we pinned this down*, never *how big is the effect*. ## The formula and what falls out of it For a simple regression of `y` on a single predictor `x`, the slope's standard error is `SE(b) = s / sqrt( sum over i of (x_i - xbar)^2 )` where `s` is the **residual standard error**, `s = sqrt(SSE / (n - 2))`, and `SSE` is the sum of squared residuals. Since the sum of squared deviations of `x` equals `(n - 1)` times the sample variance of `x`, the same thing can be written `SE(b) = s / ( s_x * sqrt(n - 1) )` Three drivers read straight off this: - **Residual noise `s`, in the numerator.** The more the outcome scatters vertically around the line, the less any one fitted line can be trusted, and the larger the standard error. - **Sample size `n`, in the denominator under a square root.** Four times the data halves the standard error — it does not quarter it. - **Spread of the predictor `s_x`, in the denominator.** A predictor that barely varies gives the line almost nothing to pivot against, so the slope is poorly determined. ## Units, and how to read one The standard error carries the same units as the coefficient: outcome units per one unit of the predictor. If a slope is `0.8` with a standard error of `0.4`, the estimate is only twice its own noise level — weak evidence. If the same `0.8` came with a standard error of `0.05`, it is sixteen times its noise level and the effect is sharply determined. The absolute size `0.8` is identical in both cases; only the precision differs. This is why interviewers ask for the standard error the moment a candidate quotes a coefficient. The standard error is also the building block for the two summaries printed next to it. The t-statistic is `estimate / SE`. The confidence interval is `estimate +/- (critical value) * SE`. Both are just the estimate expressed in units of its own standard error, so anything that inflates the standard error simultaneously flattens the t-statistic and widens the interval. ## What it is not - **It is not the residual standard error `s`.** `s` describes how far individual observations sit from the fitted line and stays roughly constant as you collect more data; `SE(b)` describes how far the *estimated coefficient* might sit from the truth and shrinks as data accumulates. They are different quantities with different behaviour, and conflating them is the most common junior error here. - **It is not a measure of effect size.** A tiny standard error next to a tiny coefficient means you have precisely established a small effect. - **It is not a guarantee.** The printed value is computed under the model's assumptions about the error term. When those assumptions fail, the printed standard error can be too small or too large; diagnosing and repairing that is separate work. ## Interview framing A clean answer names the quantity (sampling variability of the estimate), gives the units (same as the coefficient), and names the three drivers (residual noise, sample size, predictor spread). If you can add that the t-statistic and interval are both built from it, you have covered everything a screening question on this is looking for.
- How is the coefficient's standard error different from the residual standard error of the model?The residual standard error describes how far individual observations fall from the fitted line, and it does not shrink as you collect more data — it estimates the noise in the process. The coefficient's standard error describes how far the estimated coefficient may sit from the true one, and it does shrink with more data. The residual standard error sits in the numerator of the coefficient's standard error, so the two are linked but answer different questions.
- Two models on the same data give the same slope, but one has a much smaller standard error. What differs?Either the second model explains more of the outcome's variation, shrinking the residual standard error in the numerator, or it was fit on more observations or on data where the predictor varies more widely, growing the denominator. The point estimate can be identical while the precision behind it is completely different — which is why a coefficient quoted without its standard error is not yet an answer.
It is the error bar on the coefficient: the coefficient says where the needle landed, the standard error says how much the needle jitters.
saying these in an interview costs you the question
- Says it measures the spread of the raw data
- Confuses it with the residual standard error of the model
- Treats a small standard error as proof of a large effect
- Thinks it describes the size of the coefficient rather than its precision
- Assumes it shrinks whenever you add more predictors