What does a standardised regression coefficient (beta) report that a raw coefficient does not?
answer
- one yardstick for every predictor
- divide out both standard deviations
- standard deviations in, standard deviations out
- spread is a property of this sample
- significance does not move at all
basics
~20 sA standardised beta reports the change in the outcome, measured in outcome standard deviations, per one standard deviation increase in the predictor, with the other predictors held fixed. It removes the units so magnitudes can be compared.
solid answer
~50 sA standardised beta is `beta_j = b_j * s_xj / s_y`, where `s_xj` is the predictor's standard deviation and `s_y` the outcome's. It answers "how many outcome standard deviations does one predictor standard deviation buy?", so a model containing age in years, tenure in months and income in dollars produces numbers on one dimensionless scale instead of three incomparable ones. Equivalently you can standardise every variable before fitting and read the coefficients directly. What it does not change is the fit: standardising is a linear rescaling, so t-statistics, p-values and R-squared are identical. What it does not license is a stable importance ranking - the standard deviations are properties of this sample, a predictor observed over a narrow range gets a small beta regardless of the underlying relationship, and correlated predictors still split their shared association.
go deeper
Know what the number says in words: one standard deviation more of the predictor goes with this many standard deviations of the outcome, other predictors unchanged.
Be able to write the formula, explain that it is just a rescaling so nothing about significance moves, and name the sample-dependence of the standard deviations.
Show you choose the summary for the audience - betas for cross-predictor comparison, a natural step in original units for stakeholders - and that you flag correlated predictors before anyone ranks them.
Decide what your organisation reports as effect size and hold the line, so results stay comparable across models and no one upgrades a comparison into a causal claim by changing the scale.
## The definition A raw OLS slope `b_j` is in outcome units per predictor unit, which makes coefficients on differently-measured predictors incomparable in size. The standardised coefficient removes both sets of units: ``` beta_j = b_j * (s_xj / s_y) ``` where `s_xj` is the sample standard deviation of the predictor and `s_y` is the sample standard deviation of the outcome. The result is dimensionless and reads as: **a one-standard-deviation increase in this predictor is associated with `beta_j` standard deviations of change in the outcome, holding the other predictors fixed.** The equivalent procedure is to standardise every variable first - subtract each mean and divide by each standard deviation - and fit the model on the transformed variables. The coefficients that come out are the standardised betas, and the intercept is zero because all the centred variables have mean zero. A half-standardised version is sometimes more useful: multiply by `s_xj` only, leaving the outcome in its natural units, so the coefficient reads as "outcome units per standard deviation of the predictor". Stakeholders often find that easier than a fully dimensionless number, because the outcome stays in dollars or minutes. ## What problem it solves Put age in years, tenure in months and income in dollars into one model and the raw coefficients differ in magnitude mostly because the units differ in magnitude. Standardising asks the same question of every predictor - what does a typical-sized move in this variable buy? - so the numbers land on one scale and can at least be looked at side by side. That is a genuine gain in readability. It is the standard answer to "which of these matters more?" in fields where the predictors have no natural common unit. ## What it does not change Standardising is a linear rescaling of each variable, so it changes nothing about the fit. Fitted values, residuals and R-squared are identical. Every t-statistic and p-value is identical, because the coefficient and its standard error scale by exactly the same factor. If someone reports that standardising strengthened a result, something else changed - the model, the sample or the arithmetic. The sign and the ordering by magnitude within a single fitted model are also unchanged relative to the raw coefficients multiplied by their standard deviations; standardising just makes that comparison explicit rather than accidental. ## The limits worth knowing **Sample-dependence.** `s_xj` describes the spread of the predictor in *this* dataset. Restrict the sample to a narrow slice of a predictor and its beta shrinks, even though nothing about the underlying relationship changed. This makes betas awkward to compare across studies or across time periods with different sampling: the raw coefficient is the more portable quantity, and the beta is the more comparable-within-one-model quantity. **Binary predictors.** The standard deviation of a 0/1 indicator is `sqrt(p*(1-p))`, so it depends entirely on how common the category is, and it is at most 0.5. "A one-standard-deviation increase" in a variable that only takes two values is not a change anything can undergo. For indicators, report the effect in the model's own units instead of standardising them. **Correlated predictors.** Standardising does nothing about overlap. When two predictors carry much of the same information, the fit divides their shared association between them in a way that depends on the specification, so neither beta is a stable property of that variable alone. **Still not causal.** A standardised beta is a description of a comparison inside a fitted model. It carries no more claim about what happens if you intervene than the raw coefficient did. **Nonlinear terms.** The clean "one standard deviation in, beta standard deviations out" reading assumes the predictor enters the model in a single linear term. If it appears in more than one term, no single standardised number summarises moving it. ## Choosing between beta and a natural step In practice, decide by audience. For a technical comparison across predictors with no common unit, the standardised beta is the right summary and should be reported with a confidence interval like any other estimate. For a business audience, the effect of a concrete, decision-relevant change in the original units is almost always more useful, because a standard deviation is itself an abstraction most readers cannot picture. Reporting both is cheap and forestalls the argument. ## The interview answer Give the formula, give the reading in words, then volunteer the limitation before you are asked: it is comparable within this model and this sample, it does not change any p-value, and it is not a licence to declare one predictor causally more important than another. Volunteering the caveat is what separates someone who has used standardised betas from someone who has read about them.
- How is a standardised beta computed from a raw OLS slope?Multiply the raw slope by the predictor's sample standard deviation and divide by the outcome's: `beta = b * s_x / s_y`. Equivalently, standardise every variable before fitting and read the coefficients off directly. The intercept then comes out at zero, because centred variables all have mean zero.
- Why can two samples with the same underlying relationship give different standardised betas?Because the beta depends on the sample standard deviations. Restricting the sample to a narrow range of a predictor shrinks its standard deviation and therefore its beta, without the underlying relationship changing. Raw coefficients are the more portable quantity across samples; betas are for comparing predictors within one fit.
- Why are standardised betas awkward for binary indicator predictors?A 0/1 variable's standard deviation is `sqrt(p*(1-p))`, driven entirely by how common the category is and never exceeding 0.5. A one-standard-deviation move in a two-valued variable is not a change anything can undergo, so the number is hard to speak. Report those effects in the model's own units instead.
- Does standardising the variables change the model's t-statistics or R-squared?No. It is a linear rescaling, so fitted values and residuals are unchanged, which fixes R-squared. Each coefficient and its standard error scale by the same factor, so every t-statistic and p-value is identical. Standardising is a presentation choice with no inferential consequence whatsoever.
saying these in an interview costs you the question
- Believes standardising can change a p-value or R-squared
- Treats standardised betas as a causal importance ranking
- Forgets the betas depend on this sample's standard deviations
- Standardises binary indicators and quotes a one-SD change
- Confuses a standardised beta with a correlation coefficient