A reviewer ranks predictor importance by raw coefficient size in a table with mixed units - how do you respond?
answer
- units are a choice, not evidence
- a dollar and a year are different steps
- rescale before you rank anything
- per standard deviation or per realistic step
- overlapping intervals mean no ranking
basics
~20 sRaw coefficient size is not importance, because every coefficient is in the units of its own predictor. Put the predictors on a common footing first - per standard deviation, or over a realistic range - and report intervals alongside.
solid answer
~40 sPoint out that the ranking is measuring units, not effects. In a model with age in years, tenure in months and income in dollars, the income coefficient is tiny only because one dollar is a trivial change, and I could make it the largest number in the table by re-expressing income in millions without refitting anything. Two defensible alternatives: report the outcome change per one standard deviation of each predictor, or better for stakeholders, the change over a decision-relevant step such as a 10,000-dollar income difference or an extra year of tenure, each with a confidence interval. I would also say what neither fixes: correlated predictors split credit between themselves, the comparison is still descriptive, and a small p-value indicates precision rather than magnitude.
go deeper
Remember the rule and the reason: coefficients carry their predictor's units, so a bigger number can just mean a smaller unit, and sizes are not comparable across differently-scaled predictors.
Be able to show the arithmetic - re-express a predictor and watch its coefficient change rank - and to name at least one repair, such as reporting effects per standard deviation.
Demonstrate that you would re-present the table yourself, in decision-relevant steps with intervals, and that you can state plainly which orderings the data does not support.
Own the reporting standard so importance claims are consistent across teams, and be ready to explain to leadership why a clean ranked list of drivers is often not something the data can honestly provide.
## Why the ranking is meaningless A regression coefficient is expressed in **outcome units per predictor unit**. The size of a coefficient therefore depends on two things: how strongly the predictor relates to the outcome, and how big the unit happens to be. The second has nothing to do with the data. Take a model with age in years, tenure in months and income in dollars. The income coefficient will be minuscule - a single dollar is a negligible difference - while age gets a large-looking number because a year is a big step. Refit the model with income in millions of dollars and, without a single value in the dataset changing meaning, income becomes the biggest coefficient in the table. A quantity that can be reordered by a choice of units is not measuring importance. This is the cleanest way to end the argument in a review: offer to make any predictor the "most important" one on request. Once the reviewer sees that the ranking is under your control, the conversation moves to what ranking would actually be defensible. ## Fix one: a common statistical yardstick Multiply each coefficient by the standard deviation of its predictor. That converts every effect into "outcome change per one standard deviation of this predictor", which is at least the same kind of step for each. Dividing additionally by the outcome's standard deviation gives fully standardised coefficients, which are dimensionless and directly comparable in size. This is a real improvement and it is what most people mean by comparable coefficients. It is also sample-dependent: standard deviations describe the spread in *this* dataset, so a predictor with a narrow observed range gets a small standardised effect even if the relationship is strong. ## Fix two: a decision-relevant step Usually better for a non-technical audience. Instead of one unit or one standard deviation, quote the effect of the change someone might actually see or make: a 10,000-dollar difference in income, one extra year of tenure, a move from the 25th to the 75th percentile of the predictor. The number stays in the outcome's own units, which stakeholders can reason about, and the step is concrete. "Moving a customer from the bottom quartile of tenure to the top is associated with 180 dollars more annual spend, with a 95% interval of 120 to 240" answers a business question. "The tenure coefficient is 0.42" does not. ## Fix three: how much the predictor explains A third angle, answering a different question: how much of the variation in the outcome does this predictor account for beyond the others? Compare the model's explained variation with and without the predictor - the incremental contribution to R-squared. This is not the same as effect size and should not be conflated with it: a predictor can move the outcome a lot per unit while varying so little in the sample that it explains almost nothing, and vice versa. Decide first which question you are answering - *how much does it move the outcome* or *how much of the spread does it account for* - because they have different answers and different uses. ## Always attach uncertainty A point estimate with no interval invites false precision, and rankings are especially vulnerable: two predictors whose intervals overlap heavily should not be ordered at all. Present intervals next to every effect and say explicitly when a ranking is not supported. ## What none of these fixes Be honest about the limits, because a good interviewer will push here. **Correlated predictors share credit.** When two predictors overlap, the fit splits the association between them in a way that depends on the specification. Neither "importance" number is a stable property of either variable alone. **Comparability is not causality.** All of these are descriptive comparisons within the fitted model. Whether any of them tells you what would happen if you changed the predictor is a separate question requiring separate assumptions. **Significance is not magnitude.** A p-value combines effect size, noise and sample size. With a large enough sample a trivially small effect is highly significant; with a small sample an important one may not be. Ranking by p-value is exactly as broken as ranking by raw coefficients, for different reasons. ## The response, in one breath "That table is ranking our unit choices. I can make income first or last by choosing dollars or millions. Let me re-present it two ways: effects per standard deviation so the magnitudes are comparable, and effects over a realistic step in the original units so they are quotable, both with intervals. Where the intervals overlap I will say the order is not determined by the data." That is the answer of someone who has been in the room when a table like this drove a decision.
- How would you present a per-predictor effect to a non-technical stakeholder?In the outcome's own units, over a step they recognise, with an interval. "A 10,000-dollar increase in income is associated with about 340 dollars more annual spend, 95% interval 210 to 470, among customers alike on the other predictors." That names the step, the units, the uncertainty and the comparison, and needs no statistical vocabulary to act on.
- Does a smaller p-value mean a larger effect?No. A p-value reflects how precisely an effect is estimated, mixing the effect size, the noise and the sample size. A tiny effect measured on millions of rows can be extremely significant, while a large effect on a small sample may not be. Rank by effect size with intervals, never by p-value.
- Why doesn't putting every predictor on the same scale fully solve the ranking problem?Because standardising uses this sample's spreads, so a predictor observed over a narrow range looks weak regardless of the underlying relationship. Correlated predictors also split their shared association in a way that depends on the specification, so neither one's number is a stable property of that variable alone.
- When is comparing incremental contribution to R-squared the better summary?When the question is how much of the outcome's variation a predictor accounts for, rather than how much the outcome moves per unit of it. Those differ: a predictor with a steep slope but almost no variation in the sample explains little. Choose the summary that matches the decision, and say which one you used.
Ranking coefficients across mixed units is like ranking countries by the number on their price tags without converting currencies - the biggest number may just be the weakest unit.
saying these in an interview costs you the question
- Ranks predictors by raw coefficient magnitude across different units
- Ranks predictors by p-value or by significance stars
- Presents a ranking with no uncertainty attached
- Assumes standardising fully solves comparability between correlated predictors
- Treats an importance ranking as a list of levers to pull