Should a 1-5 satisfaction rating enter a regression as one numeric predictor or as four dummies?
answer
- one number smuggles in an assumption
- does 1 to 2 equal 4 to 5
- one degree of freedom versus four
- the models are nested, so test the restriction
- F-test on the three extra parameters
basics
~20 sEntering the rating as a single number assumes every one-point step moves the outcome by the same amount, for one degree of freedom. Four dummies let each level move freely, for four. Test that restriction rather than assuming it.
solid answer
~50 sCoding the rating 1 to 5 as a number smuggles in a strong assumption: that the gap from 1 to 2 buys the same change in the outcome as the gap from 4 to 5. Ratings rarely behave that way — the jump from 4 to 5 is often much larger than the middle steps. The dummy version drops the assumption: with five levels you add four dummies against a reference such as rating 1, and each level gets its own free contrast. The cost is degrees of freedom, four parameters instead of one, which bites when levels are thin. The honest approach is to fit both: the numeric model is nested inside the dummy model, so an F-test on the three extra degrees of freedom tests whether the equal-step restriction is tenable. Plot the contrasts too — if they rise roughly linearly, the numeric coding is a fair simplification.
go deeper
Know that a numbered rating is not automatically a number to the model, and that entering it as 1 to 5 asserts every step is worth the same change in the outcome.
Explain the tradeoff quantitatively: one parameter versus four for a five-level scale, what the extra flexibility buys, and how the equal-step restriction can be tested rather than assumed.
Show judgment on real data — reading the coefficient pattern for a threshold or top-box shape, handling levels with too few rows, and choosing a grouping on substantive grounds rather than by p-value search.
Own the standard for how survey and rating scales enter models across the team, including when a simplified coding is worth its assumption and how such simplifications are documented for anyone reading the results later.
## The two codings A satisfaction rating takes the values 1, 2, 3, 4, 5. There are two natural ways to put it into a regression. **Numeric coding.** Enter the rating as one column of numbers. The model gets one slope, `b`, and the fitted contribution is `b * rating`. **Dummy coding.** Treat the rating as a categorical variable with five levels. With an intercept you add four dummies against a reference level — say rating 1 — and each of the other levels gets its own coefficient, interpreted as that level minus rating 1. ## What the numeric coding assumes The numeric coding is not neutral. It imposes two claims at once. First, **equal spacing**: moving from 1 to 2 is the same distance as moving from 3 to 4. That is a claim about the measurement, and for a rating scale it is a convention rather than a fact — respondents do not calibrate their 2s and 4s to a common ruler. Second, **linearity in the outcome**: each additional rating point shifts the expected outcome by exactly `b`, everywhere on the scale. Real rating scales frequently violate this. Top-box behaviour is the classic case: 1 through 4 barely differ in downstream behaviour while 5 stands well apart, so the true pattern is flat then a jump. A single slope fitted to that pattern splits the difference and understates both the flatness and the jump. The dummy coding assumes neither. It lets the five levels sit at any five heights, in any order. It does *not* even impose monotonicity, which is a mild disadvantage if you have strong prior reason to expect it. ## The degrees-of-freedom tradeoff The numeric coding spends **one** parameter; the dummy coding spends **four**. Three consequences follow. With plenty of data, four parameters are cheap and the flexibility is close to free. With a small sample, or with a scale where some levels are rare — ratings of 2 often are — those four coefficients are estimated from few observations each and come back with wide standard errors. A level with a handful of rows produces a contrast so noisy it carries little information, and if a level has no observations at all the model cannot estimate it. The numeric version also concentrates all the evidence into a single slope, which is more precisely estimated and easier to communicate: 'each extra rating point is worth X'. Precision bought with an assumption is only a bargain if the assumption holds. Finally, the numeric version extrapolates naturally to a rating outside the observed range while the dummy version cannot represent an unseen level at all — a real consideration when scoring new data. ## Deciding between them The two models are **nested**: the numeric specification is the dummy specification with the restriction that the four contrasts lie on an equally spaced straight line. That makes the comparison a standard nested-model test. Fit both, and test the 3 extra degrees of freedom that the dummy model uses beyond the numeric one — 4 parameters versus 1 — with an F-test comparing the two residual sums of squares. A small p-value says the equal-step pattern is not tenable; a large one says the simpler coding is not detectably worse. Don't stop at the test. Plot the four estimated contrasts against the rating level with their confidence intervals. The shape tells you *how* the linear coding fails when it does: a jump at the top, a threshold in the middle, a U-shape. That shape is often the finding the stakeholder actually wants. A middle course is common in practice. If the pattern is flat-then-jump, collapse the scale to a small number of meaningful groups — for example a single 'top box' indicator — which costs one parameter and captures the real structure. Grouping decisions should be made on substantive grounds or on a plot of the coefficients, not by trying every grouping and keeping the one with the best p-value. ## What interviewers listen for The weak answer is a rule ('always use dummies for categories') with no mention of what the numeric coding costs or buys. The strong answer names the assumption by its content — equal spacing and a constant per-point effect — says how many degrees of freedom each coding spends, and gives a concrete way to check the restriction rather than asserting one coding is correct. Mentioning sparse levels and unseen levels at scoring time shows you have actually fitted these models on real data.
- How exactly do you test whether the numeric coding is adequate?The numeric model is nested inside the dummy model — it is the dummy model restricted to equally spaced, collinear contrasts. Fit both and compare residual sums of squares with an F-test on the 3 extra degrees of freedom, since 5 levels cost 4 parameters as dummies versus 1 as a slope. A small p-value rejects the equal-step restriction.
- The dummy contrasts come back as roughly flat for ratings 1 to 4 and much higher for 5. What would you do?That is a top-box pattern, not a linear one, so a single numeric slope would misrepresent it. Replace the scale with a single indicator for rating 5, which costs one parameter, captures the real structure and communicates cleanly. Decide the grouping from the coefficient plot or substantive knowledge, not by searching groupings for the best p-value.
- What breaks if one rating level has only three observations?Its dummy coefficient is estimated from those three rows, so the contrast has a very wide confidence interval and contributes little. With zero observations the level cannot be estimated at all, and a row with that level cannot be scored. Collapsing the sparse level into an adjacent one, or falling back to the numeric coding, buys precision at the cost of an assumption.
- Does the dummy coding assume the effect is monotone in the rating?No, and that cuts both ways. Dummies let the five levels sit at any heights in any order, which is exactly what you want when the relationship might be non-monotone. But if you have strong prior reason to expect a monotone effect, the dummy fit can show a non-monotone wobble that is pure sampling noise, particularly in thin levels.
saying these in an interview costs you the question
- Treats the numeric coding as assumption-free
- Says always use dummies without pricing the degrees of freedom
- Ignores that the two models are nested and testable
- Fits dummies on levels with a handful of observations
- Picks a grouping of levels by searching for significance