How do you decide whether to keep 199 city dummies in a regression fit on 5,000 rows?
answer
- count rows per city, not rows total
- 199 dummies plus intercept is 200 parameters
- the long tail is the problem
- adjusting for city versus reporting on it
- collapse the tail or shrink toward the mean
basics
~20 sJudge it by observations per city, not total rows. With 5,000 rows, 200 parameters leave about 4,800 residual degrees of freedom, yet a city with five rows still yields a contrast far too noisy to act on.
solid answer
~50 sThe global arithmetic looks fine: 199 dummies plus an intercept is 200 parameters, leaving about 4,800 residual degrees of freedom. The real constraint is local — each city coefficient rests on that city's own rows, so a city with five observations gives an enormous standard error, and the long tail dominates a 200-city variable. Ask what the dummies are for. If they adjust for city so another coefficient is clean, noisy per-city estimates are acceptable and never reported. If per-city effects are the deliverable, most will not survive honest uncertainty. Options: keep the well-populated cities and collapse the tail into a single Other level, drop to a coarser geography, or shrink city estimates toward the overall mean. Fix whichever rule from row counts before looking at the outcome, and decide in advance what happens when an unseen city arrives at scoring time.
go deeper
Know that a categorical with many levels becomes many columns — 200 cities means 199 dummies with an intercept — and that each of those coefficients needs its own data to be worth anything.
Be able to compute the parameter and residual degree-of-freedom cost, and explain why a level with a handful of rows yields a coefficient with a very wide confidence interval.
Show you look at the rows-per-level distribution first, handle the sparse tail with a rule fixed before seeing the outcome, and have an answer ready for levels unseen at scoring time.
Own the framing: whether city is being adjusted for or reported on, how much bias the team will accept for stability, and the standard for high-cardinality categoricals so per-level noise never reaches a decision as a ranked table.
## Do the arithmetic first, then distrust it A categorical with 200 cities enters an intercept model as 199 dummies. With 5,000 rows the model spends 200 parameters and leaves roughly 4,800 residual degrees of freedom. By the global count nothing is wrong: the model is comfortably identified and there is ample residual information for the error variance. That global count is the wrong lens. Each city's coefficient is informed only by the rows belonging to that city (and to the baseline city, since the coefficient is a contrast). If the 5,000 rows were split evenly, each city would carry 25 rows — thin but workable. Real categorical variables are never even. A handful of large cities carry most of the data and a long tail carries three, two, or one row each. A city with one observation has its residual fitted exactly and its coefficient tells you nothing about that city; it is a memorised row. So the decision rule is not 'is n large enough for 200 parameters' but 'what is the distribution of rows per level'. Look at that distribution before anything else. ## Ask what the dummies are for The right answer depends entirely on the model's job, and this is the part a lead is expected to own. **Adjustment.** If city is a nuisance variable — you want a clean estimate of some other predictor and city confounds it — the individual city coefficients are never reported and their noisiness costs little. What you are buying is the removal of between-city variation. Keeping many dummies is defensible here; the price is precision on the predictor you care about, since dummies that soak up variation correlated with it will widen its standard error. **Reporting.** If stakeholders want to know how each city performs, most of 199 contrasts will not survive honest uncertainty, and presenting a ranked table of them invites the classic error of celebrating whichever small city happens to top the list on five observations. Small-sample extremes dominate any ranking of noisy estimates. **Prediction.** If the model will score new rows, dummies have a hard operational edge: an unseen city has no column and cannot be scored. Any specification with a high-cardinality factor needs a documented answer for that case before it ships. ## The options **Keep the well-populated levels, collapse the tail.** Set a minimum row count, keep those cities as their own levels, and pool everything below the threshold into a single Other level. This is cheap, interpretable, and gives unseen cities a natural home at scoring time. The threshold must be chosen from row counts alone. Choosing it by trying several and keeping the version whose results look best is selecting on the outcome, and the resulting p-values are not trustworthy. **Move up the hierarchy.** Replace city with region, state or market tier — a coarser categorisation with a defensible number of well-populated levels. This trades resolution for stability and is often what the decision actually needs. **Shrink toward the mean.** Estimate the city effects with shrinkage, so that a city with hundreds of rows keeps close to its own average while a city with three rows is pulled most of the way toward the overall mean. This is the principled version of collapsing the tail: it uses a continuum instead of a threshold, and it degrades gracefully as row counts vary. The cost is a more complex model and estimates that are deliberately biased toward the centre in exchange for much lower variance. **Drop the variable.** If city is neither a confounder you must adjust for nor a dimension anyone will act on, 199 columns of mostly-noise is a poor trade. Being willing to say this is a signal of judgment, not laziness. ## Framing the tradeoff Every option above is the same bias-variance decision in different clothing. Full dummies are the low-bias, high-variance end: each city gets its own free parameter and its own noise. Dropping the variable is the high-bias, low-variance end. Collapsing the tail and shrinking sit in between, and the right point on that line depends on how the estimates will be used and on how expensive a confidently wrong per-city number would be. ## What to say in an interview Start with the row-count distribution rather than the total, name what the dummies are for, then choose. Mention the scoring-time problem with unseen levels, and make the point that any thresholding or grouping rule must be fixed from the predictors before the outcome is consulted. A candidate who answers only with the residual degrees-of-freedom arithmetic has answered a bookkeeping question, not the one that was asked.
- What residual degrees of freedom remain, and why is that not the binding constraint?With 199 dummies plus an intercept the model spends 200 parameters, leaving about 4,800 residual degrees of freedom from 5,000 rows — plenty in aggregate. The binding constraint is per-level: each city coefficient rests on that city's own rows, so a city with five observations gives a contrast with a huge standard error no matter how large the total sample is.
- How would you set a threshold for collapsing rare cities into an Other level?Choose it from the row-count distribution before looking at the outcome — for example the smallest count that still supports a usable standard error, or a coverage rule that keeps the cities accounting for most of the data. Trying several thresholds and keeping the one whose results look best selects on the outcome and invalidates the reported p-values.
- What happens at scoring time when a city appears that was not in the training data?It has no dummy column, so the model cannot represent it and the row cannot be scored as specified. The specification needs a decision made in advance: route unseen levels to an Other level that exists in the model, fall back to a coarser geography, or refuse to score and escalate. Discovering this in production is the avoidable failure.
- Why does shrinking city effects toward the overall mean help here?Shrinkage lets the amount of pooling depend on how much data a city has: well-observed cities stay near their own averages while sparse cities are pulled toward the overall mean, where the data cannot support a distinctive estimate. It buys a large variance reduction for a small, deliberate bias, and it degrades smoothly instead of relying on a hard threshold.
saying these in an interview costs you the question
- Judges feasibility from total rows rather than rows per city
- Ranks cities by noisy coefficients and acts on the top ones
- Chooses the rare-level threshold by which results look best
- Has no plan for a city unseen at scoring time
- Treats the choice as mechanical rather than purpose-driven