skip to content

You added 200 derived columns to a 1,500-row training table — how do you decide which earn their place?

level: seniorimportance: should knowfreq 44%

answer

  1. count rows per column, not columns
  2. chance winners appear at this width
  3. where does the selection step happen?
  4. inside every fold, against a baseline

basics

~20 s

At 200 columns and 1,500 rows there are roughly seven rows per column, so some will look predictive by chance alone. Judge them by repeated cross-validation with the selection step refitted inside every fold, against a baseline without them.

solid answer

~50 s

The number that matters first is the ratio: 200 columns over 1,500 rows is about seven rows per column, thin enough that the best-looking derived feature may simply be the luckiest one. The measurement discipline matters more than the shortlist. I would fix a baseline model on the original columns, then add the derived block and compare under repeated k-fold cross-validation, reporting the fold-to-fold spread as well as the mean — with 1,500 rows a single split is a few hundred rows and a one-point difference is noise. Critically, any selection step goes *inside* each training fold; ranking features on the full table and then cross-validating gives an estimate that has already seen the held-out rows. I would lean on regularisation rather than hand-picking, add or drop correlated columns as blocks, and require a gain clearly exceeding the fold spread before keeping a block someone must maintain and reproduce at scoring time.

go deeper

for a junior

Be ready to state the ratio and what it implies: roughly seven rows per column means the model has very little evidence per feature, and a good training score proves nothing.

for a middle

Explain why a chance winner appears when you screen 200 columns, and what regularisation does about it — L2 shrinking everything, L1 zeroing some coefficients outright.

for a senior

Show the evaluation discipline: selection refitted inside every training fold, repeated k-fold because 300 held-out rows are noisy, a baseline comparison, and a test set you touch once.

for a principal

Own the tradeoff between a wide generated feature sweep and a small domain-motivated set, including the maintenance and train-versus-serve risk each kept column adds to the team's load.

## Start with the ratio 200 columns and 1,500 rows is about 7.5 rows per column. Classical practice in regression modelling asks for something on the order of ten to twenty observations per predictor before coefficient estimates are trustworthy, so this table is on the wrong side of every rule of thumb. As the number of columns approaches the number of rows, an unpenalised linear model can fit the training data more and more closely regardless of whether any signal exists — the extra flexibility is spent on noise. The training score stops carrying information about generalisation long before the columns run out. The practical consequence is subtler and more dangerous than "it overfits". It is that **some derived columns will look genuinely predictive on this sample and be pure chance**. If you generate 200 columns of random numbers on 1,500 rows, the best-correlated one will show a correlation with the target that looks worth reporting. Chance maxima grow with the number of things you look at. ## The measurement mistake that hides all of this The classic error is: rank all 200 columns by their association with the target, keep the top 30, then run cross-validation on a model using those 30. That estimate is optimistically biased, sometimes wildly, because the ranking was computed on **every row, including the rows each fold later holds out**. The survivors are exactly the columns that happened to fit the noise in the held-out rows too, so the held-out score is no longer an honest simulation of new data. The fix is a principle worth stating plainly in an interview: **selection is part of the model**. Everything that looks at the target — filtering by correlation, recursive elimination, choosing a threshold — is refitted inside each training fold, on that fold's rows only, and the held-out fold sees the result. The estimate usually drops when you do this, and that drop is the size of the illusion you were previously reporting. One useful nuance: a *row-local* derivation cannot leak this way. A date part, a ratio of two columns on the same row, a difference of two dates — each uses only that row's own values, so computing them once for the whole table is safe. What must move inside the fold is anything fitted on a column's distribution across rows, and above all the selection step. ## What honest evaluation looks like at n = 1,500 - **Repeated k-fold**, not a single split. A 20% holdout is 300 rows; the difference between 0.84 and 0.85 accuracy there is a handful of rows changing side. Repeating the whole cross-validation with different partitions and reporting the mean and spread tells you whether a gain is real. - **A locked test set touched once.** Everything above is model selection; keep a final set that no decision has seen, and look at it at the end. - **A baseline with none of the derived columns.** "The model scores 0.87" is not a result. "The derived block adds 0.03 over the raw-column baseline, larger than the 0.01 fold-to-fold standard deviation" is. ## Let the model do the selecting Hand-picking 30 of 200 columns is itself a high-variance decision at this sample size. Regularisation is usually the better instrument. An L2 penalty shrinks all coefficients toward zero without setting any to exactly zero, and when several columns are correlated it tends to spread weight across them, which stabilises the fit. An L1 penalty drives some coefficients exactly to zero, producing a sparse model, but among a group of correlated columns it tends to keep one somewhat arbitrarily and zero the rest — so do not read its choice as a statement about which column matters. Either way, the penalty strength is a hyperparameter tuned inside the same cross-validation. Tree ensembles are not immune either. A random forest samples a subset of candidate columns at each split; drowning 20 useful columns in 180 noisy ones lowers the chance that a useful one is even offered at a given node, and accuracy can fall as a result. ## Group correlated derived features Derived columns arrive correlated by construction: a ratio and its two parents, or hour, is-weekend and day-of-week from one timestamp. Ablating them one at a time understates their value, because the rest of the group covers for the one you removed. Add and remove them as **blocks** that correspond to an idea — "the calendar block", "the affordability-ratio block" — and you get a decision you can act on and explain. ## The cost side of the ledger Every kept column is code someone maintains, a dependency (an external holiday calendar), and a place where the training computation and the scoring computation can silently diverge. On a 1,500-row table a marginal, unexplainable column is rarely worth that. The strongest answer here is a preference: a small number of derived features motivated by domain reasoning, validated as blocks against a baseline, over a generated sweep that widens the table and buys a gain indistinguishable from fold noise.

  • Why is picking the top 30 features by correlation before cross-validating optimistic?
    Because the ranking saw every row, including the rows each fold later holds out. The columns that survive are the ones that fit the noise in those held-out rows as well, so the held-out score no longer simulates new data. Move the ranking inside each training fold and the estimate typically drops — that drop is the bias you were reporting as performance.
  • Date parts and ratios are computed row by row — do those also have to sit inside the fold?
    For the honesty of the estimate, no. A row-local derivation uses only that row's own values, so it cannot carry information between rows and can be computed once for the whole table. What must go inside the fold is anything fitted across rows — a clipping percentile, a normalisation constant — and, above all, any step that looks at the target.
  • How would you decide whether a whole block of derived features is worth keeping?
    Compare it against a baseline without the block under repeated cross-validation, and require the mean gain to clearly exceed the fold-to-fold spread, then confirm on a test set touched once. Weigh that gain against the maintenance cost: the code, any external dependency such as a holiday calendar, and the risk of computing the block differently at scoring time.
  • Does adding 180 uninformative columns hurt a random forest?
    It can. A forest considers a random subset of columns at each split, so the more noise columns there are, the lower the chance a genuinely useful one is offered at a given node. Trees then split on noise, individual trees get weaker, and accuracy degrades. Trees tolerate irrelevant columns better than an unpenalised linear model does, but they are not immune.

Testing 200 columns on 1,500 rows is like letting 200 people each guess a coin sequence: someone will look like a psychic, and the trick is designing the check so their next guesses are the ones that count.

saying these in an interview costs you the question

  • Adding features until the validation score stops improving
  • Selecting features on the full dataset before cross-validating
  • Reporting a small cross-validation gain without the fold spread
  • Assuming tree ensembles are immune to hundreds of noise columns
  • Judging a correlated block of features one column at a time

context