skip to content

Feature Engineering and Data Preparation

You will learn the preprocessing that decides whether a classical model works: which algorithms need scaling, how to encode high-cardinality categoricals, how to fill gaps and rebalance a rare class.

on this pageshow

explore

questions

page 2 of 2

When would you use median/IQR robust scaling instead of z-score standardisation?

level: middleimportance: should knowfreq 46%

basics

~20 s

Use it when a column carries extreme values. Robust scaling subtracts the median and divides by the interquartile range, statistics that a handful of extreme points barely move, so the bulk of the data still lands on a usable scale.

open as a page

What goes wrong when SMOTE interpolates over one-hot encoded categorical features?

level: middleimportance: should knowfreq 46%

basics

~10 s

Interpolation produces fractional values on indicator columns, so a synthetic row comes out 0.5 on two different category flags. It represents no real category, and no such row could ever arrive at prediction time.

open as a page

Your class-weighted model outputs scores averaging 0.5 when only 2% of rows are positive — why?

level: seniorimportance: should knowfreq 46%

basics

~20 s

The weights changed the objective, so the fit targets the reweighted class mix rather than the real one. Balanced weights make both classes equally heavy, pulling the optimum toward 0.5. The scores still rank correctly but are no longer probabilities.

open as a page

You added 200 derived columns to a 1,500-row training table — how do you decide which earn their place?

level: seniorimportance: should knowfreq 44%

basics

~20 s

At 200 columns and 1,500 rows there are roughly seven rows per column, so some will look predictive by chance alone. Judge them by repeated cross-validation with the selection step refitted inside every fold, against a baseline without them.

open as a page

An account has no events in the 90-day window — which of its aggregates are zero and which are missing?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Counts and sums are genuinely zero — nothing happened, and that is an observation. Means, maxima and shares are undefined because there is nothing to summarise, and recency has no value at all. Encode those with a sentinel plus a companion flag, never a silent zero.

open as a page

Recursive feature elimination over 400 marketing-attribution columns runs for hours — what is it doing and how do you cut the cost?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Recursive feature elimination refits the model, ranks columns by its weights, drops the weakest, and repeats — hundreds of fits, times cross-validation folds. Cut cost by dropping a block per round, pre-screening with a cheap filter, or ranking with a faster model.

open as a page

Forward stepwise selection picked 8 of 60 sensor channels; why is the winning subset's cross-validated score optimistic?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Forward stepwise selection evaluated hundreds of candidate subsets and reported the winner. The maximum of many noisy estimates is biased upward, so part of the winner's margin is luck — even when the selector saw only training rows.

open as a page

In target encoding, a seller with three historical rows all converted — smooth toward the prior or bucket it as rare?

level: seniorimportance: should knowfreq 46%

basics

~10 s

Three rows cannot support a rate of 1.0. Smoothing shrinks the estimate toward the global rate, weighted by the row count, so some seller signal survives; a rare bucket discards it entirely. Prefer smoothing.

open as a page

A scoring request arrives with a null device-age field — what value do you fill and where does it come from?

level: seniorimportance: should knowfreq 46%

basics

~20 s

The fill comes from a statistic computed once on the training data and shipped with the model as a fixed parameter. Never recompute it from live traffic: the same request would then score differently depending on its neighbours.

open as a page

What does a one-hot encoder emit for a browser string never seen during training?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Either it raises an error or it emits all zeros across that column's indicators. All zeros is the dangerous case: if a reference level was dropped, that pattern already means the reference level, so the unknown browser is silently scored as it.

open as a page

Why does exponentiating a log-price model's prediction land on the median, not the mean?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Exponentiating is a convex map, so the average of log-scale predictions does not map back to the average price. Under symmetric log-scale errors it returns the conditional median; with normal log errors of standard deviation s, multiply by exp(s^2 / 2) to recover the mean.

open as a page

When does adding SMOTE samples make a classifier's decision boundary worse?

level: seniorimportance: should knowfreq 50%

basics

~20 s

SMOTE hurts when the classes overlap or the minority contains mislabelled points. Interpolating between two minority rows that sit on opposite sides of a majority cloud plants synthetic positives inside majority territory, pushing the boundary outward and costing precision.

open as a page

Your TF-IDF matrix has 50,000 columns for 20,000 documents — how do you decide what to prune?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Check the arithmetic before pruning: stored sparsely, 20,000 documents touching a hundred terms each is a few million values, not a billion. Cut with document-frequency thresholds at both ends, and let held-out folds decide how far to go.

open as a page

A scorecard's age bands break monotonic WoE in one bin — how do you fix the binning?

level: seniorimportance: should knowfreq 52%

basics

~20 s

First decide whether the dip is real or sampling noise in a thin band. If it is noise, merge that band with the neighbour it sits closest to in weight of evidence and re-cut under a monotone constraint, accepting a small loss of information value.

open as a page

With a 2% positive rate, how do you decide between class weights, resampling, and moving the operating point?

level: principalimportance: should knowfreq 58%

basics

~20 s

Decide by what is actually broken. If the model already ranks cases well and only the cut-off is wrong, move the operating point - it is free and reversible. Reweight when the fit itself ignores the rare class.

open as a page

How do you choose an encoding for a 40,000-level seller ID that must be refreshed and served daily?

level: principalimportance: should knowfreq 37%

basics

~20 s

Decide on four axes: how much signal the identity carries, what the model family can consume, what state serving can refresh, and who must explain the feature. Start cheap and escalate only when a holdout says it pays.

open as a page

When is discretising a continuous predictor into bins worth the information it destroys?

level: principalimportance: should knowfreq 36%

basics

~20 s

Rarely for accuracy, often for everything else. Binning discards within-bin variation and imposes arbitrary cutpoints, so it usually costs predictive power. It earns its place when bins buy interpretability, stable reporting, or let a linear model express a non-monotone effect.

open as a page

What does a variance-threshold filter remove from a feature matrix, and when does it drop something useful?

level: juniorimportance: nice to knowfreq 30%

basics

~20 s

A variance-threshold filter drops every column whose spread across rows falls below a cutoff, removing constant and near-constant features. It ignores the target and depends on units, so it can discard a rare binary flag that predicts strongly.

open as a page

When would you give each training row its own weight rather than one weight per class?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

When the cost of an error varies row by row, not just class by class. In insurance-claim triage, weighting each claim by the euro value at risk makes the model spend its capacity where the money is.

open as a page

Why encode a pickup hour as a sine-cosine pair instead of the integer 0-23?

level: middleimportance: nice to knowfreq 40%

basics

~20 s

Hour 23 and hour 0 are one hour apart but 23 units apart as integers. Mapping the hour onto a circle with sin(2pih/24) and cos(2pih/24) makes them neighbours again, which matters to any distance- or magnitude-based model.

open as a page

Why add share-of-total and per-day rate features next to raw per-entity event counts?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

A raw count mixes how active an entity is overall with how it splits that activity and how long it was observed. Dividing by the entity's total gives composition, and dividing by days observed gives intensity, so the model can separate a heavy user from a lopsided one.

open as a page

How does the hashing trick encode millions of ad publisher domains, and what do collisions cost?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

A hash function maps each domain to a bucket index modulo a fixed bucket count, say 2^20, and that bucket is the feature slot. Nothing is stored, so new domains need no special case. Colliding domains share one blended weight.

open as a page

When does kNN or iterative imputation beat filling a column with its median?

level: middleimportance: nice to knowfreq 38%

basics

~20 s

When the incomplete column is strongly predictable from the other columns. Both borrow that structure — kNN from similar rows, iterative imputation from a regression on the other features — where a median gives every gap the same answer.

open as a page

When do character 3-5-grams beat word unigrams as features for matching messy product titles?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

Character n-grams win when the tokens themselves are unreliable: misspelled or run-together brand names, inconsistent punctuation, model codes. A typo changes only a few of a word's character n-grams, whereas it destroys the word unigram entirely, so overlap survives.

open as a page

Should a one-hot column drop a reference level before an L2-penalised linear fit?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Usually no. Dropping a level exists to remove the exact redundancy between the full set of indicators and the intercept, which an L2 penalty already resolves on its own. Keeping every level and leaving the intercept unpenalised treats all levels symmetrically.

open as a page

Why can min-max scaling still leave one feature dominating a distance-based model?

level: seniorimportance: nice to knowfreq 31%

basics

~20 s

Min-max equalises each feature's range, not its spread. A heavy-tailed column whose maximum sits far above the bulk gets compressed near zero after scaling, so a well-spread bounded column ends up supplying almost all of the distance.

open as a page

What does rank-normalising a heavy-tailed page-load-time feature gain and cost?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Rank normalisation replaces each value with its rank, rescaled to a uniform or Gaussian shape. It flattens any tail with no parametric assumption, but it keeps only the ordering: how far apart two values were is discarded.

open as a page

How do you turn a scorecard's fitted log-odds into points at base 600 with a PDO of 20?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Apply score = offset + factor times ln(odds), where factor = PDO / ln 2, so 20 / 0.693 is about 28.85, and offset = 600 minus factor times the log of the anchor odds. Every doubling of the odds then adds exactly 20 points.

open as a page

Two teams export an active-user label under different definitions — which one do you train on?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Neither, until the definition is settled. A label whose meaning changed is a specification defect, not noise: choose one definition tied to the decision the model serves, write it down with edge cases, and rebuild the target consistently.

open as a page

showing 31–59 of 59