Feature Engineering and Data Preparation
You will learn the preprocessing that decides whether a classical model works: which algorithms need scaling, how to encode high-cardinality categoricals, how to fill gaps and rebalance a rare class.
on this pageshowhide
explore
- Numeric Features and Scaling21 questions
- Scaling and Standardization4 questions
- Skew Transforms and Binning5 questions
- Outlier Treatment4 questions
- Derived and Date Features4 questions
- Entity Aggregates from Events4 questions
- Categorical and Text Encoding17 questions
- One-Hot vs Ordinal4 questions
- High-Cardinality Categories5 questions
- TF-IDF and Bag-of-Words4 questions
- Weight of Evidence Binning4 questions
- Missingness and Label Quality8 questions
- Imputation Strategies4 questions
- Label Noise and Duplicates4 questions
- Imbalance and Feature Selection13 questions
- Class Weights and Thresholds4 questions
- SMOTE and Resampling4 questions
- Filter, Wrapper, Embedded5 questions
questions
page 2 of 2When would you use median/IQR robust scaling instead of z-score standardisation?
basics
~20 sUse it when a column carries extreme values. Robust scaling subtracts the median and divides by the interquartile range, statistics that a handful of extreme points barely move, so the bulk of the data still lands on a usable scale.
What goes wrong when SMOTE interpolates over one-hot encoded categorical features?
basics
~10 sInterpolation produces fractional values on indicator columns, so a synthetic row comes out 0.5 on two different category flags. It represents no real category, and no such row could ever arrive at prediction time.
Your class-weighted model outputs scores averaging 0.5 when only 2% of rows are positive — why?
basics
~20 sThe weights changed the objective, so the fit targets the reweighted class mix rather than the real one. Balanced weights make both classes equally heavy, pulling the optimum toward 0.5. The scores still rank correctly but are no longer probabilities.
You added 200 derived columns to a 1,500-row training table — how do you decide which earn their place?
basics
~20 sAt 200 columns and 1,500 rows there are roughly seven rows per column, so some will look predictive by chance alone. Judge them by repeated cross-validation with the selection step refitted inside every fold, against a baseline without them.
An account has no events in the 90-day window — which of its aggregates are zero and which are missing?
basics
~20 sCounts and sums are genuinely zero — nothing happened, and that is an observation. Means, maxima and shares are undefined because there is nothing to summarise, and recency has no value at all. Encode those with a sentinel plus a companion flag, never a silent zero.
Recursive feature elimination over 400 marketing-attribution columns runs for hours — what is it doing and how do you cut the cost?
basics
~20 sRecursive feature elimination refits the model, ranks columns by its weights, drops the weakest, and repeats — hundreds of fits, times cross-validation folds. Cut cost by dropping a block per round, pre-screening with a cheap filter, or ranking with a faster model.
Forward stepwise selection picked 8 of 60 sensor channels; why is the winning subset's cross-validated score optimistic?
basics
~20 sForward stepwise selection evaluated hundreds of candidate subsets and reported the winner. The maximum of many noisy estimates is biased upward, so part of the winner's margin is luck — even when the selector saw only training rows.
In target encoding, a seller with three historical rows all converted — smooth toward the prior or bucket it as rare?
basics
~10 sThree rows cannot support a rate of 1.0. Smoothing shrinks the estimate toward the global rate, weighted by the row count, so some seller signal survives; a rare bucket discards it entirely. Prefer smoothing.
A scoring request arrives with a null device-age field — what value do you fill and where does it come from?
basics
~20 sThe fill comes from a statistic computed once on the training data and shipped with the model as a fixed parameter. Never recompute it from live traffic: the same request would then score differently depending on its neighbours.
What does a one-hot encoder emit for a browser string never seen during training?
basics
~20 sEither it raises an error or it emits all zeros across that column's indicators. All zeros is the dangerous case: if a reference level was dropped, that pattern already means the reference level, so the unknown browser is silently scored as it.
Why does exponentiating a log-price model's prediction land on the median, not the mean?
basics
~20 sExponentiating is a convex map, so the average of log-scale predictions does not map back to the average price. Under symmetric log-scale errors it returns the conditional median; with normal log errors of standard deviation s, multiply by exp(s^2 / 2) to recover the mean.
When does adding SMOTE samples make a classifier's decision boundary worse?
basics
~20 sSMOTE hurts when the classes overlap or the minority contains mislabelled points. Interpolating between two minority rows that sit on opposite sides of a majority cloud plants synthetic positives inside majority territory, pushing the boundary outward and costing precision.
Your TF-IDF matrix has 50,000 columns for 20,000 documents — how do you decide what to prune?
basics
~20 sCheck the arithmetic before pruning: stored sparsely, 20,000 documents touching a hundred terms each is a few million values, not a billion. Cut with document-frequency thresholds at both ends, and let held-out folds decide how far to go.
A scorecard's age bands break monotonic WoE in one bin — how do you fix the binning?
basics
~20 sFirst decide whether the dip is real or sampling noise in a thin band. If it is noise, merge that band with the neighbour it sits closest to in weight of evidence and re-cut under a monotone constraint, accepting a small loss of information value.
With a 2% positive rate, how do you decide between class weights, resampling, and moving the operating point?
basics
~20 sDecide by what is actually broken. If the model already ranks cases well and only the cut-off is wrong, move the operating point - it is free and reversible. Reweight when the fit itself ignores the rare class.
How do you choose an encoding for a 40,000-level seller ID that must be refreshed and served daily?
basics
~20 sDecide on four axes: how much signal the identity carries, what the model family can consume, what state serving can refresh, and who must explain the feature. Start cheap and escalate only when a holdout says it pays.
When is discretising a continuous predictor into bins worth the information it destroys?
basics
~20 sRarely for accuracy, often for everything else. Binning discards within-bin variation and imposes arbitrary cutpoints, so it usually costs predictive power. It earns its place when bins buy interpretability, stable reporting, or let a linear model express a non-monotone effect.
What does a variance-threshold filter remove from a feature matrix, and when does it drop something useful?
basics
~20 sA variance-threshold filter drops every column whose spread across rows falls below a cutoff, removing constant and near-constant features. It ignores the target and depends on units, so it can discard a rare binary flag that predicts strongly.
When would you give each training row its own weight rather than one weight per class?
basics
~20 sWhen the cost of an error varies row by row, not just class by class. In insurance-claim triage, weighting each claim by the euro value at risk makes the model spend its capacity where the money is.
Why encode a pickup hour as a sine-cosine pair instead of the integer 0-23?
basics
~20 sHour 23 and hour 0 are one hour apart but 23 units apart as integers. Mapping the hour onto a circle with sin(2pih/24) and cos(2pih/24) makes them neighbours again, which matters to any distance- or magnitude-based model.
How does the hashing trick encode millions of ad publisher domains, and what do collisions cost?
basics
~20 sA hash function maps each domain to a bucket index modulo a fixed bucket count, say 2^20, and that bucket is the feature slot. Nothing is stored, so new domains need no special case. Colliding domains share one blended weight.
When does kNN or iterative imputation beat filling a column with its median?
basics
~20 sWhen the incomplete column is strongly predictable from the other columns. Both borrow that structure — kNN from similar rows, iterative imputation from a regression on the other features — where a median gives every gap the same answer.
When do character 3-5-grams beat word unigrams as features for matching messy product titles?
basics
~20 sCharacter n-grams win when the tokens themselves are unreliable: misspelled or run-together brand names, inconsistent punctuation, model codes. A typo changes only a few of a word's character n-grams, whereas it destroys the word unigram entirely, so overlap survives.
Should a one-hot column drop a reference level before an L2-penalised linear fit?
basics
~20 sUsually no. Dropping a level exists to remove the exact redundancy between the full set of indicators and the intercept, which an L2 penalty already resolves on its own. Keeping every level and leaving the intercept unpenalised treats all levels symmetrically.
Why can min-max scaling still leave one feature dominating a distance-based model?
basics
~20 sMin-max equalises each feature's range, not its spread. A heavy-tailed column whose maximum sits far above the bulk gets compressed near zero after scaling, so a well-spread bounded column ends up supplying almost all of the distance.
What does rank-normalising a heavy-tailed page-load-time feature gain and cost?
basics
~20 sRank normalisation replaces each value with its rank, rescaled to a uniform or Gaussian shape. It flattens any tail with no parametric assumption, but it keeps only the ordering: how far apart two values were is discarded.
How do you turn a scorecard's fitted log-odds into points at base 600 with a PDO of 20?
basics
~20 sApply score = offset + factor times ln(odds), where factor = PDO / ln 2, so 20 / 0.693 is about 28.85, and offset = 600 minus factor times the log of the anchor odds. Every doubling of the odds then adds exactly 20 points.
Two teams export an active-user label under different definitions — which one do you train on?
basics
~20 sNeither, until the definition is settled. A label whose meaning changed is a specification defect, not noise: choose one definition tied to the decision the model serves, write it down with edge cases, and rebuild the target consistently.
showing 31–59 of 59