skip to content

Feature Engineering and Data Preparation

You will learn the preprocessing that decides whether a classical model works: which algorithms need scaling, how to encode high-cardinality categoricals, how to fill gaps and rebalance a rare class.

on this pageshow

explore

questions

page 1 of 2

For an electricity-demand model, which features would you extract from a raw timestamp column?

level: juniorimportance: must knowfreq 70%

answer

  1. one column hides several calendars
  2. the model cannot read a timestamp
  3. weekdays, weekends and holidays differ
  4. convert to local time before extracting hour

basics

~20 s

Split the timestamp into the calendar parts demand depends on: hour of day, day of week, an is-weekend flag, month or quarter, and a public-holiday flag. The raw timestamp alone only lets a model learn a trend.

solid answer

~50 s

A raw timestamp is one ever-increasing number, so no two rows share a value and the model can only read a trend out of it. I explode it into the calendar parts the domain actually turns on: hour of day for the daily load shape, day of week plus an is-weekend flag, month or quarter for the heating and cooling season, and a public-holiday flag, because a holiday behaves like a Sunday on a Tuesday. Two details decide whether this works live: extract the hour in the *local* time of the region modelled, not UTC, so the evening peak stays in one place across daylight-saving changes; and confirm the holiday calendar covers the dates you must score. I would not keep the raw timestamp as a numeric feature — a tree cannot extrapolate past the thresholds it learned, so every future row collapses into the last region it saw.

go deeper

for a junior

Be ready to name the parts out loud — hour, day of week, is-weekend, month, holiday flag — and say in one sentence why the raw timestamp on its own is useless to the model.

for a middle

Explain the mechanism: date parts create values that repeat across rows so the model can pool evidence, and an explicit flag saves a tree several splits over a raw integer encoding.

for a senior

Show the production judgment: local time versus UTC, daylight saving, whether the holiday calendar covers future scoring dates, and why leaving the raw timestamp in breaks a tree the moment it scores tomorrow.

for a principal

Own the tradeoff between a rich calendar feature set and the dependency it creates — an external holiday calendar per region is a contract someone must maintain and reproduce identically at scoring time.

## Why a timestamp is not yet a feature A timestamp is a single number that counts upward forever — seconds since an epoch, or an equivalent encoding. A model can only use a column through the operations its hypothesis class allows: a linear model multiplies the column by one coefficient, a tree compares it against thresholds. Neither operation can express "demand is high at 18:00 on weekdays and low on public holidays", because nothing in the raw number tells the model that two rows six months apart share an hour of day. Date-part extraction fixes exactly that. It converts one non-repeating column into several columns whose values **repeat across rows**, so the model can pool every 18:00 it has ever seen into one estimate. That pooling is the whole point — it is what turns 20,000 unique timestamps into a handful of learnable regularities. ## Which parts to extract For electricity demand the useful parts follow the physics and the human calendar: - **Hour of day** — the daily load shape: a morning ramp, an evening peak, an overnight trough. - **Day of week**, plus an explicit **is-weekend** flag — industrial and office load largely disappears on Saturday and Sunday. - **Month or quarter** — heating and cooling seasons; in many grids demand is U-shaped across the year. - **Public-holiday flag** — a holiday looks like a Sunday even when it falls on a Tuesday, and no combination of day-of-week and month can say so. - Occasionally **day of year** or **week of year** for finer seasonal shape. Which parts matter is a domain question, not a completeness exercise. For a model predicting *hourly* demand, minute and second are noise. Extracting every part the calendar offers just adds columns that can only fit noise, and each one is a column you must reproduce identically when the model scores new data. ## Why keep is-weekend when you already have day of week In an information sense the flag is redundant — it is a function of day of week. In a *learnable* sense it is not. If day of week enters as an integer 0-6, a tree needs two splits to carve Saturday and Sunday out of the middle-to-end of that range, and a linear model with a single coefficient on the integer cannot say "the last two are different" at all. A single binary column makes the distinction a one-split, one-coefficient fact. The same reasoning motivates a `is_month_start` or `is_business_day` flag when the domain cares. ## Local time, not UTC Systems log in UTC; humans consume electricity on a local clock. If you extract the hour from a UTC timestamp for a grid that sits several hours away, the 18:00 peak is smeared across whatever UTC hour it happens to fall in, and it *moves* twice a year when daylight saving starts and ends. Converting to the region's local time before extracting the hour keeps the peak at a stable value and absorbs the DST shift automatically. The cost is that one local hour occurs twice on the autumn changeover and one never occurs in spring — usually acceptable, and worth knowing about when you audit row counts. ## Holidays are a data dependency, not just a column A holiday flag imports an external calendar. Before shipping it, check: does the calendar cover every region the model serves; does it extend far enough into the future to score the dates you care about; and do you want the day itself only, or also the bridge days around it, which often behave more like a holiday than a workday. Moving feasts and regionally observed days make this messier than it first looks. A feature you cannot compute at scoring time is worse than no feature. ## Do not keep the raw timestamp as a numeric column This is the trap that survives into production. A tree learns thresholds inside the range it saw; every future timestamp falls beyond the largest threshold, into the same terminal region, so the model's answer for all of next year is one constant it learned from the tail of the training window. A linear model *can* extrapolate a trend on a time column, but it will extrapolate it forever, unbounded. If a long-run trend genuinely matters, express it deliberately — a days-since-start column in a model that can extrapolate — and make that a conscious choice, not an accident of leaving the timestamp in. ## A short checklist Convert to the right time zone; extract only the parts the domain justifies; add explicit flags for the distinctions the model would otherwise need several splits to find; confirm every extracted part is computable for future rows; drop the raw timestamp. Then check the columns actually help, rather than assuming they do.

  • Why not feed the raw epoch seconds to a gradient-boosted tree and let it find the pattern?
    A tree splits on thresholds, so it can only carve the training window into date ranges — nothing repeats, and it never learns that 18:00 is special. Worse, every future timestamp exceeds the largest threshold it learned, so all future rows land in one terminal region and receive the same prediction. The tree memorises the window instead of generalising beyond it.
  • Your logs are in UTC but the grid you model sits in one local zone — does it matter which you use?
    Yes. Consumption tracks the local clock, so the evening peak has a stable local hour but a UTC hour that shifts by an hour twice a year at the daylight-saving boundaries. Extracting the hour from UTC smears the peak across two values and forces the model to relearn it after every changeover. Convert first, then extract.
  • How do you handle the public-holiday flag for dates the model has to score next year?
    The flag depends on an external calendar, so it is only usable if that calendar extends past your scoring horizon and covers every region served. Check that first. Then decide whether the flag covers only the day itself or also the bridge days around it, which frequently behave more like a holiday than a normal workday.

A timestamp is a page number in an endless book; date parts are the tags that say which page is a Sunday, which is a holiday, and which is 6pm — those are what let you compare pages at all.

saying these in an interview costs you the question

  • Feeding the raw timestamp straight into the model
  • Extracting every available date part regardless of the domain
  • Extracting the hour in UTC when behaviour follows local time
  • Assuming a holiday calendar exists for future scoring dates
  • Believing a tree extrapolates a date trend past its training window

context

open as a page

How do you turn a raw event log into one row per customer for a churn model?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Choose the entity and a reference date, keep only that entity's events inside a window ending at the reference date, then collapse them into one row: counts, sums, means, maxima, distinct counts, and days since the last event.

open as a page

20% of survey respondents skipped the income question — do you drop those rows, drop the column, or impute?

level: juniorimportance: must knowfreq 82%

basics

~20 s

Impute in most cases. Deleting the rows discards a fifth of the data and has no equivalent at scoring time, where every request must return a prediction. Drop the column only if it is nearly empty or adds no lift.

open as a page

Why is coding education level as 0-3 acceptable but coding job family the same way risky?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Education level has a real order - high-school below bachelor below master below PhD - so integer codes carry that information. Job family has no order, so codes 0-3 invent a ranking the model takes literally. Encode unordered categories as one-hot indicators instead.

open as a page

What is the difference between z-score standardisation and min-max scaling?

level: juniorimportance: must knowfreq 84%

basics

~20 s

Z-score standardisation subtracts the mean and divides by the standard deviation, producing mean 0, standard deviation 1, and no fixed bounds. Min-max scaling rescales values into a fixed range such as 0 to 1 using the training minimum and maximum.

open as a page

How do equal-width and equal-frequency bins differ on a long-tailed usage column?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Equal-width bins cut the range into intervals of the same size, so a long tail crowds almost every row into the first bin and leaves near-empty bins above. Equal-frequency bins cut at quantiles, so counts are similar and widths vary.

open as a page

What is the difference between random oversampling and random undersampling of an imbalanced training set?

level: juniorimportance: must knowfreq 74%

basics

~10 s

Random oversampling duplicates existing minority rows until the classes are closer to balanced. Random undersampling deletes majority rows to reach the same ratio. Oversampling keeps every row but repeats information; undersampling throws information away.

open as a page

In a bag-of-words matrix of movie reviews, what does TF-IDF weighting do to words like 'the' and 'and'?

level: juniorimportance: must knowfreq 78%

basics

~20 s

TF-IDF multiplies each word's count by an inverse-document-frequency factor that shrinks as the word appears in more documents. Words like 'the' and 'and' sit in nearly every review, so their columns collapse to near-zero while rarer words keep large weights.

open as a page

What does an information value of 0.015 tell you about a scorecard feature?

level: juniorimportance: must knowfreq 58%

basics

~20 s

An information value of 0.015 sits below the conventional 0.02 floor, so on its own the feature barely separates defaulters from non-defaulters. Standard practice is to drop it from the scorecard unless policy requires it or it earns its place alongside other characteristics.

open as a page

What does a class weight do to a classifier's loss when the positive class is 2% of the rows?

level: middleimportance: must knowfreq 72%

basics

~20 s

A class weight multiplies every loss term from that class by a constant, so each rare positive counts like many rows. The optimiser trades more errors on the common class for fewer on the rare one, and fitted probabilities rise.

open as a page

Why hand-build a debt-to-income ratio when the model already has both raw columns?

level: middleimportance: must knowfreq 62%

basics

~20 s

A quotient is not a linear combination of its two columns, so no coefficients on debt and income reproduce debt divided by income. A tree can only approximate it with a staircase of axis-aligned splits, which costs depth and data.

open as a page

How do filter, wrapper and embedded feature selection differ in cost and in what they can see?

level: middleimportance: must knowfreq 68%

basics

~20 s

Filters rank columns with a cheap statistic and never train the model. Wrappers train it many times on candidate subsets and keep the best. Embedded selection falls out of one model fit. Cost and model-specificity rise across the three.

open as a page

Why does target-encoding a 40,000-level seller ID leak, and how does out-of-fold encoding fix it?

level: middleimportance: must knowfreq 72%

basics

~20 s

Computing each seller's target mean over all training rows puts a row's own label inside its own feature, so the model reads the answer back. Out-of-fold encoding builds each fold's means from the other folds only.

open as a page

Why does one extreme feature value distort an OLS fit and kNN neighbourhoods but barely move a decision tree?

level: middleimportance: must knowfreq 66%

basics

~20 s

Least squares squares residuals, so a far-out point dominates the total and rotates the line. kNN puts that huge gap into every distance, so neighbourhoods break. A tree splits on rank order, so the extreme lands in its own branch.

open as a page

Why do kNN and SVM need feature scaling while decision trees do not?

level: middleimportance: must knowfreq 76%

basics

~20 s

kNN and SVM compare examples by distance, so a feature measured in large units dominates every comparison. A tree splits one feature at a time at a threshold; rescaling preserves the ordering of values, so exactly the same splits remain available.

open as a page

Which skew transform handles a right-skewed column that contains zeros and negatives?

level: middleimportance: must knowfreq 58%

basics

~20 s

Plain log fails here: log(0) is undefined and negatives have no log. For a non-negative column with zeros, use log(1+x). When values drop below zero, use Yeo-Johnson, the power transform defined on the whole real line.

open as a page

How does SMOTE create a synthetic minority sample from existing training rows?

level: middleimportance: must knowfreq 78%

basics

~20 s

SMOTE picks a minority row, finds its k nearest neighbours among other minority rows, chooses one of them at random, and places a new point at a random fraction of the way along the straight line between the two.

open as a page

Why must a TF-IDF vocabulary and its idf values be fitted on the training fold only?

level: middleimportance: must knowfreq 55%

basics

~20 s

The vocabulary and the idf values are learned parameters. Building them over the whole corpus before splitting lets held-out documents shape the features that describe them, so validation scores come out optimistically biased and overstate what production will do.

open as a page

Why do credit scorecards replace binned features with their weight of evidence?

level: middleimportance: must knowfreq 68%

basics

~20 s

Weight of evidence replaces a bin with ln(share of non-defaults / share of defaults), which is the bin's log-odds offset from the population. That makes every feature monotone, on one common scale, linear in log-odds, auditable, and lets missing values form their own bin.

open as a page

How does label noise put a ceiling on the accuracy any model can reach?

level: seniorimportance: must knowfreq 55%

basics

~20 s

If a share of labels are wrong, even a model that predicts the true answer every time scores below 100% against those labels. Two radiologists disagreeing on one scan set a floor under any model's error.

open as a page

A $180,000 B2B order sits in a consumer marketplace's GMV column; do you drop, cap or keep it?

level: seniorimportance: must knowfreq 57%

basics

~20 s

Ask whether B2B orders will be scored in production. If so, keep the row and represent that segment explicitly; deleting it trains a model blind to your most valuable orders. Drop only corrupt or out-of-scope rows.

open as a page

What does count encoding do to a 30,000-level ZIP code column, and when does the count itself carry signal?

level: juniorimportance: should knowfreq 41%

basics

~20 s

Count encoding replaces each ZIP code with how many rows carry it, turning 30,000 levels into one numeric column. It helps when frequency proxies something real, like population density. It uses no labels, so it cannot leak.

open as a page

How do exact and near-duplicate rows distort a training set and its class balance?

level: juniorimportance: should knowfreq 50%

basics

~10 s

Duplicated rows count the same information several times. They upweight those examples in the loss, inflate one class's apparent frequency, and make the real sample size much smaller than the row count suggests.

open as a page

A pressure sensor column contains occasional -999 readings; why is capping them the wrong fix?

level: juniorimportance: should knowfreq 44%

basics

~20 s

-999 is a code the logger writes when the sensor is disconnected, not a measurement. Capping folds it into the real data range and teaches a relationship that never existed. Convert it to missing instead.

open as a page

Why treat the aggregation window length for entity features as a hyperparameter?

level: middleimportance: should knowfreq 50%

basics

~20 s

Window length trades freshness against stability, and the right value depends on how fast the behaviour changes. A 7-day window reacts quickly but is sparse and noisy; a 90-day window is stable but dilutes recent change. Tune it on validation data rather than guessing.

open as a page

Why can a correlation or mutual-information filter keep two 0.97-correlated columns yet drop a feature that only matters in combination?

level: middleimportance: should knowfreq 52%

basics

~20 s

A correlation or mutual-information filter scores each column against the target on its own. Two near-duplicates both score well and survive; a feature that matters only alongside another scores near zero alone and is cut. The blind spot is univariate scoring.

open as a page

Why add a binary missing-indicator column next to a feature you have imputed?

level: middleimportance: should knowfreq 58%

basics

~20 s

Because filling destroys the fact that the value was absent, and absence is often predictive in its own right. A binary flag restores that information, and it lets the model treat filled rows differently from genuinely observed ones.

open as a page

Why report Cohen's kappa instead of raw percent agreement between two annotators?

level: middleimportance: should knowfreq 45%

basics

~20 s

Percent agreement counts the agreements two annotators would hit by chance alone. Cohen's kappa removes that baseline: kappa = (observed agreement - chance agreement) / (1 - chance agreement), so 0 means chance-level and 1 means perfect.

open as a page

What does one-hot encoding a 50-level US state column do to a linear model's fit?

level: middleimportance: should knowfreq 56%

basics

~20 s

It adds fifty binary columns that are 98 percent zeros and gives every state its own weight. Each weight is driven only by that state's rows, so small states get unstable estimates and total variance rises. Penalising the weights becomes close to mandatory.

open as a page

Why must winsorising cap bounds be computed inside each training fold rather than on the full dataset?

level: middleimportance: should knowfreq 51%

basics

~20 s

Percentile caps are parameters estimated from data. Computing them over every row lets held-out rows shape their own preprocessing, so validation scores turn optimistic. Fit the bounds on training rows and apply those stored numbers unchanged everywhere else.

open as a page

showing 1–30 of 59