skip to content

Numeric Features and Scaling

Numeric columns rarely arrive model-ready: standardisation and robust scaling, log and Yeo-Johnson fixes for skew, capping extremes, and the ratio, aggregate and date-part columns you build yourself.

on this pageshow

explore

questions

21

For an electricity-demand model, which features would you extract from a raw timestamp column?

level: juniorimportance: must knowfreq 70%

answer

  1. one column hides several calendars
  2. the model cannot read a timestamp
  3. weekdays, weekends and holidays differ
  4. convert to local time before extracting hour

basics

~20 s

Split the timestamp into the calendar parts demand depends on: hour of day, day of week, an is-weekend flag, month or quarter, and a public-holiday flag. The raw timestamp alone only lets a model learn a trend.

solid answer

~50 s

A raw timestamp is one ever-increasing number, so no two rows share a value and the model can only read a trend out of it. I explode it into the calendar parts the domain actually turns on: hour of day for the daily load shape, day of week plus an is-weekend flag, month or quarter for the heating and cooling season, and a public-holiday flag, because a holiday behaves like a Sunday on a Tuesday. Two details decide whether this works live: extract the hour in the *local* time of the region modelled, not UTC, so the evening peak stays in one place across daylight-saving changes; and confirm the holiday calendar covers the dates you must score. I would not keep the raw timestamp as a numeric feature — a tree cannot extrapolate past the thresholds it learned, so every future row collapses into the last region it saw.

go deeper

for a junior

Be ready to name the parts out loud — hour, day of week, is-weekend, month, holiday flag — and say in one sentence why the raw timestamp on its own is useless to the model.

for a middle

Explain the mechanism: date parts create values that repeat across rows so the model can pool evidence, and an explicit flag saves a tree several splits over a raw integer encoding.

for a senior

Show the production judgment: local time versus UTC, daylight saving, whether the holiday calendar covers future scoring dates, and why leaving the raw timestamp in breaks a tree the moment it scores tomorrow.

for a principal

Own the tradeoff between a rich calendar feature set and the dependency it creates — an external holiday calendar per region is a contract someone must maintain and reproduce identically at scoring time.

## Why a timestamp is not yet a feature A timestamp is a single number that counts upward forever — seconds since an epoch, or an equivalent encoding. A model can only use a column through the operations its hypothesis class allows: a linear model multiplies the column by one coefficient, a tree compares it against thresholds. Neither operation can express "demand is high at 18:00 on weekdays and low on public holidays", because nothing in the raw number tells the model that two rows six months apart share an hour of day. Date-part extraction fixes exactly that. It converts one non-repeating column into several columns whose values **repeat across rows**, so the model can pool every 18:00 it has ever seen into one estimate. That pooling is the whole point — it is what turns 20,000 unique timestamps into a handful of learnable regularities. ## Which parts to extract For electricity demand the useful parts follow the physics and the human calendar: - **Hour of day** — the daily load shape: a morning ramp, an evening peak, an overnight trough. - **Day of week**, plus an explicit **is-weekend** flag — industrial and office load largely disappears on Saturday and Sunday. - **Month or quarter** — heating and cooling seasons; in many grids demand is U-shaped across the year. - **Public-holiday flag** — a holiday looks like a Sunday even when it falls on a Tuesday, and no combination of day-of-week and month can say so. - Occasionally **day of year** or **week of year** for finer seasonal shape. Which parts matter is a domain question, not a completeness exercise. For a model predicting *hourly* demand, minute and second are noise. Extracting every part the calendar offers just adds columns that can only fit noise, and each one is a column you must reproduce identically when the model scores new data. ## Why keep is-weekend when you already have day of week In an information sense the flag is redundant — it is a function of day of week. In a *learnable* sense it is not. If day of week enters as an integer 0-6, a tree needs two splits to carve Saturday and Sunday out of the middle-to-end of that range, and a linear model with a single coefficient on the integer cannot say "the last two are different" at all. A single binary column makes the distinction a one-split, one-coefficient fact. The same reasoning motivates a `is_month_start` or `is_business_day` flag when the domain cares. ## Local time, not UTC Systems log in UTC; humans consume electricity on a local clock. If you extract the hour from a UTC timestamp for a grid that sits several hours away, the 18:00 peak is smeared across whatever UTC hour it happens to fall in, and it *moves* twice a year when daylight saving starts and ends. Converting to the region's local time before extracting the hour keeps the peak at a stable value and absorbs the DST shift automatically. The cost is that one local hour occurs twice on the autumn changeover and one never occurs in spring — usually acceptable, and worth knowing about when you audit row counts. ## Holidays are a data dependency, not just a column A holiday flag imports an external calendar. Before shipping it, check: does the calendar cover every region the model serves; does it extend far enough into the future to score the dates you care about; and do you want the day itself only, or also the bridge days around it, which often behave more like a holiday than a workday. Moving feasts and regionally observed days make this messier than it first looks. A feature you cannot compute at scoring time is worse than no feature. ## Do not keep the raw timestamp as a numeric column This is the trap that survives into production. A tree learns thresholds inside the range it saw; every future timestamp falls beyond the largest threshold, into the same terminal region, so the model's answer for all of next year is one constant it learned from the tail of the training window. A linear model *can* extrapolate a trend on a time column, but it will extrapolate it forever, unbounded. If a long-run trend genuinely matters, express it deliberately — a days-since-start column in a model that can extrapolate — and make that a conscious choice, not an accident of leaving the timestamp in. ## A short checklist Convert to the right time zone; extract only the parts the domain justifies; add explicit flags for the distinctions the model would otherwise need several splits to find; confirm every extracted part is computable for future rows; drop the raw timestamp. Then check the columns actually help, rather than assuming they do.

  • Why not feed the raw epoch seconds to a gradient-boosted tree and let it find the pattern?
    A tree splits on thresholds, so it can only carve the training window into date ranges — nothing repeats, and it never learns that 18:00 is special. Worse, every future timestamp exceeds the largest threshold it learned, so all future rows land in one terminal region and receive the same prediction. The tree memorises the window instead of generalising beyond it.
  • Your logs are in UTC but the grid you model sits in one local zone — does it matter which you use?
    Yes. Consumption tracks the local clock, so the evening peak has a stable local hour but a UTC hour that shifts by an hour twice a year at the daylight-saving boundaries. Extracting the hour from UTC smears the peak across two values and forces the model to relearn it after every changeover. Convert first, then extract.
  • How do you handle the public-holiday flag for dates the model has to score next year?
    The flag depends on an external calendar, so it is only usable if that calendar extends past your scoring horizon and covers every region served. Check that first. Then decide whether the flag covers only the day itself or also the bridge days around it, which frequently behave more like a holiday than a normal workday.

A timestamp is a page number in an endless book; date parts are the tags that say which page is a Sunday, which is a holiday, and which is 6pm — those are what let you compare pages at all.

saying these in an interview costs you the question

  • Feeding the raw timestamp straight into the model
  • Extracting every available date part regardless of the domain
  • Extracting the hour in UTC when behaviour follows local time
  • Assuming a holiday calendar exists for future scoring dates
  • Believing a tree extrapolates a date trend past its training window

context

open as a page

How do you turn a raw event log into one row per customer for a churn model?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Choose the entity and a reference date, keep only that entity's events inside a window ending at the reference date, then collapse them into one row: counts, sums, means, maxima, distinct counts, and days since the last event.

open as a page

What is the difference between z-score standardisation and min-max scaling?

level: juniorimportance: must knowfreq 84%

basics

~20 s

Z-score standardisation subtracts the mean and divides by the standard deviation, producing mean 0, standard deviation 1, and no fixed bounds. Min-max scaling rescales values into a fixed range such as 0 to 1 using the training minimum and maximum.

open as a page

How do equal-width and equal-frequency bins differ on a long-tailed usage column?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Equal-width bins cut the range into intervals of the same size, so a long tail crowds almost every row into the first bin and leaves near-empty bins above. Equal-frequency bins cut at quantiles, so counts are similar and widths vary.

open as a page

Why hand-build a debt-to-income ratio when the model already has both raw columns?

level: middleimportance: must knowfreq 62%

basics

~20 s

A quotient is not a linear combination of its two columns, so no coefficients on debt and income reproduce debt divided by income. A tree can only approximate it with a staircase of axis-aligned splits, which costs depth and data.

open as a page

Why does one extreme feature value distort an OLS fit and kNN neighbourhoods but barely move a decision tree?

level: middleimportance: must knowfreq 66%

basics

~20 s

Least squares squares residuals, so a far-out point dominates the total and rotates the line. kNN puts that huge gap into every distance, so neighbourhoods break. A tree splits on rank order, so the extreme lands in its own branch.

open as a page

Why do kNN and SVM need feature scaling while decision trees do not?

level: middleimportance: must knowfreq 76%

basics

~20 s

kNN and SVM compare examples by distance, so a feature measured in large units dominates every comparison. A tree splits one feature at a time at a threshold; rescaling preserves the ordering of values, so exactly the same splits remain available.

open as a page

Which skew transform handles a right-skewed column that contains zeros and negatives?

level: middleimportance: must knowfreq 58%

basics

~20 s

Plain log fails here: log(0) is undefined and negatives have no log. For a non-negative column with zeros, use log(1+x). When values drop below zero, use Yeo-Johnson, the power transform defined on the whole real line.

open as a page

A $180,000 B2B order sits in a consumer marketplace's GMV column; do you drop, cap or keep it?

level: seniorimportance: must knowfreq 57%

basics

~20 s

Ask whether B2B orders will be scored in production. If so, keep the row and represent that segment explicitly; deleting it trains a model blind to your most valuable orders. Drop only corrupt or out-of-scope rows.

open as a page

A pressure sensor column contains occasional -999 readings; why is capping them the wrong fix?

level: juniorimportance: should knowfreq 44%

basics

~20 s

-999 is a code the logger writes when the sensor is disconnected, not a measurement. Capping folds it into the real data range and teaches a relationship that never existed. Convert it to missing instead.

open as a page

Why treat the aggregation window length for entity features as a hyperparameter?

level: middleimportance: should knowfreq 50%

basics

~20 s

Window length trades freshness against stability, and the right value depends on how fast the behaviour changes. A 7-day window reacts quickly but is sparse and noisy; a 90-day window is stable but dilutes recent change. Tune it on validation data rather than guessing.

open as a page

Why must winsorising cap bounds be computed inside each training fold rather than on the full dataset?

level: middleimportance: should knowfreq 51%

basics

~20 s

Percentile caps are parameters estimated from data. Computing them over every row lets held-out rows shape their own preprocessing, so validation scores turn optimistic. Fit the bounds on training rows and apply those stored numbers unchanged everywhere else.

open as a page

When would you use median/IQR robust scaling instead of z-score standardisation?

level: middleimportance: should knowfreq 46%

basics

~20 s

Use it when a column carries extreme values. Robust scaling subtracts the median and divides by the interquartile range, statistics that a handful of extreme points barely move, so the bulk of the data still lands on a usable scale.

open as a page

You added 200 derived columns to a 1,500-row training table — how do you decide which earn their place?

level: seniorimportance: should knowfreq 44%

basics

~20 s

At 200 columns and 1,500 rows there are roughly seven rows per column, so some will look predictive by chance alone. Judge them by repeated cross-validation with the selection step refitted inside every fold, against a baseline without them.

open as a page

An account has no events in the 90-day window — which of its aggregates are zero and which are missing?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Counts and sums are genuinely zero — nothing happened, and that is an observation. Means, maxima and shares are undefined because there is nothing to summarise, and recency has no value at all. Encode those with a sentinel plus a companion flag, never a silent zero.

open as a page

Why does exponentiating a log-price model's prediction land on the median, not the mean?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Exponentiating is a convex map, so the average of log-scale predictions does not map back to the average price. Under symmetric log-scale errors it returns the conditional median; with normal log errors of standard deviation s, multiply by exp(s^2 / 2) to recover the mean.

open as a page

When is discretising a continuous predictor into bins worth the information it destroys?

level: principalimportance: should knowfreq 36%

basics

~20 s

Rarely for accuracy, often for everything else. Binning discards within-bin variation and imposes arbitrary cutpoints, so it usually costs predictive power. It earns its place when bins buy interpretability, stable reporting, or let a linear model express a non-monotone effect.

open as a page

Why encode a pickup hour as a sine-cosine pair instead of the integer 0-23?

level: middleimportance: nice to knowfreq 40%

basics

~20 s

Hour 23 and hour 0 are one hour apart but 23 units apart as integers. Mapping the hour onto a circle with sin(2pih/24) and cos(2pih/24) makes them neighbours again, which matters to any distance- or magnitude-based model.

open as a page

Why add share-of-total and per-day rate features next to raw per-entity event counts?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

A raw count mixes how active an entity is overall with how it splits that activity and how long it was observed. Dividing by the entity's total gives composition, and dividing by days observed gives intensity, so the model can separate a heavy user from a lopsided one.

open as a page

Why can min-max scaling still leave one feature dominating a distance-based model?

level: seniorimportance: nice to knowfreq 31%

basics

~20 s

Min-max equalises each feature's range, not its spread. A heavy-tailed column whose maximum sits far above the bulk gets compressed near zero after scaling, so a well-spread bounded column ends up supplying almost all of the distance.

open as a page

What does rank-normalising a heavy-tailed page-load-time feature gain and cost?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Rank normalisation replaces each value with its rank, rescaled to a uniform or Gaussian shape. It flattens any tail with no parametric assumption, but it keeps only the ordering: how far apart two values were is discarded.

open as a page