For an electricity-demand model, which features would you extract from a raw timestamp column?
answer
- one column hides several calendars
- the model cannot read a timestamp
- weekdays, weekends and holidays differ
- convert to local time before extracting hour
basics
~20 sSplit the timestamp into the calendar parts demand depends on: hour of day, day of week, an is-weekend flag, month or quarter, and a public-holiday flag. The raw timestamp alone only lets a model learn a trend.
solid answer
~50 sA raw timestamp is one ever-increasing number, so no two rows share a value and the model can only read a trend out of it. I explode it into the calendar parts the domain actually turns on: hour of day for the daily load shape, day of week plus an is-weekend flag, month or quarter for the heating and cooling season, and a public-holiday flag, because a holiday behaves like a Sunday on a Tuesday. Two details decide whether this works live: extract the hour in the *local* time of the region modelled, not UTC, so the evening peak stays in one place across daylight-saving changes; and confirm the holiday calendar covers the dates you must score. I would not keep the raw timestamp as a numeric feature — a tree cannot extrapolate past the thresholds it learned, so every future row collapses into the last region it saw.
go deeper
Be ready to name the parts out loud — hour, day of week, is-weekend, month, holiday flag — and say in one sentence why the raw timestamp on its own is useless to the model.
Explain the mechanism: date parts create values that repeat across rows so the model can pool evidence, and an explicit flag saves a tree several splits over a raw integer encoding.
Show the production judgment: local time versus UTC, daylight saving, whether the holiday calendar covers future scoring dates, and why leaving the raw timestamp in breaks a tree the moment it scores tomorrow.
Own the tradeoff between a rich calendar feature set and the dependency it creates — an external holiday calendar per region is a contract someone must maintain and reproduce identically at scoring time.
## Why a timestamp is not yet a feature A timestamp is a single number that counts upward forever — seconds since an epoch, or an equivalent encoding. A model can only use a column through the operations its hypothesis class allows: a linear model multiplies the column by one coefficient, a tree compares it against thresholds. Neither operation can express "demand is high at 18:00 on weekdays and low on public holidays", because nothing in the raw number tells the model that two rows six months apart share an hour of day. Date-part extraction fixes exactly that. It converts one non-repeating column into several columns whose values **repeat across rows**, so the model can pool every 18:00 it has ever seen into one estimate. That pooling is the whole point — it is what turns 20,000 unique timestamps into a handful of learnable regularities. ## Which parts to extract For electricity demand the useful parts follow the physics and the human calendar: - **Hour of day** — the daily load shape: a morning ramp, an evening peak, an overnight trough. - **Day of week**, plus an explicit **is-weekend** flag — industrial and office load largely disappears on Saturday and Sunday. - **Month or quarter** — heating and cooling seasons; in many grids demand is U-shaped across the year. - **Public-holiday flag** — a holiday looks like a Sunday even when it falls on a Tuesday, and no combination of day-of-week and month can say so. - Occasionally **day of year** or **week of year** for finer seasonal shape. Which parts matter is a domain question, not a completeness exercise. For a model predicting *hourly* demand, minute and second are noise. Extracting every part the calendar offers just adds columns that can only fit noise, and each one is a column you must reproduce identically when the model scores new data. ## Why keep is-weekend when you already have day of week In an information sense the flag is redundant — it is a function of day of week. In a *learnable* sense it is not. If day of week enters as an integer 0-6, a tree needs two splits to carve Saturday and Sunday out of the middle-to-end of that range, and a linear model with a single coefficient on the integer cannot say "the last two are different" at all. A single binary column makes the distinction a one-split, one-coefficient fact. The same reasoning motivates a `is_month_start` or `is_business_day` flag when the domain cares. ## Local time, not UTC Systems log in UTC; humans consume electricity on a local clock. If you extract the hour from a UTC timestamp for a grid that sits several hours away, the 18:00 peak is smeared across whatever UTC hour it happens to fall in, and it *moves* twice a year when daylight saving starts and ends. Converting to the region's local time before extracting the hour keeps the peak at a stable value and absorbs the DST shift automatically. The cost is that one local hour occurs twice on the autumn changeover and one never occurs in spring — usually acceptable, and worth knowing about when you audit row counts. ## Holidays are a data dependency, not just a column A holiday flag imports an external calendar. Before shipping it, check: does the calendar cover every region the model serves; does it extend far enough into the future to score the dates you care about; and do you want the day itself only, or also the bridge days around it, which often behave more like a holiday than a workday. Moving feasts and regionally observed days make this messier than it first looks. A feature you cannot compute at scoring time is worse than no feature. ## Do not keep the raw timestamp as a numeric column This is the trap that survives into production. A tree learns thresholds inside the range it saw; every future timestamp falls beyond the largest threshold, into the same terminal region, so the model's answer for all of next year is one constant it learned from the tail of the training window. A linear model *can* extrapolate a trend on a time column, but it will extrapolate it forever, unbounded. If a long-run trend genuinely matters, express it deliberately — a days-since-start column in a model that can extrapolate — and make that a conscious choice, not an accident of leaving the timestamp in. ## A short checklist Convert to the right time zone; extract only the parts the domain justifies; add explicit flags for the distinctions the model would otherwise need several splits to find; confirm every extracted part is computable for future rows; drop the raw timestamp. Then check the columns actually help, rather than assuming they do.
- Why not feed the raw epoch seconds to a gradient-boosted tree and let it find the pattern?A tree splits on thresholds, so it can only carve the training window into date ranges — nothing repeats, and it never learns that 18:00 is special. Worse, every future timestamp exceeds the largest threshold it learned, so all future rows land in one terminal region and receive the same prediction. The tree memorises the window instead of generalising beyond it.
- Your logs are in UTC but the grid you model sits in one local zone — does it matter which you use?Yes. Consumption tracks the local clock, so the evening peak has a stable local hour but a UTC hour that shifts by an hour twice a year at the daylight-saving boundaries. Extracting the hour from UTC smears the peak across two values and forces the model to relearn it after every changeover. Convert first, then extract.
- How do you handle the public-holiday flag for dates the model has to score next year?The flag depends on an external calendar, so it is only usable if that calendar extends past your scoring horizon and covers every region served. Check that first. Then decide whether the flag covers only the day itself or also the bridge days around it, which frequently behave more like a holiday than a normal workday.
A timestamp is a page number in an endless book; date parts are the tags that say which page is a Sunday, which is a holiday, and which is 6pm — those are what let you compare pages at all.
saying these in an interview costs you the question
- Feeding the raw timestamp straight into the model
- Extracting every available date part regardless of the domain
- Extracting the hour in UTC when behaviour follows local time
- Assuming a holiday calendar exists for future scoring dates
- Believing a tree extrapolates a date trend past its training window