Why treat the aggregation window length for entity features as a hyperparameter?
answer
- freshness against stability
- short window, few events per entity
- long window dilutes recent change
- tune it like any other hyperparameter
- several windows plus their ratio
basics
~20 sWindow length trades freshness against stability, and the right value depends on how fast the behaviour changes. A 7-day window reacts quickly but is sparse and noisy; a 90-day window is stable but dilutes recent change. Tune it on validation data rather than guessing.
solid answer
~60 sThe window is a modelling choice with a real bias-variance flavour, so I treat it like any other hyperparameter. A short window — 7 days for a mobile-game reactivation model — reflects the player's current state, but many players have only a handful of events in it, so counts are noisy and a large share of rows are empty. A 90-day window gives every entity enough events for a stable estimate, but a player who stopped last week still looks busy because the older activity dominates. Which is better is an empirical question about how fast the behaviour turns over, so I evaluate candidate windows under the same validation protocol and compare. In practice the strongest answer is usually not one window but several — 7, 30 and 90 day versions of the same aggregate — plus their ratios, because `count_7d / count_30d` tells the model whether the entity is accelerating or decaying. The cost is a wider feature table and more compute, so I stop adding windows when validation stops improving.
go deeper
Know that the look-back period is a choice you make, not a property of the data, and be able to say that short windows are fresher while long windows are steadier.
Explain both failure directions concretely: too short means few events per entity and noisy aggregates, too long means a recent behaviour change is drowned by older activity.
Demonstrate the tuning discipline — fix the entity set before varying the window, evaluate on held-out data, and reach for multiple windows plus their ratios when one horizon cannot serve both freshness and stability.
Weigh the cost side: every extra window multiplies the feature table, the recompute bill and the explanation burden. Decide where the team's default sits and when a simpler single-window set is the better product decision.
## The window is a knob, not a convention When you summarise an entity's events into one row, you choose how far back to look. Thirty days is the reflex answer, but nothing about the data makes 30 correct; it is a hyperparameter in exactly the sense that a tree depth or a regularisation strength is, and it deserves the same treatment — a small grid, evaluated under the validation protocol you already trust, with the choice recorded. ## What shortening the window buys and costs **Buys: freshness.** The whole point of an activity feature is to describe the entity's state now. A 7-day window is dominated by what just happened, so a player who quit last Tuesday shows a collapsing session count immediately. **Costs: sample size per entity.** Each entity's aggregate is an estimate computed from however many events fell in the window. Seven days might contain two sessions, so a mean session length computed from them swings wildly, a maximum is essentially one draw from the tail, and a distinct count is bounded by how few chances the entity had. Worse, a large fraction of entities have *no* events in a short window, so the feature is constant across a big block of rows and stops discriminating within it. ## What lengthening the window buys and costs **Buys: stability and coverage.** A 90-day window averages over more events, so aggregates are less noisy, more entities are non-empty, and rare-but-informative event types actually appear. **Costs: staleness and dilution.** A long window mixes the entity's past with its present. If a shopper ordered weekly for two months and then stopped three weeks ago, a 90-day order count still looks healthy — the very change you want to detect is diluted by the history that precedes it. Long windows also require long history: an account that is 20 days old cannot have a genuine 90-day aggregate, so either its window is truncated (making its counts incomparable to a mature account's) or the account is excluded. ## The tuning Treat candidate windows as a grid — say 7, 14, 30, 90, all-history — and compare validation performance with everything else held fixed. Two cautions make the comparison honest: - **Compare like with like.** Changing the window changes how much history each row needs, which can silently change *which entities* survive into the dataset. Fix the entity set first, then vary the window, or you are comparing two different problems. - **Do not tune on the data you report on.** The window is chosen from validation performance, so its benefit is already optimised-for; the held-out estimate must come from data that took no part in the choice. ## The usual best answer: several windows at once A single window forces a compromise between freshness and stability. Computing the same aggregate over multiple windows sidesteps it: give the model `sessions_7d`, `sessions_30d` and `sessions_90d` and let it decide which horizon matters. The real payoff is in the derived comparisons: - **Ratio** — `count_7d / count_30d`. For a steady entity this sits near 7/30; well above means accelerating, well below means decaying. This trend signal is invisible to any single window. - **Difference of rates** — events per day in the last week minus events per day over the quarter, which has the same interpretation on an additive scale. Guard the denominators: a ratio is undefined when the long-window count is zero, so define that case explicitly rather than letting an infinity or a silent null propagate. ## Where to stop Each extra window multiplies the feature count by the size of the aggregate vocabulary, and every column costs compute, storage and a reviewer's attention. The discipline is to add windows while validation improves and stop when it flattens — and to remember that a model with three windows of every feature is harder to explain to the business than one with a well-chosen single window and a couple of trend ratios. ## How the window interacts with the label One structural rule holds regardless of length: the window ends at the reference date and looks only backwards. Extending it forwards to gather more events is not a longer window, it is a different and much more serious problem. Lengthening a window always means reaching further into the past.
- What does a 7-day over 30-day count ratio tell the model that neither count does alone?Trend. For an entity with steady activity the ratio sits near 7/30; a value well above that means the entity is accelerating, well below means it is going quiet. Both raw counts can look ordinary while the ratio is extreme, which is exactly the early-warning case a churn or reactivation model wants.
- How does a short window interact with entities that have very few events?Badly. The aggregate is an estimate from whatever fell in the window, so with two or three events the mean and maximum are dominated by chance and the distinct count is capped by opportunity. Carry the underlying event count as its own feature so the model can learn to discount the unstable summaries.
- Why can't you compare window lengths by whichever gives the best training score?Because a longer window adds information the model can memorise; training score tends to improve with more columns regardless of whether the extra history generalises. The comparison has to be on held-out data, and the final reported estimate should come from data that took no part in choosing the window.
saying these in an interview costs you the question
- Picks 30 days because it is the usual number
- Claims a longer window is always more informative
- Compares windows on training performance
- Never notices that a short window leaves many entities empty
- Adds ten windows without checking validation gain