skip to content

An account has no events in the 90-day window — which of its aggregates are zero and which are missing?

level: seniorimportance: should knowfreq 38%

answer

  1. blank hides three different states
  2. nothing happened is an observation
  3. the mean of an empty set
  4. never active is not very stale
  5. sentinel needs a companion flag

basics

~20 s

Counts and sums are genuinely zero — nothing happened, and that is an observation. Means, maxima and shares are undefined because there is nothing to summarise, and recency has no value at all. Encode those with a sentinel plus a companion flag, never a silent zero.

solid answer

~60 s

I separate three states that all look like 'blank'. **Structurally zero**: the account existed and was observable, and did nothing — its event count and spend sum really are 0, and writing 0 is the honest value. **Undefined**: mean order value, maximum session length and share-of-total have no value over an empty set; writing 0 there tells the model 'this account's typical order is worth nothing', which is a different and false claim. **Not yet observable**: an account created after the window opened has a short exposure, so even its counts are not comparable to a mature account's; that needs a tenure column and a rate, not a zero. For the undefined group I use an explicit sentinel plus a boolean flag — a tree can split on the sentinel, but a linear model reads it as a number on the same scale, so the flag is what makes it safe. Recency is the sharp case: 'never active' is not a large number of days, it is a different category.

go deeper

for a junior

Know that an empty window is not one situation: a count of zero is a real measurement, while an average with nothing to average is simply undefined.

for a middle

Explain why writing zero into a mean, a maximum or a recency column makes a false claim, and describe the sentinel-plus-flag and missing-plus-indicator encodings.

for a senior

Show the diagnosis instinct — separate structurally zero, undefined and not-yet-observable rows, choose an encoding that suits the model family, and make sure training and scoring apply the same convention.

for a principal

Own the convention itself. Decide the house rule for empty and thin aggregates, get it into the data dictionary, and make it reviewable, because inconsistent blank handling across teams is a defect that no single model review catches.

## Three states that all render as blank When you left-join aggregates onto an entity roster, some rows come back with nothing. Treating them all the same is one of the most common quiet defects in an entity feature table, because three genuinely different situations produce the same empty result. **1. Structurally zero.** The entity existed, was observable for the whole window, and simply did nothing. Its event count is 0, its spend sum is 0, its distinct-category count is 0. These are not missing values — they are measurements, and they are frequently the most informative rows in a churn model. Writing 0 is correct, and any imputation applied here destroys real signal. **2. Undefined.** Some aggregates have no value over an empty set. The **mean** order value of zero orders is not zero; it is the average of nothing. The **maximum** session length of zero sessions does not exist. A **share-of-total** with a zero denominator is undefined. **Recency** — days since the last event — has no value when there is no last event. Writing 0 into these columns makes an affirmative false statement: zero mean order value says the account's typical order is worthless, and zero recency says the account was active today, which is the exact opposite of the truth. **3. Not yet observable.** An account created 20 days before the reference date has only 20 days of possible activity inside a 90-day window. Its counts are small for a structural reason, and comparing them to a three-year-old account's counts is comparing a partly filled window with a full one. This needs a tenure column and a per-observed-day rate so the model can tell short exposure from genuine quiet. ## How to encode the undefined group There are three workable approaches, and the right one depends on the model family. **Sentinel plus flag.** Fill the undefined column with an out-of-range constant — recency of -1, or a value larger than any real recency — and add a boolean `never_active` column beside it. A tree can isolate the sentinel with a single split, so this works well for tree ensembles. A linear model, however, reads the sentinel as a number on the same scale and will happily interpolate through it, which is exactly why the flag matters: the flag gives the linear model a clean indicator to attach its own coefficient to. **Explicit missing plus indicator.** Leave the value genuinely missing, let whatever imputation the pipeline applies fill it, and carry an indicator column recording that it was imputed. This keeps the honest statement ("no value") and still lets the model learn that missingness itself predicts the outcome — which, for activity features, it usually does. **Decompose the feature.** Sometimes the cleanest answer is two columns that are each always defined: `has_any_activity` (0/1) and, for the active subset, the summary. This makes the two-population structure explicit instead of hiding it in a sentinel. What is never acceptable is silently writing 0 into an undefined column and then forgetting, because it creates a group of rows whose feature value is a lie, and the model will learn from it. ## Recency deserves its own paragraph Recency is the most valuable column in most activity feature sets and the most frequently mis-encoded. "Never active" is not "active a very long time ago" — an account that has never ordered and an account that ordered once two years ago behave differently, and a large-number sentinel merges them. If you pick a sentinel, pick one clearly outside the observed range and always ship the flag with it. ## Thin windows, not just empty ones The same reasoning applies one step up. An entity with two events in the window has a *defined* mean and maximum, but they are estimates from two draws: the mean swings, the maximum is essentially a single sample from the tail, and the distinct count is capped at two by opportunity rather than by preference. Always carry the underlying event count alongside the summaries so the model can learn to discount unreliable ones. Where the ratio matters more than the level, shrinking an entity's value toward the population value in proportion to how few events it has is the standard remedy. ## A checklist for the empty row - Counts, sums and distinct counts: **0**, and say so in the data dictionary. - Means, medians, maxima, minima, shares: **undefined** — sentinel plus flag, or missing plus indicator. - Recency and tenure: recency undefined for never-active entities; tenure is defined if you know a signup date, and it is the column that separates "new" from "gone quiet". - Short-tenure entities: keep them, but add exposure-normalised versions of the counts so the comparison is fair. - Whichever convention you choose, apply it identically when the model is scored later; a feature that means one thing in training and another at scoring time is worse than a missing feature.

  • Why is a large-number sentinel for recency risky in a linear model but tolerable in a tree?
    A tree isolates the sentinel with one split, so the never-active group gets its own leaf. A linear model multiplies the sentinel by a coefficient and interpolates through it, so a value of 9999 drags the fitted line and makes the coefficient meaningless. Ship a boolean flag alongside so the linear model has an indicator to use instead.
  • How do you tell a brand-new account apart from one that has gone quiet?
    Tenure. Both have low counts and possibly no recent events, but the new account's window is barely open while the quiet one has a long history and a large recency. Carrying tenure, plus counts normalised by days observed, lets the model separate the two without you hand-coding a rule.
  • An entity has two events in the window. Is its maximum event value trustworthy?
    Not on its own. A maximum over two draws is close to a single sample from the tail, so it varies enormously between otherwise similar entities. Keep the event count as a feature so the model can discount thin summaries, and consider shrinking ratio-style features toward the population value in proportion to how few events back them.

A restaurant that served no customers last month had zero covers — a real number. But its average bill for the month does not exist; writing zero there claims every diner spent nothing, which is a different story entirely.

saying these in an interview costs you the question

  • Fills every empty aggregate with zero and moves on
  • Encodes never-active as recency zero
  • Imputes the population mean over genuine zero counts
  • Uses a 9999 sentinel with no companion flag
  • Compares raw counts of a new account and a mature one
  • Trusts a maximum computed from two events

context