skip to content

Why add share-of-total and per-day rate features next to raw per-entity event counts?

level: middleimportance: nice to knowfreq 32%

answer

  1. the count bundles size and mix
  2. divide by the entity's own total
  3. unequal observation windows across entities
  4. keep the level, add the ratio
  5. watch the zero denominator

basics

~20 s

A raw count mixes how active an entity is overall with how it splits that activity and how long it was observed. Dividing by the entity's total gives composition, and dividing by days observed gives intensity, so the model can separate a heavy user from a lopsided one.

solid answer

~60 s

Raw aggregates confound three different things: overall size, composition and exposure. A player who spent 60 on one game title looks the same as another who spent 60 — until you know the first spent 200 across all titles and the second spent 60 on that one title alone. Share-of-total, `60 / 200 = 0.30`, is the composition feature; it says where the entity's activity concentrates, independent of scale. A per-day rate does the same job for exposure: an account observed for 12 days since signup and an account observed for the full 90 cannot be compared on raw counts, but events per observed day can. Distinct counts add the third axis, breadth. I keep the raw level alongside the normalised version rather than replacing it, because scale itself is predictive — a whale really is different from a casual user. And I guard the denominators: a share is undefined when the total is zero, so that case needs an explicit value plus a flag rather than a silent null.

go deeper

for a junior

Recall that a count on its own can mean 'very active' or 'active for a long time', and that dividing by a total or by days observed gives a number you can compare across entities.

for a middle

Explain the three axes concretely: share-of-total for composition, per-day rate for exposure, distinct count for breadth, and why the raw level stays in the table too.

for a senior

Show the operational care — zero and near-zero denominators handled explicitly, sparse entities shrunk toward the population value or accompanied by their event count, extreme rates capped by an agreed rule.

for a principal

Take a position on how far normalisation should go as a house style, and on the readability cost: a table full of unexplained ratios is expensive for analysts and for the business partners who must trust the model's drivers.

## Three things hidden in one number An aggregate like "60 units of spend on title X in the last 30 days" bundles together several facts that a model would benefit from seeing separately: 1. **Scale** — how much this entity does in total. 2. **Composition** — what fraction of its activity went to this particular thing. 3. **Exposure** — how long the entity was actually observable inside the window. A raw count or sum reports the product of these, and the model has to disentangle them from the data. Sometimes it can, if you also supply the totals and the exposure; often it cannot, particularly for linear models, which would need an explicit interaction to express a ratio at all. ## Share-of-total: the composition axis Divide an entity's aggregate over one slice by the same aggregate over all slices. A player's spend on one title as a fraction of their total spend that month is the canonical example: `share = 60 / 200 = 0.30`. Two players with identical 60-unit spend on that title now separate cleanly — one is a generalist who happens to play it, the other is devoted to it. Share-of-total features are naturally bounded in `[0, 1]`, which makes them comparable across entities of wildly different sizes and stable when the population's overall activity drifts. They are the right feature whenever the question is "what kind of entity is this?" rather than "how much does it do?" ## Per-day rate: the exposure axis A window is a fixed calendar span, but entities are not observable for all of it. An account created 12 days before the reference date could only ever generate 12 days of events, so its 90-day count is structurally small. Comparing it to a three-year-old account's count is comparing a partly filled window with a full one. The fix is to divide by exposure: `events per observed day = count / min(window_length, days_since_signup)`. This turns the count into an intensity that is comparable across entities with different histories. Carrying tenure as its own feature is the complement — it tells the model *why* the exposure was short, which is often predictive in its own right. ## Distinct counts: the breadth axis The third normalisation is not a division at all. Distinct counts — how many merchant categories a cardholder touched in the last 90 days, how many product features a SaaS account used — measure variety rather than volume. A distinct count is bounded by the size of the category set rather than by activity level, so it is naturally less dominated by the heaviest users than a raw count is. ## Keep the level as well as the ratio A common overcorrection is to replace the raw aggregates with normalised ones. Don't. Scale itself carries signal: a customer who spends 20,000 a month is a different commercial proposition from one who spends 20, even if their composition is identical. Ratios also destroy information about certainty — a 100% share computed from one event and a 100% share computed from four hundred events are the same number with very different reliability, and the raw count is what lets the model tell them apart. Supply both and let the model choose. ## Guard the denominators Every normalised feature has a denominator that can be zero or tiny: - **Zero total** — an entity with no activity at all has no meaningful share. Give the column an explicit sentinel value and add a companion indicator flag rather than letting a division by zero become a silent missing value. - **Tiny denominators** — one event in the window produces shares of exactly 0 or 1 and rates that swing enormously. Some teams shrink the ratio toward the population value in proportion to the entity's event count, so sparse entities are pulled toward the average and dense ones keep their own value; at minimum, carry the event count so the model can learn the discount itself. - **Clipping** — a rate computed over a one-day exposure can be an extreme outlier that dominates a scale-sensitive model, so decide up front whether such rows are capped, excluded, or kept. ## Reading the result Well-named normalised columns explain themselves: `spend_share_title_x_30d` and `events_per_active_day_30d` say what was divided by what. That matters more here than for raw aggregates, because a ratio's meaning lives entirely in its denominator, and a column called `spend_ratio` is unreadable six months later.

  • Why keep the raw count once you have the share-of-total?
    Because scale is predictive on its own, and because the ratio hides certainty. A 100% share from one event and a 100% share from four hundred are the same number with very different reliability; the raw count is what lets the model distinguish them. Supply both and let it decide.
  • How do you handle a share-of-total when the entity's total is zero?
    The share genuinely does not exist, so I give the column an explicit sentinel and add a companion flag marking the entity as having no activity at all. Silently writing zero conflates 'no activity' with 'activity, none of it here', which are different states and often have opposite implications.
  • What does dividing an event count by days observed protect against?
    Unequal exposure. An account created twelve days before the reference date could only generate twelve days of events, so its 90-day count is structurally small for a reason that has nothing to do with engagement. Events per observed day makes it comparable, and tenure alongside it tells the model why the exposure was short.

saying these in an interview costs you the question

  • Replaces every raw aggregate with a ratio
  • Compares raw counts across entities with different tenures
  • Divides without checking for a zero denominator
  • Trusts a share of 1.0 computed from a single event
  • Names the column ratio with no denominator in the name

context