skip to content

What changes when an active-user metric uses a 28-day activity window instead of 7?

level: middleimportance: should knowfreq 41%

answer

  1. how much evidence of life counts
  2. smoother versus faster
  3. one action keeps you counted for weeks
  4. four weeks is a whole number of weeks
  5. the window must fit inside the test

basics

~20 s

A 28-day window counts anyone with a single qualifying action in the past four weeks, so it reads higher, moves slowly and hides recent declines. A 7-day window reacts faster, is noisier, and is far more exposed to day-of-week effects.

solid answer

~50 s

The window length decides how much evidence of life a user needs to count, so it trades responsiveness against stability. A 28-day window has a large qualifying period: the metric is higher, smoother, and can take weeks to reflect a drop, because a user who churned yesterday keeps counting for 27 more days. A 7-day window turns over quickly and surfaces changes fast, but it is noisier and inherits weekly seasonality directly. In an experiment the window also has to fit: to attribute a 28-day active read to a treatment, every counted user needs a full 28 days of post-exposure observation, otherwise part of each user's window predates exposure and the measured effect is diluted toward zero. That usually means either running long enough to cover a whole window after the last enrolment, or choosing a shorter window and saying so up front. Pin the window in the definition before launch; it is not a knob to tune after seeing results.

go deeper

for a junior

Know that the window length is part of the metric, that a longer window always yields a bigger number, and that figures from different windows are not comparable to each other.

for a middle

Explain the tradeoff mechanically: the window is a memory that delays declines, whole-week lengths hold weekday mix constant, and a low bar over four weeks saturates and loses sensitivity.

for a senior

Show you check that the window fits the measurement period — a full post-exposure window per user — and that you can spot dilution from pre-exposure days making a null result meaningless.

for a principal

Own the choice of the company's canonical window and the rule that it is not re-tuned per experiment. Be ready to justify running both a fast and a stable definition rather than letting teams pick a length after seeing their numbers.

## What an activity window is doing An activity-window metric is a binary per-user indicator with a lookback attached: a user counts as active on a given day if they performed at least one qualifying action at some point in the preceding N days. Three choices define it completely — the qualifying action, the window length N, and the reference date the window looks back from. Changing any one of them changes the number, so all three belong in the metric definition. ## Longer versus shorter **Level.** A longer window can only include more users, so a 28-day count is always at least as large as a 7-day count on the same action and the same date. The two are not different measurements of the same quantity; they are different quantities, and comparing a 28-day figure against a 7-day figure is a category error. **Responsiveness and memory.** The window is a memory. With N = 28, a user who stops today continues to count for the next 27 days, so a genuine collapse in activity decays out of the metric gradually rather than appearing at once. That smoothing is useful for reporting a stable trend and dangerous for detecting a regression quickly. With N = 7, a stop shows up within a week, which is what you want for an incident or a fast experiment, at the cost of a much bouncier series. **Seasonality.** Weekly rhythm is the dominant cycle in most consumer and business products. A window that is a whole number of weeks contains the same number of each weekday, so weekday composition stays constant as the window slides. This is the practical reason 28 days is preferred over 30: 28 is exactly four weeks, whereas a 30- or 31-day window keeps changing which weekdays it double-counts, producing a sawtooth that is easy to misread as a trend. A 7-day window shares the whole-week property but has far less averaging behind it. **Sensitivity to a change.** A long window sets a low bar — one action in four weeks — so it saturates. Many products have most of their base clearing that bar, which leaves little headroom for a treatment to move the number and makes the metric insensitive as an experiment readout. A shorter window sets a higher bar and typically discriminates better between arms, which is why experiment readouts often use tighter windows than executive reporting does. ## The experiment-duration problem This is the part candidates most often miss. To claim a treatment caused a change in a 28-day active metric, the 28 days being examined must all lie after exposure. If the experiment runs for 14 days and you compute a 28-day window at the end, roughly half of each user's window predates their exposure. Activity in that pre-exposure half is identical in expectation across arms, so it adds signal-free mass to both, and any real effect is diluted toward zero. A null result under those conditions tells you almost nothing. The options are: - Run long enough that every enrolled user has a full window observed after exposure — for a 28-day window, that is 28 days after the last enrolment, not after the first. - Use a shorter window that fits the test, and state that the readout is a 7-day or 14-day activity metric rather than the reporting metric the company usually quotes. - Read the window-based metric as a directional secondary and decide on a metric that is defined within the test's horizon. A related trap is left-censoring at the start of a product's life or a new cohort's life: users who have not existed for 28 days cannot have a full window, and including them mechanically depresses the metric. ## Getting the definition right Write down: the qualifying action (opened the app? completed a meaningful action? passive background sync definitely should not count); the window length and its justification; whether the window rolls daily or snaps to calendar boundaries; and how a user with activity exactly on the boundary day is treated. Then leave it alone. Because the boundary is arbitrary, changing N after seeing results is one of the easiest ways to manufacture a favourable number, and it is visible to anyone who reads the history of the definition. Pre-register the window with the experiment, and if you genuinely need both a fast and a stable read, publish both, permanently labelled, rather than switching between them.

  • Why do teams prefer a 28-day window over a 30-day one?
    Because 28 days is exactly four weeks, so every window contains the same count of each weekday. A 30- or 31-day window shifts weekday composition as it slides, adding a sawtooth of a few percent that has nothing to do with user behaviour and is easily mistaken for a trend. The same logic makes 7 and 14 sensible short windows.
  • What goes wrong if a 14-day experiment reports a 28-day active-user metric?
    Half of each counted user's window predates their exposure, and pre-exposure activity is identical across arms in expectation. That signal-free mass dilutes the estimated effect toward zero, so a null is uninformative rather than reassuring. Either extend the test to a full window past the last enrolment or switch to a window that fits, labelling it clearly as a different metric.
  • How would you pick the window length for a new product?
    Start from the product's natural usage cadence: a tool used daily justifies a short window, one used monthly does not, and a window shorter than the typical gap between visits makes engaged users look churned. Then check what the window must support — fast regression detection wants short, stable executive reporting wants a whole number of weeks — and fix the choice before it can be tuned against results.

saying these in an interview costs you the question

  • Compares a 28-day figure directly against a 7-day figure
  • Reads a 28-day metric over a two-week experiment as conclusive
  • Thinks a longer window is simply more accurate
  • Ignores weekday composition when picking the length
  • Changes the window after seeing the experiment result

context