In a churn model for a subscription service, what must one row of the training table fix before any feature is picked?
answer
- a row is a question already answered
- unit, horizon, as-of stamp
- one account in one billing month
- horizon names a window, not a flag
- features stop at the as-of stamp
basics
~20 sOne row must fix three things: the prediction unit (one account in one billing month), the horizon the label covers (cancels within the next 30 days), and the as-of timestamp that separates feature territory from label territory.
solid answer
~40 sA training row is a question that has already been answered, so it needs an entity, an as-of timestamp, features drawn only from before that stamp, and a label drawn only from after it. For a subscription churn ask, that means deciding the **prediction unit** (one paying account in one billing month is typical, but account-week and account-session are different tables), the **horizon** the label covers (`cancels within 30 days of the as-of stamp`, not the bare word churn), and the **as-of stamp** itself. The unit should match the cadence of the action the score feeds: a weekly outreach batch cannot use a once-per-lifetime score. Unit and horizon together set the base rate, so a churn rate quoted without both is meaningless.
go deeper
Recall the three parts of a row: which entity and period it covers, which forward window the label covers, and the timestamp that cuts features from label. Say them out loud before naming any feature.
Explain how the unit and the horizon jointly set the base rate and the row count, and why the same account appearing in many rows forces an account-level split rather than a random one.
Show you drive the unit from the intervention: the cadence of the retention action, the window it needs to work in, and the volume of scores the process can actually consume.
Frame the unit and horizon as a commitment that outlives the model - it fixes the label pipeline, the scoring volume and the operating cadence every later design decision has to fit.
## What a training row actually is A supervised training table is a list of questions that have already been answered. Each row carries an **entity**, an **as-of timestamp**, a set of **features** computed only from what was known at that timestamp, and a **label** recording what happened afterwards. Until those four parts are pinned down, 'predict churn' is a sentence, not a specification: two engineers can build two tables from it that share no rows, no positive count and no base rate. Three decisions fix the row. 1. **The prediction unit** - the entity-plus-timestamp pair. One paying account observed in one billing month is the common choice for a subscription business; one account per week, one account per session, or one account once over its lifetime are all defensible and all produce different tables. 2. **The horizon** - the forward window the label covers. `cancels within 30 days of the as-of stamp` is a horizon; 'will churn' is not. 3. **The as-of timestamp** - the instant separating feature territory from label territory. Everything at or before it may be a feature; everything after it belongs to the label. ## Choosing the unit The unit should match the cadence of the action the prediction feeds, not the convenience of whatever aggregate table already exists. | Prediction unit | One row is | Rows per account per year | Fits which action | |---|---|---|---| | Account lifetime | one account, scored once | 1 | a one-off segmentation study | | Account-month | one account in one billing month | 12 | a monthly outbound retention campaign | | Account-week | one account in one calendar week | 52 | a weekly call list with limited seats | | Account-session | one account on one visit | hundreds | an offer shown while the visit is live | A retention team that sends one batch of offers a week cannot act on a lifetime score, and an in-product intercept that must fire during a visit cannot wait for a monthly batch. The unit also decides how many rows one account contributes, and therefore how correlated the rows are: the same account appearing in twelve monthly rows is twelve near-copies, so a split that puts some of them in training and others in evaluation lets the model recognise the account rather than the pattern. Split by account, not by row. ## What the horizon buys and costs The horizon is a product decision before it is a modelling one, because it has to cover the time an intervention needs to work. Its effects run in opposite directions: - **A longer horizon raises the per-row base rate.** If roughly 2% of accounts cancel in a month, roughly 6% cancel over a quarter. More positives per row makes the table easier to learn from. - **A longer horizon delays every label.** Nothing inside the last horizon-length of data can be labelled yet, so the training cut moves further into the past and the model learns from a staler world. - **A longer horizon blurs the action.** 'Will leave sometime in the next six months' does not tell a retention team whether to call today. The unit and the horizon together set the base rate, which makes the base rate a design output rather than a fact about the business. Quoting 'our churn is 2%' without saying per what and over what window says nothing usable about the table. ## The as-of timestamp Every feature is an aggregate over a window that ends at the as-of stamp: sessions in the trailing 28 days, tickets opened in the previous fortnight, tenure measured from signup to the stamp. A feature whose window crosses the stamp is not a feature; it is a piece of the label wearing a feature's name. Writing the stamp into the row explicitly, rather than leaving it implied by the file the row came from, is what makes that rule checkable months later by someone who did not build the table. ## Where this goes wrong - A table with one row per account and no time column at all: nothing states when the prediction is made, so the leakage question cannot even be asked. - A horizon chosen after training, adjusted until the numbers look better. - A unit chosen because a monthly aggregate already exists, while the retention process runs weekly. - Rows split randomly for evaluation, so the same account lands on both sides of the split. - A base rate quoted with no unit or horizon attached, then compared against another team's number.
- How does switching the row from one account-month to one account-week change the churn base rate?It lowers the per-row base rate and raises the row count in roughly inverse proportion. At about 2% of accounts cancelling per month, an account-week row sits near 0.5% positive while each account contributes about four times as many rows. The total number of cancellations is unchanged; only their density per row moves, which is why a base rate is only meaningful once the unit and horizon are stated.
- Does the prediction unit have to match the cadence of the action it feeds?It should. Scoring per account-month while the retention team draws a call list weekly means three of every four lists are made from a stale score, and the horizon no longer lines up with the window the outreach can influence. Matching the unit to the action also makes the row count, and therefore the serving volume, a number you can defend rather than inherit.
saying these in an interview costs you the question
- One row per account, with no timestamp anywhere in the table
- Treating churn as a flag rather than an event inside a stated window
- Picking the horizon after the model trains, to improve the numbers
- Assuming the base rate is a fixed business fact, independent of the row
- Splitting train and evaluation rows at random, so an account appears on both sides