skip to content

With a 30-day churn horizon plus a 14-day grace period, which subscription rows cannot be labelled yet?

level: middleimportance: should knowfreq 49%

answer

  1. the outcome has not happened yet
  2. horizon plus grace, not horizon alone
  3. 44 days of unlabellable rows
  4. training cut sits six weeks back
  5. cadence does not move the frontier

basics

~20 s

Every row whose as-of timestamp is newer than 44 days ago. The horizon plus the grace period is the label maturity lag, so the training cut sits 44 days in the past and the most recent six weeks of data carries features but no trustworthy outcome.

solid answer

~40 s

The outcome a row predicts is only observable once its whole forward window has passed and the grace period has closed, so the maturity lag is `horizon + grace` - here 30 + 14 = 44 days. Any row stamped inside the last 44 days has features but no final label, and including it labels genuine future churners as negatives. The practical consequence is that every retrain, however often it runs, learns from a world at least 44 days old: a price change or a campaign launched last month is invisible to training even though it is already shaping live traffic. Retraining more frequently does not shorten the lag, because cadence and maturity are independent. The levers that do are a shorter horizon, a shorter grace period, or a separate earlier-maturing target.

code

pseudocode · 14 lines
pseudocode
HORIZON_DAYS = 30
GRACE_DAYS   = 14
LABEL_LAG    = HORIZON_DAYS + GRACE_DAYS   // 44

for each row in candidate_rows:
    if row.as_of + LABEL_LAG > today:
        row.label = null                   // outcome window still open
        park row in the pending table
        continue

    window_end = row.as_of + HORIZON_DAYS
    row.label  = paid_access_ends_between(row.account_id, row.as_of, window_end)
                 and no_new_paid_period_by(row.account_id, window_end + GRACE_DAYS)
    emit row into the training table

go deeper

for a junior

Remember that a row cannot be labelled until its whole forward window has passed, so the newest data in the warehouse is not training data yet.

for a middle

Compute the lag as horizon plus grace and explain why labelling immature rows as negatives biases the table towards the accounts that were about to leave.

for a senior

Show what the lag does to operations: which regime changes training cannot see, why cadence is not a fix, and how the evaluation window has to move back with the cut.

for a principal

Own the trade the lag forces - shorten the horizon, shorten grace, adopt an earlier-maturing target, or accept that the model always reasons about a world six weeks old.

## Why the newest rows carry no label A row's label answers a question about its future. For an as-of date `T` with horizon `H`, the question is whether the account's paid access ends within `(T, T + H]`, and that cannot be answered before `T + H` has arrived. Silent non-renewals need longer still, because a lapse is only confirmed once the grace period `G` has closed without a new paid period. So the earliest moment the row is trustworthy is `T + H + G`. Turn that around and it becomes a cut on the table: 1. Today is `D`. 2. The **label maturity lag** is `H + G`, here 30 + 14 = **44 days**. 3. Every row with `T > D - 44` is unlabellable. Those rows can be built, featurised and inspected, but their outcome is still in flight. There is also a partial-maturity stage that trips people up. At `T + H`, explicit cancellations inside the window are already known; only the silent lapses are pending. A table labelled at `T + H` is therefore not unlabelled - it is **biased**, systematically missing exactly the silent portion of churn, which is often the larger half. ## What the lag costs - **The model always learns from an old world.** The freshest training row describes conditions from six weeks ago, no matter when the run happens. - **Recent regime changes are invisible.** A price rise, a packaging change or a big acquisition campaign launched last month shapes live traffic but appears in no labelled row. - **Retraining cadence does not help.** Running the job nightly rather than monthly refreshes which rows are included; it does not move the frontier of what is knowable. - **Evaluation inherits the same cut.** A holdout period must also sit entirely before the cut, which pushes the evaluated window further back again. - **Incidents are hard to confirm.** If something broke three weeks ago, no matured labels exist for the period yet. ## The levers, and what each one trades | Lever | What it buys | What it costs | |---|---|---| | Shorten the horizon to 14 days | labels mature in 28 days instead of 44 | fewer positives per row, and a window too short for slow interventions | | Shorten the grace period to 7 days | labels mature a week sooner | late renewals get labelled positive, so the label gets noisier | | Add an earlier-maturing target | a signal available in weeks | it is no longer the business outcome, and the gap has to be tracked | | Accept the lag | the label stays honest and reconstructible | training never sees the last six weeks | None of these is free, and the one that is quietly worst is the unstated fourth option: labelling recent rows as negatives because nothing has happened yet. That converts every not-yet-churned account into a confident negative and biases the table towards exactly the accounts that were about to leave. ## Making the cut explicit The cut belongs in the table-building step as an explicit guard, not as a comment in a runbook. Each row is emitted only when its entire outcome window, including grace, has closed; everything newer is retained with a null label so it can be relabelled later rather than silently dropped and rebuilt from scratch. ## Where this goes wrong - Labelling rows up to yesterday and calling the empty outcomes negatives. - Setting the cut at the horizon alone, forgetting the grace period, and importing partially matured silent lapses as negatives. - Assuming a faster retraining cadence closes the gap. - Reporting an evaluation window that overlaps the immature region, so the numbers look better than they are. - Quoting the lag as a modelling detail rather than telling the business that the newest six weeks cannot be learned from.

  • Would retraining weekly instead of monthly shorten the 44-day lag?
    No. Cadence decides how often the table is rebuilt; maturity decides how far back the newest usable row sits. A weekly run repeatedly trains on data whose frontier is still 44 days old, which is worth doing for other reasons but buys nothing here. The only things that move the frontier are a shorter horizon, a shorter grace period, or a different target that matures sooner.
  • What breaks if you cut at the horizon alone and ignore the grace period?
    The last 14 days of rows are labelled before silent lapses can be confirmed, so accounts that quietly failed to renew enter the table as negatives. The bias is not random: it removes exactly one kind of churn from the positive class in the most recent slice, teaching the model that recent non-renewers look healthy.

saying these in an interview costs you the question

  • Rows up to yesterday can be labelled; nothing happened, so they are negatives
  • The training cut sits at the horizon, and grace is only a reporting detail
  • Retraining more often gives the model fresher labels
  • A campaign launched last month is already represented in training data
  • The evaluation window can overlap the immature region if the model is good