A predictive-maintenance model predicts failure within 30 days - why does this week's precision chart read optimistically high, and what corrects it?
answer
- two clocks, not one
- outcomes are still open
- failures resolve fast, survival slowly
- bucket by prediction date
- publish only matured cohorts
basics
~20 sRecent cohorts are still right-censored: a failure resolves the moment it happens while survival is only confirmed at day 30, so dropping unresolved rows over-counts failures. Bucket the metric by prediction date, publish only matured cohorts, and restate as labels land.
solid answer
~50 sA prediction made today has no adjudicable outcome for 30 days plus the lag it takes for a failure to be recorded and reconciled, so every row in this week's cohort is right-censored. Resolution is asymmetric: a failure announces itself on day 2 as readily as on day 29, while a healthy asset is only confirmed healthy once the whole window has elapsed. A metric job that quietly drops rows with no outcome yet therefore computes precision on a subset enriched with failures - 200 flags can read `21 / 30 = 0.70` after five days and `76 / 200 = 0.38` once the same cohort matures, with nothing about the model having changed. The fix is cohort accounting: bucket by prediction date, hold a cohort's figure until `scored_at + horizon + reconciliation_lag`, mark younger points provisional, and backfill the metric table as outcomes land.
code
pseudocode · 18 linesHORIZON_DAYS = 30
RECONCILE_DAYS = 7
function cohort_precision(cohort_date, now):
if now < cohort_date + HORIZON_DAYS + RECONCILE_DAYS:
return PENDING // every row is still right-censored
flagged = predictions where scored_at == cohort_date
and score >= alert_threshold
failed = 0
for each p in flagged:
outcome = failure_events.lookup(p.asset_id,
from = cohort_date,
to = cohort_date + HORIZON_DAYS)
if outcome exists:
failed = failed + 1
return failed / count(flagged) // denominator is every flag, not the resolved onesgo deeper
Remember that a prediction about the next 30 days has no outcome to compare against yet, and that a metric computed over such rows is measuring which labels arrived first.
Explain the mechanics: the horizon plus the reconciliation lag defines the maturation window, resolution is asymmetric because failures self-report and survival does not, and cohorts are bucketed by prediction date.
Show the operating discipline - a publication gate, a rewritable metric table, provisional versus final points on the chart, and a refusal to compare cohorts at different maturity in a review.
The tradeoff to name is how much reporting lag the business will absorb against how much bias it will tolerate, and which decisions are allowed to run on a provisional estimate at all.
## Two clocks, not one A prediction from the maintenance model says *this asset will fail within 30 days*. Two different durations follow from that sentence, and confusing them is the root defect in most delayed-label dashboards: - the **prediction horizon** is the 30 days of future time the claim is about; - the **label-maturation window** is the wall-clock delay before the claim can be adjudicated - the horizon **plus** the reconciliation lag it takes for a failure to be noticed, written against the right asset, closed out, and landed in the store the metric job reads. A plant on a weekly paperwork cycle can easily run a 30-day horizon and a 37-day maturation window. Measure that lag rather than assuming it. Until a row's maturation window has elapsed, its outcome is **right-censored**: you know the asset had not failed as of the last time anyone looked, and nothing more. ## Why the young end of the chart reads high Resolution is not symmetric in time. A failure announces itself the moment it happens, on day 2 as readily as on day 29. Survival announces itself once, at the end of the window, when the asset has gone the whole 30 days intact. So the subset of a young cohort whose outcome is already known is **enriched with failures**, and a metric job that silently drops rows with no outcome yet reports a number computed on that enriched subset. Worked, on one day's flags, with a row counted as resolved when a failure is recorded or a technician closes the asset as healthy: - 200 assets are flagged on the same day; - five days later 30 rows are resolved - 21 failures and 9 closed healthy - and the chart shows precision `21 / 30 = 0.70`; - at day 37 all 200 rows are resolved, 76 of them failures, and the same cohort reads `76 / 200 = 0.38`. Nothing about the model changed between those two readings. The first number measured label arrival order. ## Two naive handling rules, two opposite errors | rule for unresolved rows | what a young cohort reads | why | |---|---|---| | drop them from the metric | too high | the resolved subset over-represents early failures | | count them as negatives | too low | flagged assets have not had their full 30 days to fail | | hold the cohort until it matures | correct, and late | every row has had the whole window | Both naive rules yield a confident number and differ only in the direction they lie, so "ours was pessimistic, we were being conservative" is not a defence either. ## What the pipeline has to carry 1. Stamp every prediction with its scoring time, its horizon and an explicit **outcome state** - `pending`, `failed`, `survived`, `censored` - never an implicit negative inferred from a missing row. 2. Bucket the metric by **prediction date**, not by the date the metric job ran; the cohort, not the calendar day, is the unit that matures. 3. Gate publication so a cohort's figure is emitted only once `now >= scored_at + horizon + reconciliation_lag`. 4. **Backfill and restate**: a figure for a prediction date stays provisional until its cohort matures, so that row of the metric table must be rewritable and the chart must show which points are final. 5. Never compare a young cohort with a matured one - that comparison is the most common route by which this bias reaches a decision. ## What you can honestly claim before the labels land - Matured cohorts give a **final** number, so the model is never unmeasurable; it is measured with a known lag, and the lag is what you state next to the figure. - A young cohort can carry a **censoring-aware estimate** that uses partial windows instead of forcing a binary: a survival-style estimator such as the **Kaplan-Meier** estimator treats an asset watched for 6 of its 30 days as six days of evidence rather than as a survivor. Publish it with an interval and the share of the cohort already matured. - Every **label-hungry** measure inherits the same lag - precision, recall, **PR-AUC**, the **Brier score** and **expected calibration error** all need matured outcomes, so a claim that the live model's calibration has drifted can only be made on a cohort past the horizon, and drift in calibration is detected with exactly the same delay as drift in accuracy. - Segments mature at different speeds: a line whose work orders are closed weekly matures later than one closed daily, so a single global gate publishes a mixture unless the gate is applied per segment. - A figure that carries an as-of stamp and a maturity percentage survives being forwarded to someone who was not in the room; a bare number does not.
- What can you say about the live model's calibration today, given the same delay?Nothing about today. Calibration measures such as the Brier score and expected calibration error are label-hungry, so they can only be computed on cohorts past the maturation window. A recent slice will look well calibrated purely because its known outcomes are enriched with failures, which is an artefact of arrival order, not evidence about calibration.
- Where does the reconciliation lag come from, on top of the horizon?From the operational path a real outcome takes: someone has to notice the failure, attribute it to the right asset, close the paperwork, and let the batch land in the store the metric job reads. Measure that distribution from historical rows and set the gate at a high percentile of it rather than at its mean.
- Should one maturation gate apply to every segment of the fleet?Not automatically. If one production line closes work orders daily and another weekly, their labels mature at different speeds, so a single global gate publishes a mixture of matured and half-matured rows. Either apply the gate per segment with its own measured lag, or set one gate at the slowest segment's lag and accept the extra delay.
saying these in an interview costs you the question
- Treating today's precision chart as a measurement of today's model
- Counting unresolved predictions as negatives and calling it conservative
- Assuming the label arrives the instant the 30-day horizon ends
- Believing a published figure for a date should never change
- Claiming nothing about quality can be measured until all labels land