skip to content

questions

5

A predictive-maintenance model predicts failure within 30 days - why does this week's precision chart read optimistically high, and what corrects it?

level: middleimportance: must knowfreq 76%

answer

  1. two clocks, not one
  2. outcomes are still open
  3. failures resolve fast, survival slowly
  4. bucket by prediction date
  5. publish only matured cohorts

basics

~20 s

Recent cohorts are still right-censored: a failure resolves the moment it happens while survival is only confirmed at day 30, so dropping unresolved rows over-counts failures. Bucket the metric by prediction date, publish only matured cohorts, and restate as labels land.

solid answer

~50 s

A prediction made today has no adjudicable outcome for 30 days plus the lag it takes for a failure to be recorded and reconciled, so every row in this week's cohort is right-censored. Resolution is asymmetric: a failure announces itself on day 2 as readily as on day 29, while a healthy asset is only confirmed healthy once the whole window has elapsed. A metric job that quietly drops rows with no outcome yet therefore computes precision on a subset enriched with failures - 200 flags can read `21 / 30 = 0.70` after five days and `76 / 200 = 0.38` once the same cohort matures, with nothing about the model having changed. The fix is cohort accounting: bucket by prediction date, hold a cohort's figure until `scored_at + horizon + reconciliation_lag`, mark younger points provisional, and backfill the metric table as outcomes land.

code

pseudocode · 18 lines
pseudocode
HORIZON_DAYS = 30
RECONCILE_DAYS = 7

function cohort_precision(cohort_date, now):
    if now < cohort_date + HORIZON_DAYS + RECONCILE_DAYS:
        return PENDING            // every row is still right-censored

    flagged = predictions where scored_at == cohort_date
                            and score >= alert_threshold
    failed = 0
    for each p in flagged:
        outcome = failure_events.lookup(p.asset_id,
                                       from = cohort_date,
                                       to   = cohort_date + HORIZON_DAYS)
        if outcome exists:
            failed = failed + 1

    return failed / count(flagged)   // denominator is every flag, not the resolved ones

go deeper

for a junior

Remember that a prediction about the next 30 days has no outcome to compare against yet, and that a metric computed over such rows is measuring which labels arrived first.

for a middle

Explain the mechanics: the horizon plus the reconciliation lag defines the maturation window, resolution is asymmetric because failures self-report and survival does not, and cohorts are bucketed by prediction date.

for a senior

Show the operating discipline - a publication gate, a rewritable metric table, provisional versus final points on the chart, and a refusal to compare cohorts at different maturity in a review.

for a principal

The tradeoff to name is how much reporting lag the business will absorb against how much bias it will tolerate, and which decisions are allowed to run on a provisional estimate at all.

## Two clocks, not one A prediction from the maintenance model says *this asset will fail within 30 days*. Two different durations follow from that sentence, and confusing them is the root defect in most delayed-label dashboards: - the **prediction horizon** is the 30 days of future time the claim is about; - the **label-maturation window** is the wall-clock delay before the claim can be adjudicated - the horizon **plus** the reconciliation lag it takes for a failure to be noticed, written against the right asset, closed out, and landed in the store the metric job reads. A plant on a weekly paperwork cycle can easily run a 30-day horizon and a 37-day maturation window. Measure that lag rather than assuming it. Until a row's maturation window has elapsed, its outcome is **right-censored**: you know the asset had not failed as of the last time anyone looked, and nothing more. ## Why the young end of the chart reads high Resolution is not symmetric in time. A failure announces itself the moment it happens, on day 2 as readily as on day 29. Survival announces itself once, at the end of the window, when the asset has gone the whole 30 days intact. So the subset of a young cohort whose outcome is already known is **enriched with failures**, and a metric job that silently drops rows with no outcome yet reports a number computed on that enriched subset. Worked, on one day's flags, with a row counted as resolved when a failure is recorded or a technician closes the asset as healthy: - 200 assets are flagged on the same day; - five days later 30 rows are resolved - 21 failures and 9 closed healthy - and the chart shows precision `21 / 30 = 0.70`; - at day 37 all 200 rows are resolved, 76 of them failures, and the same cohort reads `76 / 200 = 0.38`. Nothing about the model changed between those two readings. The first number measured label arrival order. ## Two naive handling rules, two opposite errors | rule for unresolved rows | what a young cohort reads | why | |---|---|---| | drop them from the metric | too high | the resolved subset over-represents early failures | | count them as negatives | too low | flagged assets have not had their full 30 days to fail | | hold the cohort until it matures | correct, and late | every row has had the whole window | Both naive rules yield a confident number and differ only in the direction they lie, so "ours was pessimistic, we were being conservative" is not a defence either. ## What the pipeline has to carry 1. Stamp every prediction with its scoring time, its horizon and an explicit **outcome state** - `pending`, `failed`, `survived`, `censored` - never an implicit negative inferred from a missing row. 2. Bucket the metric by **prediction date**, not by the date the metric job ran; the cohort, not the calendar day, is the unit that matures. 3. Gate publication so a cohort's figure is emitted only once `now >= scored_at + horizon + reconciliation_lag`. 4. **Backfill and restate**: a figure for a prediction date stays provisional until its cohort matures, so that row of the metric table must be rewritable and the chart must show which points are final. 5. Never compare a young cohort with a matured one - that comparison is the most common route by which this bias reaches a decision. ## What you can honestly claim before the labels land - Matured cohorts give a **final** number, so the model is never unmeasurable; it is measured with a known lag, and the lag is what you state next to the figure. - A young cohort can carry a **censoring-aware estimate** that uses partial windows instead of forcing a binary: a survival-style estimator such as the **Kaplan-Meier** estimator treats an asset watched for 6 of its 30 days as six days of evidence rather than as a survivor. Publish it with an interval and the share of the cohort already matured. - Every **label-hungry** measure inherits the same lag - precision, recall, **PR-AUC**, the **Brier score** and **expected calibration error** all need matured outcomes, so a claim that the live model's calibration has drifted can only be made on a cohort past the horizon, and drift in calibration is detected with exactly the same delay as drift in accuracy. - Segments mature at different speeds: a line whose work orders are closed weekly matures later than one closed daily, so a single global gate publishes a mixture unless the gate is applied per segment. - A figure that carries an as-of stamp and a maturity percentage survives being forwarded to someone who was not in the room; a bare number does not.

  • What can you say about the live model's calibration today, given the same delay?
    Nothing about today. Calibration measures such as the Brier score and expected calibration error are label-hungry, so they can only be computed on cohorts past the maturation window. A recent slice will look well calibrated purely because its known outcomes are enriched with failures, which is an artefact of arrival order, not evidence about calibration.
  • Where does the reconciliation lag come from, on top of the horizon?
    From the operational path a real outcome takes: someone has to notice the failure, attribute it to the right asset, close the paperwork, and let the batch land in the store the metric job reads. Measure that distribution from historical rows and set the gate at a high percentile of it rather than at its mean.
  • Should one maturation gate apply to every segment of the fleet?
    Not automatically. If one production line closes work orders daily and another weekly, their labels mature at different speeds, so a single global gate publishes a mixture of matured and half-matured rows. Either apply the gate per segment with its own measured lag, or set one gate at the slowest segment's lag and accept the extra delay.

saying these in an interview costs you the question

  • Treating today's precision chart as a measurement of today's model
  • Counting unresolved predictions as negatives and calling it conservative
  • Assuming the label arrives the instant the 30-day horizon ends
  • Believing a published figure for a date should never change
  • Claiming nothing about quality can be measured until all labels land
open as a page

Technicians repair every asset the 30-day failure model flags, so it never fails - why does the label pipeline then record those flags as false positives?

level: seniorimportance: must knowfreq 64%

basics

~20 s

Maintenance destroyed the counterfactual. The alert caused a repair, no failure event was written inside the window, and a pipeline that equates 'no failure recorded' with 'negative outcome' books a prevented failure as a false positive.

open as a page

Instead of waiting 30 days for a breakdown, a team labels its failure model from closed maintenance work orders - how does that proxy lie?

level: middleimportance: should knowfreq 52%

basics

~20 s

A work order records that maintenance happened, not that an asset was failing. Alerts cause orders, preventive work opens them on healthy assets, coverage is incomplete, and closure timestamps are paperwork dates rather than fault times.

open as a page

Only flagged machines get inspected, so misses stay invisible - how do you design a sample that estimates the failure model's false-negative rate?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Draw a random sample of unflagged assets stratified by score band with known inclusion probabilities, inspect it, then weight each find by the inverse of its stratum's sampling rate to estimate misses across the whole fleet.

open as a page

Plant leadership wants one monthly quality number for the 30-day failure model while most of the month's predictions are unresolved - what do you publish, and what do you commit to?

level: principalimportance: nice to knowfreq 31%

basics

~20 s

Publish the last fully matured cohort as the official figure, plus a clearly marked provisional estimate carrying an interval and a maturity percentage, and commit to a restatement policy that says when a number becomes final and who is told when it moves.

open as a page