skip to content

Technicians repair every asset the 30-day failure model flags, so it never fails - why does the label pipeline then record those flags as false positives?

level: seniorimportance: must knowfreq 64%

answer

  1. the action destroyed the evidence
  2. a prevented failure leaves no event
  3. censored, not negative
  4. success looks like crying wolf
  5. randomised hold-back restores the denominator

basics

~20 s

Maintenance destroyed the counterfactual. The alert caused a repair, no failure event was written inside the window, and a pipeline that equates 'no failure recorded' with 'negative outcome' books a prevented failure as a false positive.

solid answer

~50 s

The label rule is usually "a failure event exists inside the 30-day window", and a successful intervention guarantees none exists. The row is not a negative; it is **censored at the moment of intervention**, because the outcome the label is supposed to describe was prevented rather than observed. The better the response, the worse the model measures: on a line with 300 flags in a month where crews act on 285 and hold back 15 for capacity, 6 failures among the acted rows and 6 among the held-back rows read as `12 / 300 = 0.04` precision, while the only rows whose counterfactual was observable read `6 / 15 = 0.40`. The repair has to be recorded as a first-class outcome state so downstream consumers can filter on it, and an honest denominator needs rows where the alert was deliberately not acted on.

code

json · 16 lines
json
{
  "prediction_id": "p-88213",
  "asset_id": "line-3-pump-07",
  "scored_at": "2026-02-01",
  "horizon_days": 30,
  "score": 0.81,
  "outcome_state": "censored_intervened",
  "observed_through": "2026-02-04",
  "intervention": {
    "acted_at": "2026-02-04",
    "action": "bearing_replaced",
    "technician_verdict": "degradation_confirmed"
  },
  "holdback_arm": "acted",
  "usable_as_training_label": false
}

go deeper

for a junior

Grasp the core shape: if someone fixes the machine because of the alert, the failure never happens, so there is no outcome recorded and the flag can look wrong on paper.

for a middle

Explain censoring as a third outcome state with its own timestamp, and why a label rule written as 'no failure event means negative' produces the error automatically.

for a senior

Demonstrate that you would notice this in a real fleet, quantify the censored share by segment, and design a hold-back or delayed-action arm that the plant would actually approve.

for a principal

The judgment is what evidence is worth buying: the cost and risk of leaving some alerts unserviced against running a system whose measured quality is permanently unreadable.

## The prevented failure has no label The maintenance model flags a pump; a crew replaces the bearing on day 4; the asset runs cleanly for the remaining 26 days. The window closes with no failure event against that asset, and a pipeline whose label rule is "a failure event exists inside the 30-day window" writes `survived` - a false positive against a flag that was very likely right. This is **censoring by intervention**. The outcome the label claims to describe - what the asset would have done if left alone - was destroyed by the action the model exists to trigger. The prediction is not wrong and the label is not right; the row simply carries no evidence about which it is, and coding it as a negative asserts evidence that does not exist. ## What it does to the numbers One month on a line where the crew acts on essentially every alert: - 300 assets flagged; 285 acted on inside the window, 15 left unserviced because the crew ran out of capacity; - among the 285 acted rows, 6 failed anyway; - among the 15 unserviced rows, 6 failed. Code every non-failing acted row as a negative and measured precision is `12 / 300 = 0.04`, which reads as a broken model. The only rows in which the counterfactual was actually observed are the 15 unserviced ones, giving `6 / 15 = 0.40` - and that estimate is trustworthy only to the degree the hold-back was **random**. If the crew skipped the alerts it personally distrusted, or the assets it considered unimportant, the 15 are a self-selected sample and the estimate inherits their judgment. ## Three outcome states, not two | outcome state | what happened | safe as a training label? | in the precision denominator? | |---|---|---|---| | `failed` | a failure event landed inside the window | yes | yes | | `survived` | the full window elapsed with no intervention and no failure | yes | yes | | `censored_intervened` | the asset was acted on before the window closed | no | no - excluded, or carried by a censoring-aware estimator | Excluding censored rows is not free: it leaves a denominator of rows nobody chose to fix, which tilts the other way. The defensible options are to exclude them **and publish the censored share** alongside the figure, or to keep them with an estimator that treats intervention as a censoring time. Both are better than silently coding them negative, which is the one choice that asserts a fact nobody observed. ## Why the damage does not stop at the dashboard - The label store is not private to the dashboard; anything that reads it inherits the distortion, so a censored row coded as a negative teaches every downstream consumer that this pattern was safe. - The distortion concentrates where the system works best: the alerts crews trust most are acted on fastest, so the strongest signals are the most reliably mislabelled. - It is segment-dependent. Critical assets get same-shift response and high-censoring rates; low-criticality assets sit, so their rows resolve honestly. A fleet-level precision number is then a weighted average over segments with completely different label quality. - It is invisible in any count of alerts or failures on their own - only the join between a prediction, an intervention record and an outcome exposes it. ## Getting an honest denominator back 1. **A randomised hold-back arm.** A small share of alerts, chosen at random, is not shown to the crew, and those rows resolve on their own merits. This is the only design that yields an unbiased estimate, and it is an operational decision, not an analytical one: cap it, restrict it to low-criticality assets, and write in an explicit safety veto. 2. **Delayed action instead of suppression.** Act at day 7 instead of day 0, and you observe a real if truncated window on every delayed row - weaker than a hold-back arm, far easier to authorise. 3. **A verdict label at the moment of intervention.** The technician records what was actually found - a degraded bearing or a healthy one. It is available on every acted row and costs nothing extra, but the technician has seen the alert, so it is a proxy that leans toward confirming it; blind the work order where the process allows. 4. **Capacity-limited natural experiments.** Days when alerts went unserviced give observed outcomes for free, but they are not random - they cluster on busy shifts and specific lines, so treat them as a sanity check rather than as the measurement. The design rule underneath all four: the pipeline must record **that an intervention happened and when**, because after the fact nothing else distinguishes a prevented failure from a model that was simply wrong.

  • How do you size a hold-back arm without risking equipment?
    Restrict it to low-criticality assets, cap the share and the duration, and prefer delaying the alert to suppressing it so a partial window is still observed. Write a safety veto into the design that lets an engineer pull any asset out of the arm, and accept that the arm then yields a slightly optimistic estimate rather than none at all.
  • What has to be recorded so a censored row is never mistaken for a negative later?
    Three things on the prediction row: the outcome state as an explicit enumeration, the time through which the asset was actually observed, and a reference to the intervention. Downstream consumers then filter on the state rather than inferring a negative from the absence of a failure event, which is the inference that caused the problem.
  • The technician's verdict at repair time is available on every acted row - why not just use it as the label?
    Because the technician saw the alert before opening the machine, so the verdict leans toward confirming it, and a verdict of 'degrading' is not the same claim as 'would have failed within 30 days'. Use it as a proxy whose agreement with matured outcomes you have measured, not as a drop-in replacement for the real outcome.

A building where the sprinklers always fire before anyone reports a fire: the fire log stays empty, so on paper the alarm system has never once been right.

saying these in an interview costs you the question

  • Assuming no failure inside the window always means the flag was wrong
  • Feeding intervened rows back into the label store as negatives
  • Believing censoring is a modelling nicety that never touches production labels
  • Letting crews self-select which alerts go unserviced and calling it a control
  • Thinking enough alert volume removes the need for a hold-back arm