skip to content

Instead of waiting 30 days for a breakdown, a team labels its failure model from closed maintenance work orders - how does that proxy lie?

level: middleimportance: should knowfreq 52%

answer

  1. it records paperwork, not physics
  2. downstream of the alert
  3. preventive work on healthy assets
  4. closure date is not fault time
  5. measure agreement on matured rows

basics

~20 s

A work order records that maintenance happened, not that an asset was failing. Alerts cause orders, preventive work opens them on healthy assets, coverage is incomplete, and closure timestamps are paperwork dates rather than fault times.

solid answer

~50 s

The proxy answers a different question from the one the model was asked. Four distortions matter: **contamination**, because the alert itself causes the work order, so the proxy partly mirrors the model's own output; **coverage bias**, because failures fixed under a generic order or during scheduled downtime never attach to the asset; **timestamp skew**, because a closure date is when paperwork finished, which can place a row in the wrong cohort; and **semantic mismatch**, because preventive work opens orders on perfectly healthy assets. The only honest way to use one is to measure it: on the cohort where both the proxy and the matured outcome exist, 145 proxy positives against 90 matured positives with 68 overlapping gives proxy precision `68 / 145 = 0.47` and recall `68 / 90 = 0.76`. That licenses the proxy as a directional early read with a stated bias, not as a substitute label.

go deeper

for a junior

Hold on to the distinction: a maintenance record says work happened, while the model claimed the asset would fail, and those are not the same statement about the world.

for a middle

Name the specific distortions - contamination by the alert, coverage gaps, timestamp skew, preventive work on healthy assets - and show how to quantify them against matured rows.

for a senior

Show the operating rules you would impose: separate columns, separate chart series, a re-measurement cadence, and a preference for signals the model does not cause.

for a principal

The call to own is how much bias buys how much speed, and which decisions may run on a fast biased reading versus waiting for the matured cohort.

## What the proxy is being asked to stand in for The real outcome is a fact about an asset: *did it fail within 30 days of being scored?* A work-order proxy substitutes a fact about an organisation: *did someone open and close a maintenance record against this asset?* Those two facts correlate, which is what makes the proxy tempting when the real one is a month away, and they differ in specific, directional ways that a design conversation should be able to name. ## Four ways the work-order proxy lies - **Contamination by the model's own output.** The alert is why the crew opened the order. Scoring the model against a signal it caused measures the alerting workflow's compliance rate, and it flatters exactly the predictions that were acted on most eagerly. A proxy downstream of the model is a mirror, not a measurement. - **Coverage bias.** Failures get fixed without a matching record: a fault cleared during scheduled downtime, work booked against a line rather than an asset, a swap logged as a parts issue. Those rows look like healthy assets, and they are concentrated in exactly the corners of the fleet whose paperwork is weakest. - **Timestamp skew.** A work order carries several times - opened, worked, closed - and they can be days or weeks apart. Joining on the closure date pulls a record into the wrong cohort and can make an order that was opened before the prediction appear to follow it. The label's event time must be the fault-observation time, chosen once and documented. - **Semantic mismatch.** Preventive maintenance opens orders on assets in perfect health, and a free-text severity field written differently by each technician cannot be mapped to a clean binary. "A repair happened" is simply not "this asset was going to fail within 30 days". ## Measuring the lie instead of arguing about it The proxy's error is estimable, because on any cohort old enough to have matured you hold both labels. Build the two-by-two directly: | | matured outcome: failed | matured outcome: survived | |---|---|---| | **proxy says failed** | 68 | 77 | | **proxy says survived** | 22 | the rest of the cohort | On 90 matured positives and 145 proxy positives with 68 in common, the proxy's precision against the truth is `68 / 145 = 0.47` and its recall is `68 / 90 = 0.76`. That is a proxy that catches most real failures and roughly doubles the positive count - useful as a direction, dangerous as a label. State both numbers whenever the proxy is quoted. ## Rules for using a proxy you have measured 1. **Direction, not level.** A proxy-based figure is a nowcast of movement; the level belongs to matured cohorts. Publish the proxy series and the matured series as separate lines, never blended into one. 2. **Keep the columns separate.** The proxy label and the matured label live in different fields on the row, so nothing downstream can silently consume one as the other. 3. **Re-measure on a schedule.** The proxy's bias is an artefact of current operations; a new preventive-maintenance policy or a change in how orders are coded moves it without anyone touching the model. 4. **Prefer a proxy the model did not cause.** An unplanned-stoppage log or a downstream product-defect rate on the line is a weaker signal than a work order but is not downstream of the alert, so it can contradict the model rather than echo it. ## When a proxy is worth having at all The honest case for a proxy is time. A 30-day horizon plus reconciliation means a real regression is invisible for over a month, and a measured, biased, fast signal that moves when something breaks is worth more than an unbiased signal that arrives six weeks late. The failure mode is forgetting that it was a compromise: the proxy gets quoted without its measured bias, then promoted into the label column, and a month later nobody can say what the number meant. The discipline is cheap - a named owner for the proxy definition, a stored agreement measurement with the date it was taken, and a rule that no irreversible decision runs on a proxy reading alone.

  • How do you decide whether a proxy is good enough to act as an early read?
    Measure it against matured outcomes on the overlap cohort and publish its precision and recall against them, then use it only for direction with that bias stated. Re-measure whenever maintenance policy or coding practice changes, because the proxy's error is a property of current operations rather than a constant.
  • Which timestamp should a work-order-derived label carry?
    The fault-observation time, not the closure time. Closure drifts with paperwork cycles and varies by shift, so joining on it can move a row into the wrong cohort and can even make an order that predates the prediction look like its consequence. Pick one time field, document it, and use it everywhere.
  • Why is a proxy that the model itself triggers worse than a noisier independent one?
    Because it cannot disagree with the model in the direction that matters. If alerts drive work orders, a model that starts flagging the wrong assets still produces matching orders, so the proxy tracks compliance rather than correctness. An independent signal such as an unplanned-stoppage log is noisier but can actually contradict the model.

saying these in an interview costs you the question

  • Treating a closed work order as proof the asset was about to fail
  • Assuming proxy bias averages out given enough volume
  • Reading a high agreement rate as evidence the proxy is safe
  • Joining labels on the paperwork closure date
  • Promoting a proxy into the label column once it looks reasonable