Only flagged machines get inspected, so misses stay invisible - how do you design a sample that estimates the failure model's false-negative rate?
answer
- one side of the matrix is unobserved
- you must buy the misses
- stratify by score band
- known inclusion probabilities, drawn first
- inverse-probability weighting, wide interval
basics
~20 sDraw a random sample of unflagged assets stratified by score band with known inclusion probabilities, inspect it, then weight each find by the inverse of its stratum's sampling rate to estimate misses across the whole fleet.
solid answer
~40 sPrecision can be read off the flagged cohort once it matures, but the false-negative rate cannot: nothing ever looks at the assets the model stayed quiet about. You have to buy that information with review capacity. Stratify the unflagged population by score band, fix a sampling rate per stratum, draw the sample randomly and **before** any inspection, then weight finds by the inverse rate. With 600 near-threshold assets sampled at 10% yielding 6 finds and 3,400 low-score assets sampled at 1% yielding 1, the estimated misses are `6 / 0.10 + 1 / 0.01 = 60 + 100 = 160`, so recall against 76 caught failures is `76 / (76 + 160) = 0.32`. Report an interval, not a point: one extra find in the 1% stratum moves the estimate by 100.
code
pseudocode · 14 linesstrata = [
{ name: "near_threshold", range: [0.40, alert_threshold), size: 600, rate: 0.10 },
{ name: "low", range: [0.00, 0.40), size: 3400, rate: 0.01 }
]
missed_estimate = 0
for each s in strata:
frame = unflagged assets whose score is in s.range // frozen as of the scoring date
sample = random_draw(frame, n = s.size * s.rate) // drawn BEFORE any inspection
found = count(a in sample where inspect(a) == "degrading")
missed_estimate = missed_estimate + found / s.rate // inverse-probability weight
// near_threshold: 6 / 0.10 = 60 ; low: 1 / 0.01 = 100 ; total 160
recall_estimate = caught / (caught + missed_estimate) // 76 / (76 + 160) = 0.32go deeper
Take away the asymmetry: the model's mistakes on machines it stayed quiet about are never discovered unless somebody deliberately goes and looks at a sample of them.
Explain stratification by score band, why the inclusion probability must be fixed in advance, and how inverse-probability weighting scales finds back to the fleet.
Show the operational awareness - inspection budget, the audit's own censoring effect, self-selection creeping back in, and reporting an interval rather than a falsely precise number.
The tradeoff is how much inspection capacity a recall number is worth, and whether a wide interval on the misses changes any decision the plant actually makes.
## Precision is observable, recall is not A flagged asset gets looked at, so its story is eventually written down. An unflagged asset is simply left alone, and unless it breaks loudly enough to generate its own record, the system never learns whether the model should have flagged it. That asymmetry means the two halves of the quality picture have completely different costs: precision on the flagged cohort arrives by itself once the cohort matures, while the false-negative rate has to be **purchased** with technician time. A monitoring design that reports only precision is reporting the half that was free. ## Designing the audit sample - **Stratify by score band.** Misses are far denser just under the alert threshold than at the bottom of the distribution, so sampling those bands at different rates buys more information per inspection hour than one uniform rate. Uniform sampling is unbiased too - it is simply wasteful, because most of its draws land where nothing is ever found. - **Fix the inclusion probability per stratum in advance**, and store it on every sampled row. The weighting later is only valid if the rate that produced each row is known exactly. - **Draw randomly inside the stratum and before anyone inspects.** If a shift supervisor picks which machines to audit, inclusion probabilities are unknown and unequal, and the weighting becomes arithmetic on a number nobody can define. - **Freeze the sampling frame.** Assets enter and leave service; the frame is the set of unflagged assets as of the scoring date, not whatever is on the floor when the auditor arrives. ## Weighting back to the fleet | stratum | unflagged assets | sampling rate | finds | weighted misses | |---|---|---|---|---| | near threshold | 600 | 10% | 6 | `6 / 0.10 = 60` | | low score | 3,400 | 1% | 1 | `1 / 0.01 = 100` | | **total** | 4,000 | - | 7 | **160** | Against 76 confirmed failures inside the flagged cohort, recall is `76 / (76 + 160) = 0.32`. Two health warnings travel with that number. First, the numerator carries whatever censoring the flagged cohort suffered, so it is a floor rather than a fact. Second, the estimate is dominated by the thin stratum: one more find among those 34 inspections adds another 100 to the miss count and drops recall to about `76 / 336 = 0.23`. ## The variance problem, and where to spend the next inspection 1. **Report an interval.** A point estimate built on seven finds implies a precision the data does not have. Attach an interval and the per-stratum find counts that generated it, so the reader can see which stratum is carrying the number. 2. **Rebalance deliberately.** Raising the low-score rate shrinks the swing that dominates the estimate; raising the near-threshold rate sharpens a number that is already reasonably precise. Spend on the stratum whose weight is largest, not the one where finds are easiest. 3. **Do not collapse the strata after the fact.** Pooling finds and dividing by a blended rate throws away the design and produces a number whose bias nobody can characterise. ## What the audit costs, and what quietly breaks it - **Technician hours are the real budget.** Every sampled inspection competes with responding to alerts, so the sampling rate is negotiated against operations, not chosen by an analyst alone. - **The audit intervenes on what it measures.** An inspector who finds a degrading bearing will usually fix it, which censors that row's 30-day outcome exactly as an alert-driven repair does. Record the audit verdict as an observation made at inspection time, not as a matured outcome. - **Inclusion probabilities drift.** If the alert threshold moves, the unflagged population changes shape and last quarter's strata no longer mean the same thing, so the rates must be restated with the threshold. - **Self-selection creeps back in.** "We inspected the ones we were passing anyway" quietly replaces the random draw; storing the drawn sample as its own record, before inspection, makes that substitution visible. - **Rare positives need volume.** At a low base rate the sample that would pin recall down precisely can exceed the whole fleet's inspection capacity. The honest answer is then a wide published interval, not a confident number computed from too few finds.
- Why must the audit sample be drawn before anyone inspects?Because the weighting is arithmetic on the inclusion probability, and a supervisor choosing machines makes that probability unknown and correlated with the outcome. Freeze the frame, draw randomly inside each stratum, store the drawn sample and its rate as a record, and only then inspect.
- What does an audit inspection do to the outcome it is measuring?It often changes it: an inspector who finds a degrading asset fixes it, so the row is censored by the audit itself in the same way an alert-driven repair censors a flagged row. Record the inspection verdict as an observation at inspection time rather than as a 30-day outcome, and keep the two fields separate.
- Does this audit replace the matured labels, since it gives an answer immediately?No. It estimates a quantity matured labels cannot give you at all, namely how many misses the quiet half of the fleet is hiding, but its verdict is an expert judgment at a point in time rather than an adjudicated 30-day outcome. The two are complements, and the audit's own agreement with matured outcomes is worth measuring.
saying these in an interview costs you the question
- Assuming recall is as measurable as precision from flagged assets alone
- Calling uniform sampling the efficient design for rare misses
- Letting a shift supervisor choose which machines get audited
- Reporting a point recall figure from a handful of finds
- Ignoring that an audit inspection often repairs what it inspects