skip to content

A sensor logs occasional 250C readings; how do you decide if they are contamination or genuine extremes?

level: seniorimportance: should knowfreq 56%

answer

  1. the rule flags unusual, not wrong
  2. ask where the number came from
  3. an independent source that agrees or does not
  4. flag rather than delete; run it both ways

basics

~20 s

No statistical rule can answer this. Check provenance: the instrument's rated range, sentinel codes, unit mix-ups and an independent source. A statistic says a value is unusual; only evidence about its origin says it is wrong.

solid answer

~40 s

Outlier rules detect unusualness, and unusualness has two completely different causes: a broken measurement process, and a real but rare event. Deciding between them is a data-provenance question. I would check whether the value exceeds the instrument's rated range — 250C from a sensor rated to 60C is physically impossible, which is decisive evidence of a fault. I would look for sentinel or rail values that a failing device emits, unit or scale mix-ups, and whether the accompanying fields are simultaneously implausible. Then corroboration: does an independent sensor or a nearby station show anything similar? A genuine record heatwave reading arrives with supporting context; an instrument fault arrives alone. When the evidence is inconclusive, I flag rather than delete, keep the raw value, and report the analysis both with and without the disputed points.

go deeper

for a junior

Know that a flagged point is a prompt to look at the raw record, not permission to delete it, and that raw values should be preserved.

for a middle

Explain concretely how you would investigate: rated instrument range, sentinel codes, unit mix-ups, and whether the other fields on the same record are also implausible.

for a senior

Show the process discipline of someone who has been burned: flag rather than delete, seek independent corroboration, and report results with and without disputed points.

for a principal

Own the policy. Decide where validation belongs in the pipeline, who adjudicates disputed records, and how the asymmetric cost of discarding real extremes shapes the default.

## The question a statistic cannot answer Every outlier rule — fences, z-scores, robust scores — answers exactly one question: is this value unusual relative to the rest of the column? That is not the question anyone actually cares about. The decision-relevant question is whether the value is a faithful measurement of something rare, or a corrupt record of something ordinary. Those two cases demand opposite treatment: a real extreme is often the most valuable row in the table, and a corrupt one poisons every estimate it touches. No amount of arithmetic on the column itself separates them, because both look identical from inside the distribution. ## Evidence that points to contamination - **Physical impossibility.** A value outside what the measured quantity or the instrument can produce settles the matter. An air-temperature sensor rated to 60C reporting 250C is not measuring air. - **Rail and sentinel values.** Failing devices commonly emit a fixed extreme: the top of the register, a repeated digit, or a coded placeholder such as 999 or -9999 that a downstream loader silently treated as a number. If the suspicious values are all identical, that is a strong signal of a code rather than a measurement. - **Unit and scale mix-ups.** A stream that switched units mid-collection, or a decimal point shifted by a factor of ten or a thousand, produces extremes with suspiciously round ratios to the bulk. - **Incoherent companions.** Check the other fields recorded at the same moment. If humidity, pressure and battery voltage are simultaneously nonsensical, the device was malfunctioning, not the weather. - **Process metadata.** Calibration and maintenance logs, firmware changes, a battery replacement, a relocation. Extremes that begin at a known intervention and stop at the next one are a device story. - **Isolation.** A single value far outside plausible range with immediate return to normal behaviour, and no other channel reflecting anything, describes a glitch rather than an event. ## Evidence that points to a genuine extreme - **Physical plausibility.** The value is remarkable but possible for the quantity being measured. - **Corroboration.** An independent instrument, a neighbouring station, or a separate system recorded consistent conditions. This is the single strongest piece of evidence available, and it is why redundant measurement is worth its cost. - **Coherent context.** The related fields move together in a way the underlying physics or business process would predict. - **A known cause.** Someone can point to an event — a heatwave, a promotion, an incident — that explains it. - **Precedent.** The historical record contains comparable values, so the observation is rare rather than unprecedented. ## The decision procedure 1. **Use the rule as a trigger only.** A fence or robust score selects candidates for inspection. It never issues a verdict. 2. **Look at the raw records, not the summary.** Read the actual rows, with their timestamps, identifiers and neighbouring fields. Most contamination is obvious once you look at it and invisible in aggregate. 3. **Ask someone who knows the instrument.** Domain and operations staff usually recognise a device failure signature instantly. This step resolves more cases than any analysis. 4. **Flag, never silently delete.** Add a validity indicator alongside the raw value. Deleting destroys the evidence and makes the decision unauditable and unrepeatable. 5. **Run a sensitivity analysis when doubt remains.** Report the result with the disputed points included and excluded. If the conclusion is the same either way, the debate is moot and you can say so. If it flips, that is the headline finding — your conclusion rests on a handful of records whose validity is unresolved, and everyone downstream needs to know it. 6. **Prefer robust estimators while unresolved.** Rank-based summaries let you produce a usable answer without having settled every case, which buys time to settle them properly. 7. **Write the rule down.** Once you know the failure signature, encode it as an explicit validation check at ingestion with a documented rationale, so the next occurrence is caught mechanically rather than re-litigated. ## The asymmetry of the mistake Deleting genuine extremes is the more insidious error. It is invisible — nothing looks wrong afterwards, the histogram gets tidier, the fit improves — and it systematically removes exactly the tail behaviour that risk, capacity and safety analysis exists to characterise. Keeping contamination at least tends to announce itself as an implausible result. Given genuine uncertainty and no way to resolve it, retaining the point and reporting robustly is usually the safer default. ## What interviewers listen for The explicit statement that a statistical rule cannot classify a point as an error, a concrete evidence checklist rooted in the measurement process, and process discipline: flag rather than delete, keep the raw value, run it both ways, and document the decision.

  • The extreme values are all exactly identical. What does that tell you?
    Almost certainly a code rather than a measurement. Genuine physical readings vary continuously, so repeated identical extremes usually mean a rail value at the top of the instrument's register, a fixed fault code, or a sentinel such as 999 or -9999 that a loader parsed as a real number. The check is fast: count distinct values among the flagged points and compare them against the documented sentinel and error codes for that source.
  • You cannot resolve whether a handful of extremes are errors. What do you report?
    Both versions. Run the analysis with and without the disputed points and state the difference explicitly. If the conclusion holds either way, say so and move on. If it flips, that is the most important thing you know: the result depends on a few records of unresolved validity, and the honest deliverable is that dependency plus a plan to settle it, not a single number chosen quietly.
  • Why is deleting a genuine extreme worse than keeping a contaminated one?
    Because it is silent. A retained bad value tends to produce a visibly implausible result that someone questions. A deleted real value leaves a tidier distribution and a better-looking fit while quietly removing the tail behaviour that capacity planning, risk limits and safety margins are meant to characterise. The damage is invisible in every diagnostic and shows up only when the rare event recurs in production.

saying these in an interview costs you the question

  • Deletes every point flagged by an automated rule
  • Believes a statistical test can prove a value is an error
  • Overwrites raw values in place instead of adding a validity flag
  • Never checks the instrument range or the other fields on the record
  • Reports one version without disclosing that points were removed

context