skip to content

How would you turn an autoencoder's reconstruction error into a machine anomaly score?

level: seniorimportance: should knowfreq 45%

answer

  1. fit normal, score the residual
  2. what gets into the training set matters
  3. capacity decides whether anything separates
  4. threshold from held-out normal error
  5. widen the code and the alarms stop

basics

~20 s

Train the autoencoder on normal data only, score each window by its reconstruction error, and threshold at a high quantile of the error measured on held-out normal data. Keep the code narrow enough that the model cannot rebuild behaviour it never saw.

solid answer

~50 s

The premise is that a narrow autoencoder fitted to normal operation has spent its capacity on normal structure, so anything unlike normal reconstructs badly. Concretely, take 128-channel turbine vibration windows compressed to an 8-unit code: standardise each channel first, or the loudest sensor dominates the squared-error sum and the score becomes a proxy for that one channel. Train on windows you are confident are healthy — any failure that leaks into training gets reconstructed well and will never be flagged. Set the threshold from a high quantile of reconstruction error on held-out normal windows, chosen against an alarm budget, never from training error, which is optimistically low. The critical knob is code width: widen 8 units to 64 and the decoder starts rebuilding unseen behaviour too, so the score separates nothing. Recalibrate as the machine ages, and decompose the error per channel to tell the operator where the deviation is.

go deeper

for a junior

Know the basic recipe: train only on normal data, score each window by how badly it is reconstructed, and raise an alarm when that error crosses a threshold.

for a middle

Explain why a wider code weakens the detector, why channels must be scaled before a squared-error score, and why the threshold has to come from data the model did not train on.

for a senior

Demonstrate that you have operated one of these: curating the normal training set, selecting code width by separation rather than loss, converting a quantile into an alarm budget, and recalibrating under drift.

for a principal

Own the decision of whether this detector belongs in the system at all — the false-alarm cost to the operations team, the baseline it must beat, and the change-control story for a threshold that will need to move over the asset's life.

## The assumption you are making Reconstruction-based anomaly detection rests on one claim: a model with limited capacity, fitted only to normal data, reconstructs normal inputs well and abnormal ones badly. Every design decision below is really about protecting that claim, and every failure mode is a way the claim quietly stops being true. A worked setting: a turbine instrumented with 128 vibration channels, sampled into fixed-length windows, each window compressed to an 8-unit code and decoded back. The per-window score is the mean squared difference between the window and its reconstruction. ## Getting the training set right Train on normal-only data. This is the requirement people violate first, usually by pointing the job at 'all historical data' because that is what the warehouse holds. If the two known failure episodes sit in the training set, the model allocates capacity to reconstructing them, their error drops to normal levels, and the detector is blind to exactly the events it exists to catch. Curate: exclude known failure windows, exclude maintenance and shutdown periods unless you want those treated as normal, and be explicit about which operating regimes are represented. A regime the model never saw — a colder season, a new load profile — will produce high error and a stream of false alarms that look like a fault. ## Scaling, or the loudest channel wins A squared-error score sums over all 128 channels in their raw units. A channel that swings over hundreds contributes orders of magnitude more than one that swings over hundredths, so both the training objective and the score are effectively about the loud channels. Standardise each channel using statistics computed on the training window set only, or attach deliberate per-channel weights that encode which sensors you actually care about. Either way, make the choice consciously — leaving it implicit means the physics of your alarm is decided by sensor gain. ## Capacity is the whole game The most instructive failure: someone widens the code from 8 units to 64 to 'improve reconstruction', watches the training loss fall, and ships it. The detector then fires almost never. With enough capacity, the decoder learns to rebuild a broad range of inputs including patterns it was not trained on, so the error gap between normal and abnormal windows collapses. This is the same mechanism that makes an overcomplete autoencoder learn the identity, and here it destroys a production system rather than merely wasting a training run. The consequence is that you cannot select the code width by reconstruction error, which always favours the widest model. Select it by **separation**: on a validation set that contains the few labelled abnormal windows you have, choose the width that best separates their scores from the normal ones. If you have no labelled abnormalities at all, hold the code as narrow as the normal-window error will tolerate and treat the width as a risk parameter reviewed with operations. ## Setting the threshold Use a high quantile — the 99th, say — of reconstruction error on **held-out** normal windows the model never trained on. Training-set error is systematically lower because the model fitted those windows, so a threshold derived from it sits too low and floods the operators with false alarms on the first day. Translate the quantile into an alarm budget before you pick it: at one window per minute, a 99th-percentile threshold means roughly fourteen alarms a day on healthy operation. If the team can absorb two, you need a much higher quantile, temporal aggregation (flag only when several consecutive windows exceed the threshold), or a smoothed score. Alarm rate is a product decision, not a modelling one. ## What the score is not Reconstruction error is not a probability, not a likelihood, and not calibrated. A score of 0.4 does not mean a 40 percent chance of failure, and the number is not comparable across machines, retrains or preprocessing changes. Treat it as an ordinal signal: useful for ranking windows and for exceeding a threshold you calibrated, useless as a quantity to report to an operator on its own. ## Operating it Machines drift — bearings wear, ambient conditions shift, sensors are replaced. Error on healthy operation creeps upward and the fixed threshold slowly becomes an alarm generator. Monitor the distribution of scores on presumed-normal windows and recalibrate on a schedule, with a change-control record of when the threshold moved and why, so a later investigation can reconstruct what the detector was doing at the time of an incident. Finally, give the operator attribution. The per-window score is a sum over channels and time steps, so you can rank the channels by their contribution and say 'the deviation is concentrated in these three channels near the end of the window'. That is attribution, not diagnosis — it says where the model was surprised, not what broke — but it is what turns an unexplained number into an actionable alert. ## The baseline you should always run Before any of this, check whether a simple threshold on channel amplitude, or a per-channel control limit, catches the same events. If it does, the autoencoder is added complexity with an operational burden. The model earns its place only when the anomalies are about the *relationship between channels* rather than any single channel leaving its range — which is exactly the kind of structure a bottleneck learns.

  • Your alarms stopped firing after the code was widened from 8 to 64 units — what happened?
    The extra capacity let the decoder rebuild inputs it was never trained on, so abnormal windows now reconstruct almost as well as normal ones and the score gap collapses. Training reconstruction error looks better than ever, which is why the change seemed like an improvement. Narrow the code and select width by score separation, not by reconstruction error.
  • Why shouldn't you set the threshold from reconstruction error on the training set?
    The model fitted those exact windows, so their error distribution is optimistically low and a quantile taken from it sits below the error the model produces on genuinely unseen healthy data. Deploy that threshold and normal operation exceeds it constantly. Use held-out normal windows and convert the quantile into an expected alarms-per-day figure before committing.
  • How do you tell an operator which part of the machine tripped the alarm?
    The score is a sum of squared residuals over channels and time steps, so rank the channels by their share of the total and report the top contributors with the part of the window where the residual peaked. Present it as attribution — where the model was surprised — rather than as a diagnosis of the physical fault.
  • What kind of anomaly will this approach systematically miss?
    Anything that stays inside the normal manifold: a fault whose vibration signature resembles a legitimate operating regime the model learned, or a slow drift the model was retrained on. It also misses faults present in the training data. Reconstruction error detects unfamiliarity, which is not the same as danger.

saying these in an interview costs you the question

  • Trains on all historical data including known failure windows
  • Assumes a bigger autoencoder makes a better anomaly detector
  • Reads reconstruction error as a probability of failure
  • Sets one fixed threshold and never recalibrates as the machine ages
  • Skips per-channel scaling so the loudest sensor dominates the score
  • Never checks whether a simple amplitude limit catches the same events

context