skip to content

A supplier's backdoor-scan report says 'no anomaly detected' - what do you ask for next?

level: seniorimportance: should knowfreq 38%

answer

  1. read the coverage, not the verdict
  2. which family, which cap, which classes
  3. an underpowered run reports nothing too
  4. who chose the clean reference inputs
  5. detection rate on planted controls

basics

~20 s

Ask for the coverage fields, not the verdict: which trigger family and size cap were searched, which classes and at what per-class budget, whose clean inputs were used, and the tool's detection rate on planted controls.

solid answer

~50 s

Read the report as a coverage statement. Four things decide what it is worth. **Family and cap**: which shape of key was searched and to what size - anything outside it was never examined. **Completeness and budget**: all classes or a sample, and how many optimisation steps each got, since an underpowered run and a clean model produce the same output. **Reference inputs**: several methods measure a candidate key against clean data, and if the supplier chose that data they influenced the baseline, so ask for a run on inputs you collected. **Sensitivity**: a detection rate on planted controls of the family it claims to search - without a known true-positive rate, a negative is not a measurement. Then treat the residual as residual: acceptance evaluation on your own labelled inputs, sampled human re-check of production decisions, and limits on what one model verdict can do unchecked.

code

text · 10 lines
text
scan_summary:            no anomaly detected
method:                  per-class trigger reconstruction
search_family:           additive static patch, contiguous
max_patch_area:          4% of input
classes_scanned:         12 of 12
steps_per_class:         500
reference_inputs:        200 clean samples supplied by vendor
training_data_available: no
detection_rate_on_planted_controls: not reported
...

go deeper

for a junior

Know that a scan report has a scope: which key shapes were searched and which classes were covered. The verdict line alone does not tell you what was checked.

for a middle

Be able to explain why an underpowered run and a clean model produce the same summary line, and why the searched family and size cap define what the negative actually covers.

for a senior

Demonstrate the full read: family, cap, class coverage, per-class budget, whose reference data, sensitivity on planted controls - then name the compensating controls you would fund for the residual.

for a principal

Decide what the organisation is allowed to conclude from a supplier's scan report, and set the standing requirement: which fields must be present before a report is accepted into a review file at all.

## Why the verdict line is the least informative part A scan report has one line everybody reads and several lines that decide what that line means. The party who published the checkpoint chose the key after reading the same detection literature the tool implements, so the only defensible reading of a negative is: *no key of this shape, at this budget, on these classes, against this reference data*. Everything you need to say that is in the coverage fields, and it is routinely absent from the summary a supplier forwards. ## The four things to ask for **1. The searched family and its cap.** What shape of key did the method look for, and how large was it allowed to be? A search bounded to a small contiguous region of the input never examined a condition spread across the whole input or one that varies with the input. This is not a criticism of the tool; it is its specification. Get it in writing, because it is the actual scope of the negative. **2. Completeness and per-class budget.** Were all output classes scanned, or a subset chosen for cost? How much optimisation did each class get? Reconstruction-style methods will report nothing when under-resourced, and an underpowered run is indistinguishable from a clean model in the summary line. If the report shows classes scanned but no per-class budget, you are reading an unmeasured negative. **3. Whose reference data was used.** Several method families evaluate a candidate key relative to a set of clean inputs - to establish what normal activation or normal confidence looks like. Whoever supplies that data influences the baseline, and the supplier is exactly the party whose artefact is in question. Ask for a run against inputs you collected and labelled yourself. On an industrial line that means parts off your own line, in your own lighting, not the vendor's sample set. **4. Sensitivity on planted controls.** Does this tool, at this budget, on this architecture and this number of classes, actually find keys of the family it claims to search? Without a stated true-positive rate on deliberately planted controls, a clean result and a misconfigured run look the same. A negative with no known detection rate is not a measurement, and this is the single question that most often has no answer. A useful secondary question: was the same tool run twice on the same artefact, or two different tools? Two independent tools widen the union of families searched, but the overlap is large because most published methods share the small, static, local premise. It is a wider net of the same weave. ## What you do with what is left The residual after a well-documented clean scan is: a conditional whose shape lies outside the searched family. Nothing in the report reduces it, so the controls have to be ones that do not require knowing the key: - **Acceptance evaluation on your own inputs.** Data you collected and labelled, covering the decision boundary you actually care about, including the classes where a wrong call is expensive. - **Sampled re-check in production.** A human or an independent check on a fraction of decisions, weighted toward the direction that costs you most - on a pass/fail inspection line, that is the passes. - **Distribution monitoring.** A conditional that fires in the field changes the decision mix on inputs from a particular source or period, even when nobody knows what the key is. - **Bounding the blast radius.** Decide what a single model verdict is allowed to do alone. If a pass ships a part with no further check, the model is a single point of failure regardless of any scan. ## Saying it in the room Interviewers are listening for whether you know that a negative has a scope and whether you ask for it. The strongest short answer names the coverage fields, states plainly that a clean scan on a supplier artefact bounds trigger shape rather than model behaviour, and moves immediately to controls that do not depend on knowing the key. The weakest answer accepts the summary line and files it.

  • The vendor supplied the clean reference inputs. Why does that matter?
    Several method families judge a candidate key relative to what clean data looks like, so whoever chooses that data sets the baseline. The party supplying it is the party whose artefact is under review, which is a conflict of interest even without any deliberate selection. Ask for a re-run against inputs you collected yourself, covering the conditions the model will actually see.
  • The report has no detection rate on planted controls. What does that cost you?
    You lose the ability to distinguish a clean model from a run that could not have found anything. Without a true-positive rate for the family the tool claims to search, at that budget and on that architecture, the negative has no sensitivity attached, and a negative with unknown sensitivity is not evidence. It is the field most often missing and the one worth insisting on.
  • Two independent tools both came back clean. How much does that add?
    The union of the families they search, minus the overlap - and the overlap is large, because most published methods share the premise of a small, static, input-agnostic key. You get a modestly wider negative and no change in kind: the residual is still a conditional whose shape nobody searched for. Report it as two families cleared, not as corroboration.
  • What would make you refuse the checkpoint outright rather than compensate?
    When the decision it makes cannot be compensated: no independent check downstream, a cost of a wrong pass that you cannot absorb, and no way to run acceptance evaluation on inputs you control. At that point the honest options are training the model yourself on data you own, or keeping a human in the loop on every decision until you can.

saying these in an interview costs you the question

  • Accepts the summary line without asking its scope
  • Never asks who supplied the clean reference inputs
  • Ignores the per-class compute the scan was given
  • Treats two overlapping tools as independent corroboration
  • Files the clean scan as an acceptance criterion

context