skip to content

How much assurance can you claim from a backdoor scanner a year after its method went public?

level: principalimportance: nice to knowfreq 26%

answer

  1. the decay is not uniform across adversaries
  2. cheap and careless keys are still caught
  3. weakest where the artefact is worth attacking
  4. decide the sentence that leaves the room
  5. spend on controls that need no key

basics

~20 s

Less against a supplier who read it, but not nothing. A public method still prices out careless and copied artefacts; its negative decays exactly where you are targeted, so treat it as an entry bar, never an acceptance criterion.

solid answer

~50 s

The value of a published detector decays in one specific direction: against a party who chose the key after reading it. It does not decay against the careless supplier, the reused off-the-shelf artefact, or the accidental case, so keep running it - it is cheap and it produces a documented negative on a named family. The organisational call is what claim leaves your team. Write the scoped sentence into the review file - no key of the searched family, at this budget, on these classes - and refuse to let it become "the model is clean". Then fund the controls that do not depend on knowing the key: acceptance evaluation on inputs you collected, a sampled independent re-check weighted toward the costly direction, and a limit on what one model verdict does unchecked. And note the perverse shape: the negative is weakest precisely where the artefact is worth attacking.

go deeper

for a junior

Know that a scan result gets weaker over time against suppliers who have read the method, and that it is a step in a process rather than a sign-off.

for a middle

Be able to explain why the decay is directional: cheap and copied keys are still caught, while a supplier who designed against the published method is not.

for a senior

Show that you would keep the scan, record the scoped negative, and spend the next increment on acceptance testing and sampled re-checks rather than on more tools of the same kind.

for a principal

Own the wording that leaves your team and the budget behind it. Be ready to tell a stakeholder that no scan answers 'is it backdoored', and to defend a scoped negative plus a funded compensating control as the correct deliverable.

## The decay is directional, and that is the whole judgment A detection method's assumptions become public the day it is published. From then on, any party choosing a key for a checkpoint they intend to hand over can read them. That is not a defect of the tool and it is not fixed by a newer tool, which will also be published. So the honest model of a scanner's value over time is not "it degrades" but **it degrades against one class of adversary and not against the others**: | Who you are facing | What a public scan is still worth | | --- | --- | | A supplier who never thought about this | Nearly its original value - cheap keys are found | | A reused or copied artefact from elsewhere | High - off-the-shelf conditionals sit inside the searched family | | An accidental correlation learned from data | Some - it can surface a class reachable too easily | | A supplier who read the method and is targeting you | Little, and it falls further each quarter | The uncomfortable consequence, and the sentence worth saying out loud in a principal interview: **the assurance is weakest exactly where the artefact is worth attacking.** A checkpoint on a low-stakes internal task is well covered; the one running a high-value production decision is the one an adversary would bother to design against. ## The claim you let leave the team The deliverable of a review is a sentence somebody else will act on, often months later and without the context. The failure mode this leaf exists to prevent is a bounded negative turning into an assurance claim on the way out of the room: "scanned clean" becomes "verified", and by the time it reaches a risk register the family, the cap, the class coverage and the budget are gone. So set the wording as policy, not as a habit: - what goes in the file is the scoped negative, with the family and the budget attached; - "no backdoor" is not a conclusion any scan can produce, and nobody may write it; - the compensating control and its rate are recorded next to the negative, because the negative alone is not an acceptance decision. This is genuinely an organisational call rather than a technical one: you are deciding what your team is permitted to assert, and you will be arguing it with someone who wants a yes. ## Where the next unit of budget goes Given that the scan's coverage is fixed and public, additional spend on the same axis buys little. The controls that keep working are the ones that never needed to know the key: - **Acceptance evaluation on data you own** - collected under your conditions, labelled by you, weighted toward the decisions that hurt when they are wrong. - **A sampled independent re-check in production**, sized by consequence rather than by convenience, and biased toward the direction of loss. On a pass/fail line, sampling the passes is what matters. - **Monitoring the decision distribution** by source, batch, and period. A conditional that fires in the field moves the mix even when nobody can name the key. - **Reducing what one verdict does alone.** If a model decision has an independent check behind it, a hidden conditional degrades a metric instead of causing the loss directly. - **Not accepting the artefact** for the small set of decisions where none of the above is affordable - training on data you control, or keeping a person in the loop. ## Two traps in the periodic-review version of this question First, **re-running the same tool on the same unchanged artefact next quarter adds nothing**. It produces the same bounded negative. What would add something is a method whose searched family is genuinely different, which is a procurement question, not a scheduling one. Second, **stacking tools is not corroboration**. Two published methods that both assume a small, static, local key overlap heavily, so the union of coverage grows far less than the count of tools suggests. Report families cleared, not tools run. ## The answer to the stakeholder who wants a yes or no Give a scoped yes and the control that covers the rest: no key of the searched family was found, at this budget, across these classes, against reference data we supplied; the residual is a conditional of a different shape, and it is covered by an acceptance set we own and a sampled re-check at a stated rate. Saying that a scan cannot answer "is it backdoored" is not evasion - it is the accurate scope of the measurement, and owning that distinction is what the level is being tested on.

  • Does this argument mean you should stop scanning entirely?
    No. It is cheap, it catches copied and careless artefacts, and it forces a determined author into a costlier and less reliable key. A documented negative on a named family is a real fact for a review file. What you stop is quoting it as clearance and letting it stand in for acceptance testing - the practice changes, not the tool.
  • How do you answer a risk owner who insists on a yes or no?
    Give a scoped yes plus the compensating control: no key of the searched family, at this budget, on these classes, and here is the acceptance set and the sampling rate that covers the rest. Then say plainly that no scan can produce a verdict on the model itself, so a bare yes would be a claim nobody measured. Offering the scope is the answer, not a refusal to answer.
  • Would scheduling a quarterly re-scan of the same checkpoint help?
    Not on its own. The same search over an unchanged artefact returns the same bounded negative. What changes coverage is a method whose searched family is different - a procurement decision - or a change to the artefact itself, which resets the acceptance evaluation anyway. Scheduling repetition mostly buys the appearance of diligence.
  • Where is the perverse incentive in relying on published scanners?
    Coverage is best on artefacts nobody bothered to attack and worst on the ones worth attacking, because the second group's authors have read the method. A programme that measures itself by scans passed will therefore look strongest precisely where its assurance is weakest, which is why the acceptance decision has to rest on controls that do not depend on knowing the key.

saying these in an interview costs you the question

  • Treats a scan pass as an acceptance criterion
  • Lets a scoped negative leave the room as 'verified'
  • Assumes newer tools close the gap permanently
  • Counts tools run instead of families covered
  • Schedules re-scans of an unchanged artefact as diligence

context