skip to content

A supplier reports a 95% catch rate for an adversarial-input detector - what do you ask before believing it?

level: seniorimportance: should knowfreq 45%

answer

  1. ask which attack produced it
  2. was the attack told the detector existed
  3. a rate is not a property
  4. the missing row is the finding
  5. want a ratio, plus the false-alarm bill

basics

~20 s

Ask which attack produced it. A catch rate from attacks that did not know the detector existed shows those attacks were caught, not that the detector holds. Ask for the detector-aware number and the false-alarm rate.

solid answer

~50 s

The number tells you the attacks that were run were mostly caught; it does not tell you the detector catches crafted inputs. The first question is whether the attack was aware of the detector - if it optimised only against the mail classifier, it was never asked to satisfy the second condition, and the catch rate is measuring the wrong adversary. Then ask for the pieces that make any number comparable: what an edit budget was, how much effort each attack was given, and the same measurement run against both models jointly. The figure I actually want is the multiplier - effort to one delivered message with the detector over effort without it, under an attack that knows both are there. Finally ask what fraction of legitimate mail is quarantined at that operating point, because that is the bill the business pays for whatever the multiplier turns out to be.

code

text · 9 lines
text
detector evaluation - inbound mail (as reported)

attack                          detector-aware   edit budget   tries/msg   caught
replayed public crafted sample   no              <= 8 edits        1         97%
search against classifier only   no              <= 8 edits      400         95%
search against both models        -               -               -          not run
...
legitimate mail quarantined at this threshold: not reported
effort multiplier (with detector / without): not reported

go deeper

for a junior

Know to ask what was attacked and how, before repeating any catch rate. A number without the attack that produced it is not yet information about the detector.

for a middle

Explain why an attack that targeted only the classifier cannot measure the detector, and name the columns a comparable result needs: models attacked, edit budget, effort given, outcome.

for a senior

Show you would demand the joint detector-aware row, convert the result into an effort multiplier at a stated budget, and put the legitimate-mail quarantine rate beside it before anyone quotes the defence.

for a principal

Decide what the organisation may claim on this evidence: a priced increase in attacker effort with a stated false-alarm cost, never a statement that crafted inputs no longer reach the classifier.

### What a catch rate can and cannot establish A catch rate is a result about a specific attack run under specific conditions. Reported as "95% of adversarial inputs caught", it establishes exactly one thing: 95 percent of the messages produced by *that procedure* were flagged. It does not establish a property of the detector, and it certainly does not establish a property of the pipeline, because the adversary you care about gets to choose their procedure after reading the datasheet. The specific defect this leaf is about: if the messages were produced by an attack aimed only at the mail classifier, the detector was tested against an adversary who was never asked to satisfy it. That adversary does not exist once the detector is deployed and known. The measurement is not slightly optimistic; it is measuring a different threat model from the one the deployment faces. ### The columns that turn a number into a claim A report you can act on names, for every row: which models the attack optimised against, the budget it worked inside (for text, how many behaviour-preserving edits, and what counts as behaviour-preserving), how much effort it was given (candidates tried, submissions made), and what fraction of messages reached delivery. Two rows are comparable only when everything but one column matches. The row that is almost always absent is the joint one - an attack that knows both models are present. Its absence is the finding, and asking for it is the whole senior move here. ### The number that is worth having Once the joint row exists, the detector's value is a ratio, not a rate: effort to get one message delivered with the detector in place, divided by effort without it, at a stated budget. If it is fifteen, that is a real and quotable result - the detector made this attack fifteen times more expensive. If it is 1.2, the second model is being served, monitored and licensed for very little. Either answer is useful; the catch rate against an unaware attack is useful for neither. ### The cost column nobody volunteers A detector's operating point sets two things at once, and only one of them appears in a datasheet. At the threshold that produced the catch rate, some share of legitimate mail is quarantined. That share is not spread evenly: it lands on unusual-but-legitimate traffic - automated notifications, unfamiliar languages, odd formatting - which is precisely the mail whose loss is hardest to notice and most annoying to the affected user. Any comparison of detector settings that reports catch rate without the corresponding false-alarm rate is comparing points on a curve while hiding the axis. ### Direction of the claim, stated carefully Say these in the right direction and the rest of the conversation stays honest. A high catch rate proves the attack that was run was mostly caught, not that the detector catches crafted inputs. A clean result on a public sample of crafted messages bounds the sample, not the space. Flat aggregate accuracy on ordinary mail proves the detector is not disrupting normal traffic, not that it is contributing security. And a good result obtained while the attacker was assumed ignorant of the detector is a measurement of your obscurity, which is why an evaluation should grant the adversary knowledge of the defence on purpose. ### What to do with the answer If the joint number does not exist, the correct statement to the person deciding is that the detector's contribution is unmeasured, not that it is zero - it certainly stops unadapted traffic replaying public samples, and that has value. If the joint number exists and the multiplier is meaningful, the claim to make is a price claim with a budget attached, and never a claim that crafted inputs no longer reach the classifier. The distinction survives contact with an adversary; the catch rate does not.

  • They answer that a real attacker would not know the detector is there. What do you say?
    That the value depending on that assumption is obscurity, not robustness, and it should not be counted in a security claim. A bought-in detector is knowable to anyone who can buy or download it, and deployments announce their screens in product pages and error messages. An evaluation should grant the adversary knowledge of the defence deliberately, so the result measures the defence rather than the secret.
  • What single number would you accept in place of the catch rate?
    The effort multiplier under a detector-aware attack at a stated edit budget: candidates tried or submissions made to get one message delivered with the detector, over the same figure without it. Quoted with the budget definition and the share of legitimate mail quarantined at that operating point, it is a claim someone can act on and compare across settings.
  • The joint attack succeeds only once in five runs. Is the detector working?
    It means the search is unreliable at that budget, which is a cost result and worth recording as one - but a one-in-five reproduction is a working attack with a retry loop, not a defence. Report it as effort per success, note the variance, and resist closing it as unreproducible; the adversary is not billed for failed attempts the way your triage queue is.

saying these in an interview costs you the question

  • Accepts a catch rate without asking which attack produced it
  • Reads a high catch rate as a property of the detector
  • Ignores that the attack never targeted the detector
  • Compares two rows with different budgets or effort
  • Reports catch rate with no false-alarm rate on real mail

context