A nine-day adaptive attack on a shipped malware classifier failed - what can the red-team report honestly claim?
answer
- a bounded negative, not a proof
- say what was spent and seen
- your vantage was below the ceiling
- list the designs you abandoned
- wording: not broken by, never robust
basics
~20 sOnly that these designs, from this vantage, at this effort, did not break it. A failure to break is bounded by who tried and how hard, so the report states access, days spent and designs abandoned - never robust.
solid answer
~50 sA negative result is the weakest evidence this field produces, and the report has to carry its own bounds. Three of them matter. **Vantage**: an offline copy of the shipped engine returning a verdict and a coarse confidence band, with no derivatives through the ensemble — an adversary with weights sits strictly above you, so your failure caps your search, not the model. **Effort**: nine of twelve contracted days, and the number of formulations designed and dropped. **Coverage**: which attack designs were tried, which were abandoned and why, and which one you would have written next with more days. The claim then reads as a sentence somebody can argue with: not broken by these designs, at this access, for this effort, on this date. It is also worth saying explicitly what the result does not license the vendor to print.
go deeper
Take away the direction of the claim: not breaking something is not the same as showing it is safe. The number of days someone spent is part of the result.
Be able to name the bounds that belong beside a non-break: what the attacker could see, how much effort was spent, and which attack designs were tried. Explain why a weaker vantage caps what the failure proves.
This is your question. Show that you would write the claim as a dated, bounded sentence, that you would document abandoned formulations, and that you would push back on a customer rewriting a negative into a verification.
Decide the reporting standard before the engagement starts: what wording the firm will sign, what happens when a client wants more, and whether a bounded negative is worth buying at all for this class of product claim.
### Why the negative case is the hard one to write Breaking something is self-documenting. You produce the artifact, it reproduces, the claim is over. Failing to break something documents nothing on its own — the same empty result is produced by a strong defense, by a weak defense attacked from a poor vantage, and by a competent team that ran out of contracted days on the wrong idea. So the entire informational content of a negative result lives in the bounds you attach to it, and writing those bounds is the deliverable. ### Bound 1: the vantage you actually had In this engagement the adversary held a trial licence and an offline copy of the shipped engine, could run it locally without limit, and got back a verdict plus a coarse confidence band. No derivatives through the ensemble. That is a specific and thin vantage. Two consequences must appear in the report: - Anyone holding the weights sits *above* you. Your failure bounds what a decision-and-band adversary managed, and says nothing about what a weights-holding adversary would manage. The white-box result is the ceiling every weaker adversary sits under, and you did not measure it. - The coarse confidence band is itself a limit. Coarser feedback makes searches that rely on small changes in returned confidence much more expensive, and the report should say so, because a future product change that returns finer scores invalidates your result without anyone touching the model. ### Bound 2: the effort actually spent Effort — not a norm, not a query bill — is the unit here, because the engagement was fixed-fee and days were the scarce thing. Report the number spent, the number contracted, and how they were distributed across formulations. Nine of twelve days is a very different claim from twelve of twelve, and both are different from a colleague spending a month unpaid because the problem interested them. Be honest about the shape of the spend too. Six days on one idea that never worked, plus three days on two others, is a thinner result than nine days across five formulations, and a reader who is deciding whether to fund a follow-up needs that distinction. ### Bound 3: the designs, including the dead ones List what was designed against this defense and what was abandoned. Two reasons, and both are practical rather than ceremonial: 1. It tells the next reader where the remaining probability mass sits, so the next engagement does not re-spend the same days on the same dead ends. 2. It is what makes the bound auditable. A reader who thinks you missed the obvious formulation can say so, by name, and that argument is exactly the value the customer bought. Also state what you would have written next. A red-teamer who stopped with a live idea unexplored is telling the customer something specific about how much of the space remains, and it is dishonest to leave that out of a report that will be read as reassurance. ### The wording, and the fight about the wording The customer's incentive is to quote the engagement as *independently verified robust*. That sentence takes an effort-bounded negative and inflates it into an unbounded positive, which is precisely the inflation the whole adaptive convention exists to prevent. The counter is not to argue in the abstract but to supply the sentence you will stand behind, in a form suitable for their deck, and to say you will not stand behind rewrites of it: *Not broken by five defense-aware attack designs written against the published mechanism, from an offline copy returning verdicts and a coarse confidence band, over nine analyst-days, in [month]. An adversary with model weights, more days, or finer confidence output is not covered.* That sentence has a date, an actor, a vantage and a price. It can be re-tested. It also expires naturally, which is the property you want, because the defense's exposure changes when the product changes what it returns or when somebody publishes a better formulation. ### The useful thing you produce despite failing A report like this is not a null result to the customer if you frame it as a *price list*. You have established that a specific, realistic adversary — the one who buys a licence and reads the whitepaper — did not get there in nine funded days. That is genuinely worth knowing when the defense's job is to raise cost against exactly that adversary. What it is not is a statement about the model. Keeping those two apart in the executive summary, not just in the appendix, is the actual craft. ### The trap to avoid in the interview Do not slide into claiming your failure validates the vendor's original stock-suite number. It does not touch it. Their number came from attacks that never knew the defense existed; yours came from attacks that did. Yours is better evidence and still weak, and both being weak for different reasons is a distinction worth being able to state cleanly under pressure.
- The vendor wants to quote your engagement as independently verified robust. What do you tell them?That I will not stand behind that wording, and here is the sentence I will stand behind — naming the attack designs, the access I had, the days spent and the date. A bounded failure to break is not verification of anything; the quote silently converts my nine days into a claim about all adversaries. I would also offer to re-scope the claim as a cost statement, which is what they can actually defend.
- Why list the attack formulations you tried and abandoned rather than only the outcome?Because the strength of a non-break scales with effort and coverage, and neither is visible from an empty result. The list tells the next reader where the unexplored space is, stops a follow-up engagement re-spending days on known dead ends, and lets a critic point at the formulation I missed. Without it, my report asks to be trusted rather than checked.
- Does having no derivatives through the ensemble weaken your negative result?Yes, materially. I searched from a thinner vantage than the ceiling, so my failure bounds what a verdict-and-band adversary achieved in nine days, not what someone holding the weights would achieve. The report has to say that in the summary, because a reader who skips the methodology will otherwise treat my bound as the strongest one available.
saying these in an interview costs you the question
- Reports a failed break as the defense being robust
- Omits the vantage and effort that bound the result
- Treats the engagement's day limit as the adversary's limit
- Lets a customer requote a negative result as verification
- Records no abandoned designs, so coverage is invisible