Your mail detector is an off-the-shelf model anyone can download - what does that do to its value?
answer
- the multiplier assumed they had to learn it
- downloadable weights are free attempts
- only the unknowable half still charges
- labels only removes one search family
- solve once, reuse at every customer
basics
~20 sIt removes most of the cost the detector was meant to add. Anyone can solve that half offline for free, so only the classifier still charges the attacker, and one solve serves every deployment that bought it.
solid answer
~50 sThe detector's contribution was an effort multiplier, and multipliers assume the attacker has to learn the model by interacting with it. If the detector is a third-party model whose weights are downloadable, that half of the problem is solved offline at no marginal cost and with unlimited attempts - the attacker only pays for the half they cannot hold, the proprietary classifier, which here returns nothing but deliver or quarantine. That label-only reply does remove one family of search - the kind that estimates a direction by differencing returned scores - but not the kind that needs only the verdict, so it raises the bill rather than closing the door. Worse for the buyer, a shared detector means the offline work amortises: whoever solves it once has solved it for every deployment running the same product, so the multiplier you priced is divided across a whole customer base.
go deeper
Know that a defence anyone can download can be studied offline for free, so it adds far less cost than one an attacker has to learn by interacting with your system.
Explain which half of the pipeline still charges the attacker and why: the classifier they can only query, and only for the search families that need more than the final verdict.
Show you would re-derive the effort multiplier under the assumption that the detector is fully known, and that you expect a shared component's value to decay as others adapt to it.
Own the buy-versus-build argument in the right terms - correlated, decaying value from a shared component against weaker but uncorrelated value from a bespoke one - and set what the deployment may claim accordingly.
### The multiplier depends on what the attacker must buy A detector in front of a mail classifier is worth the extra effort it forces. That effort has two parts: figuring out what the detector does, and finding a message that satisfies it while still fooling the classifier. The first part is where most of the assumed cost lives - if the attacker has to infer the detector's behaviour by sending messages and watching what happens, they are paying in submissions, time and exposure. An off-the-shelf detector removes that part. If it is a product other people can license or a published model file anyone can fetch, the attacker holds it: unlimited evaluation, no rate limit, no logging, no risk, and full visibility into how it responds. The detector half of the composition becomes free to solve. ### The asymmetry that is left What remains is the proprietary classifier, and the vantage there is genuinely narrow: it returns a verdict, deliver or quarantine, and nothing else. It is worth being precise about what that narrowness does. Search that works by probing and differencing returned confidence numbers is gone - there are no numbers. Search that needs only the returned decision is not: an attacker who has a message that gets delivered and a message that gets quarantined can work along the boundary between them using verdicts alone. So withholding scores raises the price of the remaining half; it does not remove it. This is the same shape as the detector itself - a cost control being described as a boundary. The practical consequence is that the attacker's whole budget concentrates on the one model they must interact with, having already fixed the detector condition offline. The composition is only as expensive as its knowable half is cheap. ### Amortisation across the customer base The second effect is the one buyers rarely price. A bespoke detector, trained on your own traffic, has to be solved per deployment. A bought-in one is the same model at every customer that bought it. The offline work of finding messages the shared detector reads as ordinary is done once and reused everywhere, so the per-victim share of that cost tends toward nothing as the product's install base grows. Popularity, which is the reason the product looked like a safe purchase, is also what makes the solution reusable. That does not make a shared detector a bad purchase. It makes its value time-dependent and correlated: it works well against traffic that has not adapted, and its multiplier for you drops the moment anyone adapts to it, whether or not that person has ever targeted your deployment. A bespoke component has weaker average quality and an uncorrelated failure mode; that is the real tradeoff, and it should be argued on those terms rather than on catch rate. ### How to say it when someone asks A reviewer asked whether the downloadable detector still counts should answer with three statements. First, the security value that remains is the price charged on the half the attacker cannot hold offline - here, the label-only classifier. Second, that price is real but smaller than the composition's headline suggests, and it is not multiplied by the detector's accuracy. Third, whatever multiplier is measured is shared with every other deployment of the same product, so it should be expected to decay rather than hold. ### The direction of every claim here Downloadable weights do not make the classifier's boundary public - the attacker still has to work for that half, and transfer between the two models is not automatic. Equally, a private classifier does not make the composition private. And a detector's catch rate measured while the evaluator assumed the attacker could not inspect it is measuring secrecy that the product's own distribution model has already given away.
- Does the classifier returning only deliver or quarantine, with no score, close the remaining half?No. It removes search that estimates a direction by differencing returned confidence numbers, because there are none to difference. Search that needs only the verdict still works: starting from a message that gets delivered, an attacker can move toward the one they want using decisions alone. Withholding scores is a price increase on the remaining half, and it should be described as one.
- Is a detector you trained yourself on your own mail better, then?On this axis, yes - it has to be solved per deployment, so the attacker's offline work does not amortise across other customers. On other axes it is usually worse: less data, less maintenance, and quality that decays as traffic shifts. The argument is uncorrelated failure versus average quality, and it should be made in those terms rather than on catch rate.
- How should this change what the deployment claims about crafted mail?The claim stays a price claim and gets a decay note. Say that the screen raises the cost of a delivered crafted message by a measured factor at a stated edit budget, that the factor is shared with every deployment of the same bought-in model, and that it should be re-measured rather than assumed to hold. Do not claim crafted inputs are prevented from reaching the classifier.
saying these in an interview costs you the question
- Assumes the attacker must probe the detector to learn it
- Treats a downloadable detector as adding the same multiplier
- Thinks public detector weights expose the classifier's boundary
- Says label-only replies stop all black-box search
- Ignores that one solve is reused at every customer