skip to content

An engineer says stealing our malware classifier needs a corpus as big as our training set - what is wrong with that?

level: seniorimportance: should knowfreq 52%

answer

  1. two resources, only one is metered
  2. the inputs are free, the answers are not
  3. the copy inherits the target's mistakes
  4. labels matter only where the verdict flips

basics

~10 s

It confuses free inputs with paid labels. In-domain executables cost nothing and the service itself supplies the labels - and labels are only needed where the boundary is, not everywhere the data is.

solid answer

~50 s

It conflates two different resources. An attacker needs inputs and labels, and only the labels are metered. In-domain executables are free and plentiful; the endpoint supplies the labels, including the target's own mistakes, so nobody has to reproduce the original corpus or its ground truth. And labels are only worth buying where the boundary is: in regions where the copy already predicts what the service will say, another paid verdict changes nothing in it. Allocating submissions toward inputs whose verdict the partly-fitted copy cannot guess reaches usable agreement with a small fraction of the target's training-set size - commonly orders of magnitude fewer paid labels than rows in that corpus. What the objection gets right is narrower than it sounds: the copy agrees only where the attacker's inputs came from, so a corpus with a blind spot yields a copy with the same blind spot.

go deeper

for a junior

Know that an attacker copying a hosted classifier does not need its training data: the files are free and the service answers their questions for them, one paid submission at a time.

for a middle

Explain the split between free inputs and metered labels, and why a verdict the copy could already predict is worth far less than one where the target's answer flips.

for a senior

Push back on 'they would need our dataset' with the economics: what a realistic submission count buys, and which narrower part of the objection - coverage - is genuinely true and testable.

for a principal

Own the consequence: corpus size is not a moat, so the real questions become what the endpoint returns, what you can detect in query patterns, and what you would do about a copy you found.

## The objection, and why it feels right "We spent four years collecting and labelling this corpus; anyone copying us would have to do the same." It is a natural thing to say and it is one of the most common wrong answers a competent senior engineer gives about a hosted classifier. It is wrong because it prices the wrong thing. ## Two resources, and the attacker only pays for one Building a classifier needs pairs of (input, label). The objection quietly assumes an attacker must reproduce *both* halves at the original scale. They do not. - **Inputs.** Executables are one of the most abundant public artefacts in existence - project releases, packaged installers, binaries shipped with any operating system, plus whatever already passes through the attacker's own mail gateway. Cost: zero. - **Labels.** The target sells exactly this, one submission at a time. The attacker rents the labelling function they would otherwise have to build. And there is a subtlety that makes the attacker's job strictly easier than the defender's original job: they are not chasing ground truth. The copy is fitted to the *service's outputs*. Where the service is wrong, a faithful copy is wrong in the same way, and that counts as success. All the expensive parts of the original effort - analyst hours, detonation, adjudicating disagreements, correcting labels - are simply not on the attacker's bill. ## Labels are only worth buying near the boundary The second half of the objection is the more interesting error. Even granting that labels are the metered resource, the assumption is that you need them *everywhere the data is*. You do not - you need them where the copy's answer is not already determined. Think about what a paid verdict does. It constrains the copy in the neighbourhood of the submitted file. If the copy already predicts that verdict confidently, the constraint was already satisfied, and the money bought a confirmation. The verdicts that change a fitted copy are the ones it could not have guessed, and those cluster where the target's decision changes from one answer to the other. This is why the size of the *original* corpus is the wrong yardstick. The original corpus had to establish what each class looks like from scratch, densely, across the whole space. The copy only has to reproduce a decision that already exists, and a decision is characterised by where it flips. ## What that does to the bill Put concretely, for a red-teamer writing an estimate before the engagement: | allocation | what it costs, roughly | | --- | --- | | uniformly random inputs | most of the spend buys predictable answers | | a free public in-domain corpus | the strong baseline: every label lands where the target's verdict genuinely varies | | inputs the partly-fitted copies disagree about | the fewest labels for a given agreement, plus the extra submissions spent locating them | The practical statement to make in an interview is directional, not a magic number: boundary-directed allocation reaches usable agreement with a small fraction of the target's training-set size, and the gap between a well-chosen allocation and a careless one at the same spend is routinely an order of magnitude in agreement. Quoting a precise ratio would be dishonest - it depends on the task, the number of classes and how tangled the boundary is - but the ordering is stable. ## What the attacker does not get, and how it bites With a top-1 verdict and nothing else, the attacker cannot read how close a file sits to a boundary. There is no score to sort by. They infer it instead from disagreement - inputs on which their own successive copies, or two copies fitted differently, return different verdicts. That inference costs submissions of its own, so the boundary-directed allocation carries an overhead that a score-returning endpoint would not impose. It is a price rise, not a barrier, and that distinction is the one to make explicitly. ## The part of the objection that is genuinely true Do not overshoot into "the corpus is worthless". The honest version of the defender's point is about **coverage**: a copy agrees with the target only where the attacker's inputs came from. If a class of files is rare or absent from public sources, the copy is blind there whatever its headline agreement says, and the original corpus's real value is exactly the regions nobody else can easily sample. That is a claim about specific regions, and it is testable - which makes it a much better thing to say than "they would need a dataset as large as ours". ## The rest of the reflex Two related reflexes usually arrive with this one and should go the same way. "We rate-limit" - a limit stretches a fixed submission count over more time or more accounts and raises the price; it does not change what a well-chosen label is worth. "They cannot reach our weights" - true, and beside the point: an attacker who wanted a functional stand-in never needed them.

  • What part of that objection is actually right?
    The coverage part. A copy agrees only where the attacker's inputs came from, so if a region of file space is rare or absent from public sources, the copy is blind there no matter what its aggregate agreement figure says. That is the corpus's real value, and it is a specific, testable claim about regions - unlike the size argument, which does not survive contact with a free public corpus plus a paid labelling oracle.
  • With only a top-1 verdict returned, how does the attacker find inputs near the boundary at all?
    Not by reading a score - there is none. They infer it from disagreement: files on which two of their own fitted copies, or successive versions of one, return different verdicts are the ones whose paid label is not already predictable. Locating those costs submissions the attacker would not spend against a score-returning endpoint, so withholding scores raises the bill without removing the allocation strategy.
  • Does the attacker need the target's ground-truth labels?
    No, and wanting them would be a mistake. The copy is fitted to the service's outputs, so where the service is wrong the copy should be wrong identically - that is what fidelity means here. Every expensive part of the original labelling effort, from analyst adjudication to correcting mislabels, is off the attacker's bill entirely.

saying these in an interview costs you the question

  • Offers dataset size as a security argument
  • Confuses free unlabelled inputs with paid labels
  • Assumes every region of input space needs labels
  • Treats a rate limit as a boundary rather than a price
  • Says the attacker needs ground truth as well as verdicts

context