skip to content

Which Queries to Spend

Fidelity per dollar is decided by which inputs get bought — random draws, public in-domain data, or points near the boundary. Interviewers use it against the 'they'd need our whole dataset' reflex.

on this pageshow

explore

questions

4

An attacker pays per submission to a hosted malware-verdict API to copy it - what does that budget actually buy?

level: juniorimportance: must knowfreq 60%

answer

  1. the files themselves cost nothing
  2. something else on the wire is metered
  3. the endpoint answers questions for money
  4. a budget counts labels, not rows

basics

~20 s

Labels, not data. Executable files are free and plentiful; what costs money is the service's verdict on each one. The budget is a fixed count of labels, so which files receive them decides how good the copy gets.

solid answer

~50 s

They are buying labels, not files. Unlabelled executables are abundant and free - public software releases, binaries shipped with an operating system, a mail gateway's own backlog - so the scarce resource is the target's verdict on each one. The endpoint behaves as a metered labelling oracle, and a fixed budget is simply a fixed count of labels. That reframes copying as an allocation problem rather than a data-collection problem: with, say, 50,000 submissions at a fixed unit price, the only lever the attacker has is which files receive them. Fidelity here means agreement with the target's outputs on held-out in-domain files, not accuracy against ground truth - a faithful copy reproduces the target's mistakes too. Two campaigns with identical budgets can land an order of magnitude apart on agreement purely on that allocation choice.

go deeper

for a junior

Be ready to say plainly what the attacker is paying for: the verdicts, not the files. Know that success is measured as agreeing with the target, including where the target is wrong.

for a middle

Explain why a metered endpoint behaves as a labelling oracle, and why a fixed submission count turns copying into an allocation problem rather than a data-collection one.

for a senior

Show you can price a campaign: submissions, unit price, the agreement you would commit to, and how you would draw a held-out pool so that the agreement figure survives scrutiny.

for a principal

Own the framing that per-query pricing is judged against the cost of building the model, never against the value of one answer, and be able to say what that implies for what your endpoint returns.

## The situation A hosted static-analysis service accepts an executable file and returns a single word: malicious or benign. No score, no confidence, no family name, no explanation. It is billed per submission at a fixed unit price and sold to mail gateways and managed security providers. An adversary wants their own classifier that behaves like this one - so they can run it for free, without a rate limit, and without every question they ask being logged on somebody else's side. They cannot reach the weights. What they can do is submit files and read the word that comes back. ## Two resources, only one of them metered Fitting a classifier needs pairs: an input, and a label for it. Extraction splits those two apart, and the split is the whole economics of this leaf. - **Inputs are free.** Executables are one of the most abundant public artefacts there are: open-source project releases, files shipped with operating systems, packaged installers, whatever is already flowing through the attacker's own mail gateway. Nobody has to pay for a file. - **Labels are metered.** The only thing the target sells is its opinion, one submission at a time. So a query budget is not a budget for data. It is a budget for *answers about data the attacker already has*. This is the sentence that most people miss on first encounter, and it is why "our training corpus took years to assemble" is not, on its own, a defensive fact. ## The endpoint as a labelling oracle Once you see it that way, the target is a labelling service with a price list. Note what the copy is being fitted to: **the target's outputs, not the truth.** If the service calls a particular clean installer malicious, a faithful copy calls it malicious too, and that counts as success. Ground truth is irrelevant to the exercise; an extraction campaign that improved on the target's accuracy would in fact be a *worse* copy. ## What fidelity means, and against what Fidelity is normally reported as **agreement**: the fraction of held-out inputs on which the copy returns the same verdict as the target. Two things must travel with that number or it means nothing: - **Which pool it was measured on.** Agreement on odd, off-distribution inputs and agreement on real traffic are different measurements of different things. - **The target's base rate on that pool.** If the service answers benign on 97% of the evaluation pool, a copy that always answers benign scores 97% agreement while having learned nothing at all. ## Why allocation is the whole game Each paid verdict constrains the copy in the neighbourhood of the file that received it, and nowhere else. A verdict the copy could already have predicted before paying for it constrains nothing new - the money bought a confirmation. So the value of a submission is set by how surprising the answer is, and that varies enormously across inputs: | where the submission goes | what the label is worth | | --- | --- | | uniformly random or junk inputs | almost nothing: the answer is predictable in advance | | a free public corpus of real executables | a lot: it spans the regions the target actually separates | | files whose verdict the partly-fitted copy cannot guess | the most per label, but they cost extra submissions to find | The same budget spent three ways produces very different copies. That is why the question an interviewer is really asking here is an allocation question, not a "can they steal it" question. ## The chair this is asked from The person who has to answer this in a loop is usually a red-teamer writing an engagement estimate *before* the engagement: how many submissions, at what unit price, to deliver what agreement, and against which evaluation pool. Every part of that estimate is downstream of the allocation decision, and none of it depends on ever touching the target's parameters or its corpus. ## Where it stops working - The copy agrees only where the attacker's inputs came from. A free corpus that under-represents some region of file space leaves the copy blind there, and no amount of extra budget spent elsewhere fixes it. - With a top-1 verdict and nothing else, the attacker cannot read how close an input sits to a boundary; they can only infer it from disagreement between their own successive copies. That raises the submission count relative to an endpoint that returns scores - a price rise, not a barrier. - In domains where no free in-domain corpus exists, the cheap allocation is unavailable and the bill climbs sharply.

  • What does fidelity mean for a stolen copy, and against what is it measured?
    Agreement with the target's own outputs on a held-out pool of in-domain inputs - not accuracy against ground truth. A faithful copy reproduces the target's errors, and reproducing them counts as success. The number only means something alongside the pool it was measured on and the target's class base rate on that pool, since on a heavily skewed pool a constant answer already scores high.
  • Why is per-submission pricing a weak brake on this?
    Because the attacker compares the bill against the cost of building the model themselves, not against the value of one answer. If a few tens of thousands of submissions yield a usable stand-in, that spend sits far below the cost of assembling and labelling a corpus. Metering stretches the calendar and raises the price; it does not change what a well-chosen label is worth.

It is like hiring a very expensive expert to mark exam papers. Blank paper is free; the only decision that matters is which papers you send them.

saying these in an interview costs you the question

  • Claims the attacker must first obtain the training data
  • Treats every submission as carrying the same value
  • Calls the fitted copy a stolen set of weights
  • Scores the copy against ground truth instead of agreement with the target

context

open as a page

Why does spending a malware-API query budget on random byte-strings buy far less agreement than real executables?

level: middleimportance: should knowfreq 48%

basics

~20 s

Random byte-strings sit far off the distribution the classifier was fitted on, so nearly all of them come back with the same verdict. A label you could have predicted before paying for it teaches the copy nothing.

open as a page

An engineer says stealing our malware classifier needs a corpus as big as our training set - what is wrong with that?

level: seniorimportance: should knowfreq 52%

basics

~10 s

It confuses free inputs with paid labels. In-domain executables cost nothing and the service itself supplies the labels - and labels are only needed where the boundary is, not everywhere the data is.

open as a page

An extraction run reports 97% agreement with a paid malware-verdict API - what do you ask before believing it?

level: seniorimportance: nice to knowfreq 30%

basics

~10 s

Ask which pool agreement was measured on and the target's base rate there. On a pool where the service answers benign 97% of the time, a copy that always says benign also scores 97%.

open as a page