skip to content

Model Extraction & Stealing

You will learn how an attacker clones a paywalled model through nothing but prediction queries, and what it costs them relative to training from scratch. Interviewers use it to test whether you can reason about an ML API as an economic attack surface and defend it with rate limits, output hardening, and watermark-based ownership proofs.

on this pageshow

explore

questions

page 2 of 2

What does recovering a model's weights 'only up to symmetry' leave an attacker holding?

level: middleimportance: nice to knowfreq 26%

basics

~20 s

A set of parameters that computes exactly the same function as yours, but need not match yours entry by entry: hidden units can be permuted and a unit's incoming and outgoing weights rescaled against each other without changing any reply.

open as a page

Why must a latency gap measured against a hosted embedding endpoint be repeated before it means anything?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

A single call's latency is dominated by network jitter, queueing and co-tenant batching, which swamp any model-shape effect. Only a difference that survives averaging many interleaved calls carries information, and confounds that move with load never average away at all.

open as a page

A report claims 94% agreement between an extracted copy and your code endpoint — what do you ask?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Ask what agreement was measured against, on which prompt distribution, how a text match was defined, and what the un-queried base model scored. One agreement rate describes an evaluation set, not the copy — and it establishes nothing about weights.

open as a page

An extraction run reports 97% agreement with a paid malware-verdict API - what do you ask before believing it?

level: seniorimportance: nice to knowfreq 30%

basics

~10 s

Ask which pool agreement was measured on and the target's base rate there. On a pool where the service answers benign 97% of the time, a copy that always says benign also scores 97%.

open as a page

Product wants an 80% per-call price cut on your model API — what does that hand a cloner?

level: principalimportance: nice to knowfreq 26%

basics

~10 s

It divides one side of a copier's build-or-steal ledger by five while leaving their honest alternative untouched, so an endpoint nobody would copy at the old price can start paying for itself.

open as a page

A scoring API's owner asks which anti-extraction control still binds next quarter - what do you say?

level: principalimportance: nice to knowfreq 27%

basics

~20 s

None is a boundary; you are picking prices. Per-key caps are amortised away by cheap identities, coarsening taxes your own bidders, and a distribution signal fades once imitated. Only identity friction raises an unspreadable cost.

open as a page

Your face matcher ships to licensees next quarter: do you fund a planted ownership mark in that training run?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Usually no. A planted mark costs accuracy on the licensee benchmark, commits the release, and leaves an unrotatable secret. Recognising behaviour the model already has is free and decidable later — buy the mark only if you would act on it.

open as a page

You inherited a marked detector and the suspect copy was re-distilled — what is that ownership evidence worth?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Less than the team hopes, and say so plainly: a decayed mark read by a test with an unmeasured error rate is a lead, not evidence. The call is whether to fund the control run or drop the claim.

open as a page

Transparency commitments force per-field reasons on every model decision. What do you actually negotiate?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Not whether to explain, but the shape of the disclosure: how many fields it names, at what numeric precision, to which recipient, at what rate, and at what price. The obligation usually runs to the decision subject, one record at a time, not to a machine integrator pulling thousands a day.

open as a page

How do you price the per-category scores in a moderation API's response at a design review?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Report what the field buys an adversary in money and time: the annotation budget a competitor no longer funds, and how fast a stand-in reaches usable quality. Leave the keep-or-cut call to the product owner.

open as a page

showing 31–40 of 40