Model Extraction & Stealing
You will learn how an attacker clones a paywalled model through nothing but prediction queries, and what it costs them relative to training from scratch. Interviewers use it to test whether you can reason about an ML API as an economic attack surface and defend it with rate limits, output hardening, and watermark-based ownership proofs.
on this pageshowhide
explore
- Signal per Reply8 questions
- Label, Score, Everything4 questions
- Explanations as Training Data4 questions
- Assembling a Copy16 questions
- Which Queries to Spend4 questions
- Behaviour, Not Weights4 questions
- Reading the Model's Shape4 questions
- Exact Parameter Recovery4 questions
- Clone Economics8 questions
- Cheaper Than Training It4 questions
- Limits That Actually Bind4 questions
- Proving It Was Yours8 questions
- Marking or Recognizing4 questions
- What Survives Distillation4 questions
questions
page 2 of 2What does recovering a model's weights 'only up to symmetry' leave an attacker holding?
basics
~20 sA set of parameters that computes exactly the same function as yours, but need not match yours entry by entry: hidden units can be permuted and a unit's incoming and outgoing weights rescaled against each other without changing any reply.
Why must a latency gap measured against a hosted embedding endpoint be repeated before it means anything?
basics
~20 sA single call's latency is dominated by network jitter, queueing and co-tenant batching, which swamp any model-shape effect. Only a difference that survives averaging many interleaved calls carries information, and confounds that move with load never average away at all.
A report claims 94% agreement between an extracted copy and your code endpoint — what do you ask?
basics
~20 sAsk what agreement was measured against, on which prompt distribution, how a text match was defined, and what the un-queried base model scored. One agreement rate describes an evaluation set, not the copy — and it establishes nothing about weights.
An extraction run reports 97% agreement with a paid malware-verdict API - what do you ask before believing it?
basics
~10 sAsk which pool agreement was measured on and the target's base rate there. On a pool where the service answers benign 97% of the time, a copy that always says benign also scores 97%.
Product wants an 80% per-call price cut on your model API — what does that hand a cloner?
basics
~10 sIt divides one side of a copier's build-or-steal ledger by five while leaving their honest alternative untouched, so an endpoint nobody would copy at the old price can start paying for itself.
A scoring API's owner asks which anti-extraction control still binds next quarter - what do you say?
basics
~20 sNone is a boundary; you are picking prices. Per-key caps are amortised away by cheap identities, coarsening taxes your own bidders, and a distribution signal fades once imitated. Only identity friction raises an unspreadable cost.
Your face matcher ships to licensees next quarter: do you fund a planted ownership mark in that training run?
basics
~20 sUsually no. A planted mark costs accuracy on the licensee benchmark, commits the release, and leaves an unrotatable secret. Recognising behaviour the model already has is free and decidable later — buy the mark only if you would act on it.
You inherited a marked detector and the suspect copy was re-distilled — what is that ownership evidence worth?
basics
~20 sLess than the team hopes, and say so plainly: a decayed mark read by a test with an unmeasured error rate is a lead, not evidence. The call is whether to fund the control run or drop the claim.
Transparency commitments force per-field reasons on every model decision. What do you actually negotiate?
basics
~20 sNot whether to explain, but the shape of the disclosure: how many fields it names, at what numeric precision, to which recipient, at what rate, and at what price. The obligation usually runs to the decision subject, one record at a time, not to a machine integrator pulling thousands a day.
How do you price the per-category scores in a moderation API's response at a design review?
basics
~20 sReport what the field buys an adversary in money and time: the annotation budget a competitor no longer funds, and how fast a stand-in reaches usable quality. Leave the keep-or-cut call to the product owner.
showing 31–40 of 40