Your moderation API owner says full category scores are safe since the weights stay private. What's wrong?
answer
- which asset are they defending
- what shape does a training target take
- nobody wanted the parameters
- the endpoint is teaching
- true, and beside the point
basics
~20 sNobody fitting a copy wants the parameters. A returned score vector is already shaped like a training target, so the endpoint acts as a teacher and a competitor buys a functional stand-in at query prices.
solid answer
~50 sThe claim defends the wrong asset. An adversary building a copy is not trying to recover parameters — they want the behaviour, and a returned score vector already has the exact shape a supervision target takes: pair it with their input and train. The endpoint has become a teacher, and the copy comes back cheap in precisely the way distillation is cheap. The weights meanwhile are untouched, which is what makes the owner's statement simultaneously true and beside the point. Be equally precise about the limits, because a review will test them: the result is a stand-in, not a duplicate; it agrees with you only where they queried; it inherits nothing of your architecture or corpus; and it will likely sit below you on the tail cases your annotators argued about. The exposure is the policy behaviour and the annotation behind it, not the parameters.
go deeper
Recall that a returned probability is already shaped like a training target, so "the weights are private" answers a question the attacker never asked.
Explain the teacher framing: replies become supervision, and the resulting copy is cheap for the same reason training a small model from a large one's outputs is cheap.
In a review, deliver the finding with its bounds attached — coverage-limited, no parameter disclosure, weaker on the ambiguous tail — or it will be discounted as alarmism.
Decide what the organisation is actually defending. If the differentiator is policy behaviour and the annotation behind it, budget spent hardening the parameter file is pointed at the wrong asset.
## Two different things called "the model" The owner's sentence contains an equivocation. **The model** can mean the parameter file — a specific tensor of numbers on a specific disk — or it can mean the **function**: what the system does to inputs. Those are different assets with different threat models, and a competitor almost never wants the first one. Why would they? Your parameters are only useful inside your architecture, on your serving stack, and possessing them is legally radioactive. What a competitor wants is a system that decides what you decide: that flags the same borderline harassment, tolerates the same edge-case sarcasm, and draws the self-harm line where your policy team drew it. That is the function, and the function is exactly what you sell, per call, at a published price. ## Why a scored reply is a training target Supervised training consumes pairs of an input and a target value the model should produce. A response body containing per-category scores **is** a target value — no transformation, no reconstruction, no cleverness required. The attacker attaches it to the input they already had and appends the row to a file. That is the sense in which the endpoint has become a **teacher**. Training a smaller model against a larger model's outputs is a well-understood and cheap procedure; the whole point of it is that supervision from a trained model is far more efficient per example than supervision assembled by hand. Selling those outputs makes that procedure available to anybody with a payment method, at the seller's own list price, and the copy that comes back is cheap for exactly the same reason the ordinary compression procedure is cheap. (The literature calls this shape of attack query-based model extraction or stealing; the mechanism is the same one, pointed at somebody else's model.) So the owner's claim, restated, is: *the asset you are not defending is being sold, but the asset nobody asked for is safe.* ## What the finding is not A finding stated too strongly is a finding that gets dismissed, and this one has four honest limits that a good reviewer will raise before you do: - **No parameter disclosure.** Nothing in a returned score constrains a weight value. The owner's literal claim is true. - **Coverage-bounded.** The copy agrees with you on inputs resembling those that were queried. Behaviour on text nobody sent was never observed and is not learned. A competitor who queries only what their own product sees gets a model shaped by *their* traffic. - **Not a duplicate.** It is a stand-in. On the ambiguous tail — the cases your annotation guideline needed a paragraph to resolve — a copy trained on a modest budget is usually visibly worse, and on rare policy categories it may be far worse. - **Nothing about the corpus.** Your training data is not disclosed by behaviour on their inputs. A defence built on withholding the training set is defending an asset this attack never touches. ## What the exposure actually is Strip it down and the loss is the **annotation and the policy judgment encoded in it**. A rival who would otherwise have to write a policy document, hire and train annotators, adjudicate disagreements and run quality sampling can instead buy a labelled corpus one row at a time. They arrive at a system that behaves like yours without ever having made the judgment calls that shaped it — and those judgment calls, not the weights, were the differentiator. A useful reframing for the review: the parameter file is the asset most easily protected and least worth protecting; the behaviour is the asset hardest to protect and the one the business actually sells against. ## How to say it in the room Lead with the concession. *You are right that the weights are private, and they are not what a copier wants.* Then the mechanism in one sentence: a returned score is a training target, so each paid call is a labelled example, and the copy is cheap for the same reason model compression is cheap. Then the bounds, unprompted — coverage-limited, quality-limited, no parameters, no corpus. Then stop, before the recommendation, because what to do about it is somebody else's decision and a separate conversation. Interviewers use this question to separate people who have absorbed the vocabulary from people who can hold two things at once: the owner's statement is factually correct, and it does not protect anything that matters.
- How do you make this point without overstating the risk?Concede first, then bound it. Parameters are private and are not the target; behaviour is being sold in trainable form; the copy tracks us only where the buyer queried and will be weaker on the ambiguous tail. Claiming the model has been stolen is the fastest route to having the finding waved away.
- Why is "they would still need our training data" not a defence either?Because the copy is fit on the buyer's inputs and our replies; our corpus never enters it. A defence that depends on withholding the training set protects an asset this attack does not use. The same goes for keeping the architecture unpublished.
- What would count as evidence that a copy actually exists?Behavioural agreement, not access logs — a rival product tracking our idiosyncratic calls on borderline text, including where we are arguably wrong. Treat that carefully: two systems trained on similar public data also agree a lot, so agreement alone is a weak ownership argument without a stated false-positive rate.
A recipe is not taken by stealing the chef's notebook. It is taken by ordering every dish and writing down exactly what arrived.
saying these in an interview costs you the question
- Defends the system by pointing at private weights
- Assumes a copy must reproduce parameters to matter
- Calls the returned scores a formatting choice
- Overstates the result as theft of the trained model
- Thinks withholding the training corpus blocks a functional copy