Why does normal paid use of a content-moderation API also hand the caller a labelled dataset?
answer
- a dataset has two halves
- the caller already owns the inputs
- which half did you just sell
- logged responses are annotations
- they leave with data, not weights
basics
~20 sBecause the caller supplies the text and the endpoint supplies the annotation. Every paid call returns a scored row they can keep, so ordinary integration quietly accumulates the labelled corpus your annotators were paid to produce.
solid answer
~50 sSplit a supervised dataset into its two halves. The inputs are the cheap half: anyone can crawl or generate unlimited unlabelled text. The labels are the expensive half — for moderation they encode a written policy, an annotator pool trained on it, and adjudication of the borderline cases. A prediction API sells exactly that expensive half, one row at a time, at list price. A paying integrator who logs their own requests beside your responses ends a few months of ordinary traffic holding an annotated corpus, and nothing in the transaction was abusive: they used the documented contract as documented. That is the starting position for query-based extraction. Be precise about what they hold, though — annotated data, not your weights, not your architecture, and no behaviour outside the inputs they happened to send. A model still has to be fit. But the half they could not otherwise afford has been bought.
go deeper
Be ready to say plainly that a prediction is a label, and that a caller who keeps their requests beside your replies is accumulating training data. Do not reach for prompts or jailbreaks here.
Expect to explain which half of a supervised dataset the API is selling, and why the input half is nearly free to a competitor while the label half is what a company actually pays for.
Show that you can state the disclosure precisely: annotated rows, bounded by the inputs the caller sent, with parameters and architecture untouched. Overstating it as theft of the model gets the finding dismissed.
Own the framing that the response schema is a commercial decision, not only a technical one: what you publish in the body sets what a competitor no longer has to pay an annotation vendor for.
## The two halves of a supervised dataset Every supervised training set is a pile of pairs: an input, and a label attached to it. The two halves have wildly different prices. The **inputs** are close to free. For a text task, unlabelled text exists in enormous quantity — forum posts, marketplace listings, product reviews, scraped pages, synthetic variations of any of these. A competitor who wants a million pieces of text that look like the traffic a real forum operator sees can obtain them without asking anybody's permission and without a budget line. The **labels** are where the money is, and on a policy task like content moderation the gap is extreme. A label for "is this harassment" is not a fact sitting in the world waiting to be recorded. It is the output of a written policy document, an annotator pool trained on that document, a quality process, and an adjudication path for the cases two annotators disagree on. That apparatus is expensive to build, expensive to run, and it *is* the differentiated asset — two vendors with the same architecture and the same crawl differ almost entirely in how their annotation guidelines resolve the hard middle. ## What a prediction endpoint actually sells A paid classification API answers a request by attaching a label to an input. Stated that way, the security reading is immediate: the endpoint is an annotation service that bills per row. The customer brings the input half, which they had already; the seller supplies the label half, which is the half nobody can cheaply make. A caller does not have to do anything unusual to accumulate the result. Ordinary integration already logs requests and responses — for debugging, for caching, for their own analytics, for their own audit trail. After a few months of ordinary production traffic, that log **is** a supervised dataset: their text, your annotations. No scraping, no undocumented field, no error-message spelunking, no terms violation. The published contract, used as published. This is the entry point to what the literature calls query-based model extraction, and the important structural point is that the first stage of the attack is indistinguishable from being a customer. ## What they hold, precisely Overclaiming here is the classic way to lose an argument in a design review, so state the disclosure exactly: - They hold **annotated rows** — their inputs with your model's scores attached. - The corpus covers **only the region of input space they queried**. Behaviour on text they never sent was never observed. - They do **not** hold your parameters. Nothing in a returned score pins a weight value. - They do **not** hold your architecture, your training corpus, or your policy document — only the behaviour those things produce, on their inputs. - They do **not** yet hold a model at all. Fitting one is still ahead of them, and it costs compute. What has genuinely moved is the annotation, and with it the operational form of the policy judgment encoded in the annotation. That is usually the asset worth defending, which is why the disclosure matters even though it is not "the model was stolen". ## Intent does not change the disclosure A caching integrator, a customer building an offline evaluation set, and a competitor building a rival product all end up with the same artefact. The response contract decides what leaves the building; intent decides only whether anybody assembles it into something. That is why this is an API design conversation rather than an abuse-investigation one — you cannot make the disclosure conditional on motive, because motive is not a field in the request. ## Why interviewers ask this Because the wrong answer is comfortable and common: *the model is ours, the weights are private, we are just returning predictions.* The candidate who sees that a prediction is a **label** — supervision, not merely output — has the frame that every later question in this area depends on: what the response body contains is a security decision, and it is made by whoever designs the schema, usually with no security person in the room.
- At the end of such a campaign, has the model been stolen?No, and saying so costs you the room. They hold a labelled corpus and still have to train something. What comes out is a functional stand-in that agrees with you where they queried; your parameters, architecture and training corpus are not disclosed by replies. What moved is the annotation, and the policy judgment encoded in it.
- Does it matter whether the caller intended to build a copy?Not for the disclosure. The response contract decides what leaves the building; intent decides only whether anyone assembles it. A caching integrator and a competitor end up holding the same artefact, which is why this is an API design question rather than an abuse-detection one.
- What still costs the attacker money once the labels are free?The query bill, training compute, and above all input coverage. They only learn your behaviour on text they thought to send, so a copy that holds up on the traffic a real operator sees requires sourcing text that resembles that traffic. Annotation stops being their constraint; representativeness becomes it.
You are running a paid annotation service that also happens to answer questions. The customer brings the text; you write the labels.
saying these in an interview costs you the question
- Says the model is safe because the weights are private
- Thinks only a terms-of-service violation counts as extraction
- Treats a returned score as output rather than as supervision
- Claims the caller walks away with the trained model itself
- Assumes unlabelled text is the scarce half of the dataset