skip to content

In a multi-label moderation API, how does the response body set an extraction campaign's cost per labelled row?

level: middleimportance: should knowfreq 52%

answer

  1. count the numbers in one reply
  2. an annotator bills per judgment
  3. one payment, every policy
  4. your replies never disagree with themselves
  5. spend to annotated rows is a ratio

basics

~20 s

One call returns one score per policy category, so a single payment annotates that text against every category at once. The response schema, not the call count, fixes how much supervision a dollar of API spend buys an attacker.

solid answer

~50 s

Do the accounting per call rather than per query. A multi-label moderation endpoint answers one request with a whole row of per-category scores — harassment, self-harm, violence, spam — so one payment annotates that text against every policy at once, where an annotation vendor would bill separately for each judgment. The values are also real numbers rather than a single accepted-or-rejected verdict, and the consequence is that a copy reaches a given fidelity in substantially fewer calls than hard verdicts would need. A third effect gets missed: the endpoint is a deterministic function, so the corpus carries none of the inter-annotator disagreement a human-labelled corpus of the same texts would, and a clean target is easier to fit. Quote the ratio honestly, though — correlated categories carry overlapping information, so scores-per-call is an upper bound on judgments-per-call, not a measured one.

go deeper

for a junior

Know that a reply carrying several category scores annotates a text against several policies at once, so calls and labelled rows are not the same unit of account.

for a middle

Be able to walk the arithmetic out loud — numbers per reply, replies per dollar, annotated rows per dollar — and say what each step assumes.

for a senior

Demonstrate the honest caveats, correlated categories and query coverage, so the ratio you quote reads as an upper bound rather than a headline somebody can puncture.

for a principal

Own the exchange rate as a figure the business can act on: what a competitor's annotation programme would have cost against what your published response body charges them instead.

## Count the numbers, not the requests The instinct is to measure an extraction campaign in queries: how many calls did they make, how much did that cost at list price. That is the wrong unit, because the thing being bought is **supervision**, and a single call can carry a great deal of it or very little depending on what the response body was designed to contain. For a multi-label moderation classifier — text in, one score per policy category out — the arithmetic runs like this: 1. **Numbers per reply.** One request returns a score for every category the product supports. That is one text annotated against every policy in a single transaction. 2. **Replies per dollar.** Published per-call pricing converts spend into replies directly. 3. **Annotated rows per dollar.** Multiply through. The result is a labelling rate, denominated in money, that the seller set when they designed the schema. The comparison that makes the number land is the alternative. A human annotation programme is billed per **judgment**, not per document, and a policy judgment is not a cheap one: the annotator has to be trained on a guideline document, disagreements on the borderline cases have to be adjudicated, and quality has to be sampled. Four policy judgments on one text is four billable decisions there, and one paid call here. ## Why the values matter as well as the count Beyond the number of fields, the *kind* of value matters to the campaign's rate. A single accepted-or-rejected verdict is one bit of feedback about that input. A real-valued score per category carries more than that, and the well-established consequence — the reason this design choice is a security decision and not a formatting one — is that a usable copy is reached in materially fewer calls when the replies are graded rather than binary. You do not need the theory of why that is true to price the endpoint; you need the consequence, which is that the call budget for a given fidelity falls sharply. ## The effect nobody costs: the target is noise-free A human-labelled moderation corpus of the same texts would carry genuine disagreement. Two trained annotators looking at the same borderline post produce different labels a nontrivial fraction of the time, and a model fitting that corpus spends capacity on the disagreement. An endpoint does not disagree with itself. Ask twice, get the same score. So the corpus an attacker assembles from replies is *self-consistent by construction* — a clean function to fit rather than a noisy sample of human judgment. That makes it, in a specific and slightly uncomfortable sense, a **better** training target than the corpus the seller trained on, and it means the copy converges on fewer examples than a from-scratch programme would need. ## The caveats that keep the number honest Two corrections belong in any figure you quote: - **Correlated categories.** Four scores that move together do not carry four independent judgments. Abusive text often trips several policies at once, so the marginal information in the fourth field is well below the first. Quote scores-per-call as an **upper bound** on judgments-per-call. - **Coverage, not supervision, may be the real constraint.** The copy only agrees with the target on inputs resembling the ones queried. If the attacker cannot source text that looks like the traffic a real operator sees, richer replies just annotate the wrong inputs more precisely. Past that point, extra fields in the body buy them nothing. ## Where this lands The exchange rate between money and annotated rows is a property of the response schema, and it was fixed by whoever wrote the API contract — typically for good product reasons, since integrators genuinely need per-category values to route to their own review queues. The security observation is not that the fields are wrong. It is that the schema has a price attached that nobody in the design review computed, and computing it is a five-minute exercise anybody can do from the public documentation. An interviewer asking this is checking two things: whether you measure in the right unit, and whether you can state the caveats without being asked. A candidate who confidently multiplies fields by calls and stops has produced a headline; a candidate who then says "and that is an upper bound, here is why" has produced a finding.

  • Where does a richer reply stop helping the attacker at all?
    Once their bottleneck is input coverage rather than supervision. A copy only agrees with the target on text resembling what was sent, so if they cannot source queries that look like real operator traffic, extra fields refine labels on the wrong inputs. Past that point the schema is no longer the binding constraint.
  • Why does the self-consistency of the labels matter to the attacker?
    A human-labelled corpus of the same texts carries real disagreement on borderline cases, and a model fitting it burns capacity on that noise. An endpoint returns the same score for the same input every time, so the target is a clean function and the copy converges on fewer examples.
  • Does the accounting change if the policy categories are correlated?
    Yes, downward. Categories that fire together carry overlapping information, so four returned scores are worth less than four independent judgments. Quote scores-per-call as an upper bound on supervision-per-call and say so, otherwise the first person who checks will discount the whole estimate.

saying these in an interview costs you the question

  • Counts calls but never counts the numbers per call
  • Assumes one query yields one binary label regardless of schema
  • Treats query volume as the only cost driver
  • Confuses the copy's fidelity ceiling with the labelling rate
  • Ignores that correlated categories overstate the per-call gain

context