skip to content

Your moderation API returns only a verdict, no scores. Which attacks does that stop?

level: middleimportance: must knowfreq 62%

answer

  1. ask what the removed field was supplying
  2. some families never needed a number
  3. the change is a multiplier, not a wall
  4. measure the new cost, do not assume it

basics

~20 s

None outright. Removing confidences deletes a cheap signal and forces the adversary into label-only families such as decision-based search and verdict-based membership tests, at far higher query cost. Coarsening a reply moves a family's price; it closes none.

solid answer

~40 s

It stops nothing and prices several things out. Confidences are a continuous signal: an attacker sees whether an edit helped before the verdict flips, and reads a per-example magnitude that reflects how well the model fits that example. Take those away and both become discrete, so the attacker switches to families that consume only the returned symbol. Decision-based search still works, and membership tests still run on verdicts alone, both at much higher query counts. That is worth having, but the defensible claim is 'we multiplied the query bill for these families', not 'we removed the leak'. It also costs us: integrators who set their own thresholds on our confidences lose that, and some will approximate a score by sending several variants per post, raising our traffic rather than lowering our risk.

go deeper

for a junior

Remember the shape of the correct answer: removing confidences raises cost, it does not remove attacks. Be able to say that some attacks need only the returned class.

for a middle

Explain what the removed field supplied, in three parts: feedback on edits before a verdict flips, a per-example confidence magnitude, and the standing of categories that did not fire.

for a senior

Show that you would measure the multiplier rather than assert it, run the same objective against both contracts, and write the finding as a repricing with a number attached.

for a principal

Weigh the customer-side cost against the attacker-side cost and be ready to say that integrators who lose the field may reconstruct it with extra calls, raising volume without lowering exposure.

## The claim under test 'We return only the class, so there is nothing to leak' is the standard answer and it is wrong in a specific, correctable way. It confuses the **existence** of an attack family with its **price**. Almost nothing about a hosted model's exposure is binary; nearly everything is a query count. ## What confidences were doing For a hosted content-moderation service that takes a post and returns a policy decision, the per-category confidence carried three things at once: **Gradient-free feedback.** A continuous value changes under small edits to the input even when the verdict does not. That converts blind trial and error into a search with feedback after every call. **A per-example magnitude.** How strongly the model commits on one input reflects how well it fits that input, which is precisely the raw material of a membership signal. **Cross-boundary standing.** Confidences for categories that did not fire locate the input relative to several decision boundaries in one call rather than one. Deleting the field removes all three at once. That is genuinely a big cut, and it is why operators reach for it. ## What survives, and why Two well-established families need nothing but the returned symbol. **Decision-based search.** An adversary who can observe only where the verdict changes can still locate a boundary, because the verdict flipping *is* an observation. It takes many more calls than reading a score, and that is the whole difference. **Label-only privacy signals.** Membership inference is a consequence of imperfect generalisation: a model fits what it saw more tightly than what it did not. That difference shows up in confidences most cheaply, but it also shows up in how robustly a verdict holds, so a verdict-only endpoint still carries a membership signal at reduced strength. Membership results are measured as advantage over the base rate, so a weaker signal is a smaller advantage, not an absent one. And the reply body is not the only channel. Latency, error and validation text, whether the reply names the policy that fired, and whether ties resolve deterministically all remain observable after you delete a numeric field. ## Cost is a legitimate control, stated honestly None of this makes coarsening pointless. Multiplying an attacker's query count by a large factor can push an engagement past what an adversary will fund, and against an opportunistic adversary that is a win. The discipline is in how you write it down: | What you might claim | What the evidence supports | | --- | --- | | 'The endpoint no longer leaks' | 'The endpoint no longer leaks a continuous per-call signal' | | 'Extraction is prevented' | 'The families that consume confidences moved to label-only variants at a higher query count' | | 'Membership inference is closed' | 'The membership advantage should fall; measure it rather than assume it' | The right-hand column is what survives a review. The left-hand column is a claim that a red-teamer will falsify on the first engagement. ## The cost you pay for it A response contract is a customer-facing artefact. Third-party integrators who chose their own thresholds on your confidences lose the ability to trade precision for recall on their own terms and will either drop the integration or reconstruct an approximate confidence by sending several variants of each post. The second outcome increases your call volume and gives the same callers more observations per decision than they had before, which is close to the opposite of the intent. ## The interview shape When this is asked, the interviewer is checking whether you reason in prices or in walls. A strong answer names what the removed field supplied, names the families that consume only the symbol, says explicitly that the change is a multiplier on cost, and offers to measure the multiplier rather than assert it. A weak answer says the endpoint is now safe.

  • How would you actually demonstrate to a reviewer that dropping confidences bought something?
    By measuring the same goal twice: pick one attack objective, run it against the endpoint with confidences and against the coarsened one, and report the query counts side by side under identical rate conditions. That produces a multiplier you can defend. Anything else is an assertion, and an unmeasured 'this is now protected' is exactly the claim a red-team engagement is designed to break.
  • Does a verdict-only contract help against an adversary trying to learn whether one specific record was in the training set?
    It reduces the signal but does not remove it. Membership leakage comes from the model fitting what it saw more tightly than what it did not, and that difference still shows in how firmly a verdict holds under small changes. Expect a lower advantage over the base rate, and treat the harm as set by what membership means about the person rather than by the accuracy figure.
  • What does coarsening the reply body not touch at all?
    Everything outside the body. Response latency, error and validation messages, whether the reply names the category that fired, and deterministic tie-breaking remain observable, and so does behaviour at the service's operating limits. If your write-up says the observation channel is closed, it has to account for these rather than assume the field deletion covered them.

Taking the labels off a building's doors slows a stranger down considerably. It does not remove the doors.

saying these in an interview costs you the question

  • Says removing scores means the endpoint leaks nothing
  • Claims membership inference needs confidence values to work
  • Treats attack exposure as binary rather than as a query price
  • Reports the coarsening as a control without measuring the new cost
  • Ignores that integrators lose a field they were thresholding on

context