skip to content

Your moderation API now rounds scores to two decimals and returns only the fired policy. What changed?

level: seniorimportance: should knowfreq 42%

answer

  1. two changes, price them separately
  2. what does rounding keep, what does it drop
  3. the omitted categories are the bigger cut
  4. check undocumented tie-breaking

basics

~20 s

Two separate changes. Rounding coarsens the signal and creates ties, so probes must be larger or more numerous. Dropping the non-fired categories is the bigger cut: the reply now describes one boundary instead of several.

solid answer

~40 s

Price the two halves separately. Rounding to two decimals keeps ordering and direction but discards everything below the second decimal, so any read of small differences needs larger or more numerous probes; the multiplier depends on how much signal sat under that precision. Suppressing the categories that did not fire is the bigger loss, because the reply now describes one boundary instead of several, and for every other category the contract has effectively become label-only. Scoping this, I would re-plan around label-only families, budget a much larger query count, and check two things the schema omits: whether ties break deterministically, which quietly restores ordering, and whether latency or error text still varies with the decision. Then report a measured price increase, not a closed channel.

code

json · 10 lines
json
// before
{ "verdict": "block",
  "scores": { "harassment": 0.8137, "self_harm": 0.0412,
              "violence": 0.2260, "spam": 0.0031 } }

// after
{ "verdict": "block",
  "fired": "harassment",
  "score": 0.81 }
// non-fired categories: omitted. ties: undocumented.

go deeper

for a junior

Know that rounding a returned confidence and omitting other categories are two different changes, and that neither turns the endpoint into a black box that reveals nothing.

for a middle

Explain that rounding keeps ordering and direction but sets a floor on the smallest difference a caller can see, so probes get larger or more numerous.

for a senior

Demonstrate that you price the two halves separately, check undocumented tie-breaking and rounding placement, and report a measured query multiplier instead of a claim that the channel is closed.

for a principal

Own the trade: the suppressed categories are the real cut, the rounding is the part customers will push back on first, and the write-up must say what each half bought before anyone calls it a control.

## Two changes, not one A coarsening step usually arrives as a single ticket, and the review has to split it. Rounding a returned value and suppressing entries are different operations with different effects on different adversaries, and lumping them together is how teams end up unable to say what they bought. ## Rounding: precision, ties, and probe size Rounding a confidence to two decimals keeps three things: which category fired, the value's rough level, and, for anything still returned, the ordering. It discards the fine structure. What that costs an adversary depends entirely on how much of the signal they needed lived below the retained precision. The practical consequences are: - **Ties appear.** Inputs that were distinguishable at four decimals now return the same value, so a caller cannot tell whether a small edit helped. - **Probes must be larger or more numerous.** To see a difference through the rounding, the caller must either make bigger changes to the input, which risks changing its meaning, or aggregate over more calls. - **The direction survives.** Once a difference is large enough to cross a rounding step, it still points the same way. Rounding raises a floor; it does not scramble the signal. So rounding is a cost multiplier, and the multiplier is measurable rather than assumable. It is also the knob operators actually turn, because unlike deleting the field it leaves most legitimate consumers functional. ## Suppression: losing several boundaries at once Returning only the category that fired is the more consequential half. Before, one call reported where the input sat relative to every policy boundary. After, one call reports one boundary and says nothing about the rest. For every suppressed category the contract has silently become label-only, and 'label-only' here means even weaker than usual: the caller does not learn the runner-up, so they cannot even tell which boundary they are closest to crossing next. That is the change most likely to actually move an engagement's price, and it is the one teams under-credit because it looks like tidying up a response body rather than removing information. ## What the schema does not say, and you must check A red-teamer reading a response contract should treat the undocumented behaviour as part of the surface: - **Tie-breaking.** If ties among equal rounded values are broken deterministically by a fixed rule, the ordering can be partially reconstructed, and some of what rounding removed comes back. - **Rounding placement.** Rounding applied after the decision is a disclosure change only; a service that also *decides* on rounded values has changed behaviour, not just output. - **Everything outside the body.** Latency, error and validation text, and behaviour at operating limits are unaffected by either change. - **Boundary reporting.** Whether the service reports the fired category by name at all, since a named category is more information than a bare block verdict. ## How to report it The finding writes itself if you keep the two changes separate: | Change | What it costs an adversary | What it does not do | | --- | --- | --- | | Rounding to two decimals | Larger or more numerous probes to see a difference | Remove direction or ordering | | Suppressing non-fired categories | All standing against boundaries not crossed | Prevent label-only families on the fired one | Then attach a measured number: run one objective against both contracts and report the query counts. A senior answer here is not a list of consequences, it is the discipline of pricing each half separately and refusing to claim the sum removed a family. ## The trap The trap in this question is to treat rounding as the important part because it is the visibly 'security-flavoured' change. In practice the suppressed categories are the larger cut, and the rounding is the part most likely to be reversed later by a customer complaint about lost precision. Saying that out loud is what distinguishes somebody who has actually reviewed a response contract from somebody reasoning about it in the abstract.

  • Why does deterministic tie-breaking matter once values are rounded?
    Because rounding creates ties, and a fixed rule for resolving them is itself information. If equal rounded values are always ordered the same way by an underlying quantity, a caller can recover part of the ordering the rounding was meant to hide. Undocumented behaviour like this belongs in the review, since the schema alone will not tell you whether the coarsening achieved what the ticket claimed.
  • Does it matter whether the rounding happens before or after the decision is taken?
    Yes, and they are different findings. Rounding applied to an already-taken decision is purely a disclosure change. A service that decides on rounded values has changed its behaviour, which can shift borderline cases and shows up as accuracy movement rather than as an attacker cost. Establish which one you are looking at before writing either finding.
  • How would you size the engagement after reading this contract?
    Budget for label-only families against the suppressed categories, and for a probe cost on the fired one that depends on how much signal sat below two decimals. Then verify the assumption cheaply before committing: a small pilot measuring how many calls are needed to see a difference through the rounding turns a guess into a number you can put in the plan.

saying these in an interview costs you the question

  • Treats rounding and suppression as one change with one effect
  • Assumes rounding destroys the direction of the signal
  • Ignores undocumented tie-breaking as part of the surface
  • Calls the coarsened endpoint closed rather than more expensive
  • Never asks whether the decision itself now uses rounded values

context