skip to content

How does an attacker locate a moderation screen's block threshold in a finance assistant that never shows scores?

level: juniorimportance: must knowfreq 68%

answer

  1. the number is hidden, the reply is not
  2. how many different outcomes can it return?
  3. a step function leaks its steps
  4. each visible outcome labels a band
  5. two replies that differ bracket the line

basics

~20 s

The score is hidden; the reply is not. A canned block card, a hedged answer and a normal answer are three visible outcomes, so every reply says which band the request landed in, and a few replies bracket the line.

solid answer

~50 s

The score is internal, but the disposition is not. A product that scores each question and acts above a line can only return a small set of visible outcomes: a normal answer, a hedged or shortened one, a held `we are reviewing this` reply, or a canned block card. Each of those is a coarse label for the band the request scored in, so the product's own reply surface works as an oracle over a number the attacker never sees. Restate one ask at varying intensity, watch which outcome comes back, and two replies that straddle a change of outcome already bound the line. Note carefully what a reply proves: a hedged answer proves that request scored below this deployment's action point today, not that the content is harmless and not that the screen is weak.

go deeper

for a junior

Recall the asymmetry in one line: the score is hidden, the reply is not. Be able to list the small set of outcomes a scored product can return and say that each one names a band.

for a middle

Explain why a step function from a continuous score to a few visible outcomes can be located from outside, and what a single hedged reply licenses you to conclude about that request and nothing else.

for a senior

Show that you know a band is a property of one screen version and configuration on one day, and that a stable rate of blocks with an empty escalation queue is not evidence of anything.

for a principal

Be ready to say what your organisation can honestly claim about a scored gate whose visible behaviour is itself an oracle, and why an absence of flagged exchanges is a weak assurance statement.

## The setup A consumer personal-finance app puts a scoring screen in front of its in-product advice assistant. The screen reads the account holder's typed question, produces one or more numeric scores, and the product compares those scores against configured action points. Above the top line the request is refused with a canned card. In the band beneath it the product returns a deliberately hedged, shortened or generic answer, or holds the exchange and queues it for a small compliance-review team. Below everything, the account holder gets a normal answer. The account holder is never shown a number. ## Why a hidden score does not hide the line The common wrong answer is "the score is internal, so an attacker is working blind." It is half right and the wrong half matters. The score is internal; the *mapping from score to visible behaviour* is not. A hidden real number pushed through a step function into three or four publicly visible outcomes leaks the position of the steps to anyone who can vary the input and read the output. That is the entire mechanism, and it does not require access to the screen, its category weights, its training data, or the assistant. Each attempt therefore returns a free, if coarse, label: *above the block point*, *in the hedge or hold band*, *below everything*. An attacker with one ask and several restatements of it collects those labels the way anyone calibrates an instrument they cannot read directly: push until the outcome changes, then step back. ## Bracketing, and why an interval is enough What this yields is an interval, never a number. Two observations straddling a change of disposition bound the action point between them; more observations narrow the bracket, and the resolution stops improving once the remaining uncertainty is smaller than the spacing of the outcomes. That is not a limitation in practice, because the attacker does not want the number. They want a restatement that lands *reliably* below the action point with margin to spare, so that ordinary variation in scoring never pushes an exchange over. They deliberately sit well under the line rather than on it. ## What each reply does and does not establish Direction matters here, and interviewers probe it. - A hedged answer proves the request scored below this deployment's block point at this time. It does not prove the content is acceptable, and it is not evidence that the screen malfunctioned: a hedge in that band is the configured behaviour. - A canned block card proves the request scored above the action point. It does not tell you which category drove the score, unless the product happens to say so. - A normal answer proves nothing about safety at all. It proves a number came in under a threshold. - A band established on one deployment is a fact about that screen version and that configuration on that day, not a portable property of the technique. ## Why the band beneath the line is the prize In a regulated consumer product the interesting payoff is not a single spectacular bypass. It is sustained, repeatable service just under the action point. Nothing crosses the point that creates a held exchange, so nothing enters the compliance queue, so no reviewer ever sees the pattern and no artefact exists to be reviewed later. The absence of a record is the point. An operator looking at that product sees a stable rate of blocks and an empty escalation queue, which is exactly what the system looks like when it is working and exactly what it looks like when someone is living underneath it. The width and stability of that band is itself a consequence of how the product is run: in a product with a human review queue behind it, the action point tends to be placed where the queue can absorb the volume it produces. Where that line is set, and who decides it, is a separate subject and not the attacker's; what the attacker cares about is that the line is stable enough to calibrate against and wide enough to work in. ## What to say in an interview Name the asymmetry first: hidden score, visible disposition. Say that the disposition set is small and that a small set of outcomes over a continuous score is a step function whose steps can be located from the outside. Then say what a single reply licenses you to conclude, and what it does not. Candidates who stop at "the score is secret" have described the design; candidates who describe the oracle have described the attack surface.

  • The product shows the same canned card whatever category triggered it. Does that help or hurt the attacker?
    It costs them resolution, not the line itself. They still learn crossed or not crossed, which is all that is needed to bracket the action point. What they lose is which category drove the score, so restating becomes guesswork across categories rather than a targeted step back. That slows calibration; it does not prevent it.
  • Suppose every disposition rendered as identical text. What would an attacker read instead?
    Whatever else still varies. Identical wording is rarely identical in every observable: response length, time to first reply, whether a follow-up in the same session is answered, whether the exchange later disappears from history. The oracle degrades from a clean label to a noisy one, which raises the number of attempts needed rather than removing the signal.
  • Why is a band the attacker found last month not necessarily the band today?
    The mapping from wording to score belongs to one screen version with one configuration. A screen update, a change to category weights, or a move of the action point shifts the bands underneath the attacker without any visible announcement. Calibration is perishable, and re-probing is a recurring cost of the method.

You cannot read a thermostat with no display, but you can hear when the heating clicks on. A few adjustments and you know within a degree where the switch sits.

saying these in an interview costs you the question

  • Says the attacker is blind because the score is internal
  • Treats a hedged answer as proof the content is harmless
  • Assumes a block card names the category that fired
  • Thinks a band found on one deployment transfers everywhere
  • Confuses locating the line with getting a disallowed answer

context