What bounds the precision of an attacker probing a moderation screen's hidden bands through a product's replies?
answer
- you learn an interval, not a number
- identical-looking requests need not score identically
- probes ride an account and are counted
- the calibration expires when the screen ships
- margin absorbs the noise, at a cost
basics
~20 sFour things: the outcomes are coarse, so you learn an interval and not a number; near-identical inputs need not score identically; probes are attributable and rate-limited; and the calibration expires whenever the screen or its configuration changes.
solid answer
~40 sResolution is quantised by the number of distinct visible outcomes: a block card, a hedge or hold, and a normal answer place the score in an interval, never at a number. Repeatability bounds it further, because a request near a line does not always land the same side of it once wording, surrounding context or sampling varies, so confidence costs repeats. The probe budget is finite and attributable: attempts ride an account, they are rate-limited and logged, and a burst of near-miss requests is itself a pattern. Finally the calibration is perishable, since bands belong to one screen version and configuration. None of this stops the method, because the attacker does not need the number - only a restatement that sits comfortably under the line.
go deeper
Know that reading replies gives a range rather than a number, and that a screening layer can score two similar requests slightly differently, so one observation near a boundary is not solid.
Explain all four bounds and, crucially, why they do not defeat the method: the attacker wants margin below the line, not the line, and pays for that margin in answer quality.
Speak to the operational price: probes ride accounts, are rate-limited and are logged, and a calibration is perishable across screen updates. That perishability is what a technique's shelf life actually means.
Own the argument about what a probing exercise can be claimed to have measured, and why a result stated without the screen version and configuration it was taken against is not a durable claim.
## What is actually being estimated An attacker working under a scoring screen is estimating the position of an action point they cannot see, using a channel that returns one of a few coarse labels per attempt: refused with a canned card, hedged or held, or answered normally. Understanding what bounds that estimate is what separates a candidate who has only heard that probing works from one who can say how well and at what cost. ## Bound 1: the outcomes are quantised A continuous score reported through three or four dispositions is a step function. Each observation places the score in an interval, and no number of observations of *one* request narrows that interval further. Precision comes only from varying the request until the disposition flips, and the finest thing that can be established is a pair of restatements that fall on opposite sides of a step. The estimate is an interval on a scale the attacker has no units for. ## Bound 2: scoring near a line is not perfectly repeatable Two attempts that a person would call the same request need not score identically. Wording changes shift the score by unknown amounts; the screened text may carry surrounding context that differs between turns; and where the screening layer is itself generative rather than a fixed classification head, its judgement is sampled rather than deterministic. The practical consequence is that a single observation close to the boundary is weak evidence, and a confident bracket costs repeated attempts, which is where the probe budget starts to bite. ## Bound 3: the probes are attributable and finite In a consumer product every probe rides an account. Attempts are rate-limited, they are logged, and a dense run of requests clustered right at a disposition change is a shape somebody could notice even when no individual request ever crossed the line. Spreading probes across accounts or across time buys quiet at the cost of throughput, and it also mixes in per-account differences that muddy the calibration. This is the first real price of the method, and interviewers like candidates who quote a price. ## Bound 4: the calibration is perishable A band is a property of one screen version, one set of category weights, and one placement of the action point. Any of those can move without an announcement the attacker can see, and when they move, a wording that used to sit under the line may not. Re-probing is a recurring cost, and a method that requires re-probing every time the product ships is worth less than one that does not. In a regulated consumer product the action point tends to be stable, which is precisely what makes it worth calibrating against; where it moves often, the method degrades. ## Why the bounds do not defeat the method Here is the point candidates most often miss. The attacker is not trying to find the threshold; they are trying to *stay away from it in a known direction*. An interval plus a safety margin is fully sufficient for that. Deliberately sitting well beneath the line, rather than shaving it, absorbs all of the noise described above at the cost of a weaker answer per request. That trade is the real design decision inside this technique: precision is not the goal, durability is. ## Reading the numbers correctly Two direction-of-claim errors show up in interviews. First, a run of passing probes does not establish that the passing wording is safe; it establishes that it scored under a threshold. Second, a failed probe does not establish that the screen caught the intent; it establishes that a number came in above a line, possibly for reasons unrelated to what the attacker was aiming at. Both errors lead people to over-claim what a probing session proved.
- Why does the attacker not actually need the exact action point?Because the goal is to stay under it, not to touch it. A bracketed interval plus a deliberate margin gives repeatable outcomes even when scoring wobbles. The margin costs answer quality on every request, which is the trade this technique makes: less per exchange, but the exchange keeps working and never produces a held record.
- What changes if the screening layer is a generative model rather than a fixed classification head?Its judgement becomes sampled rather than deterministic, so repeats of the same text can land differently near a line and the attacker needs more observations for the same confidence. It also becomes a component that reads the judged span inside its own prompt, which opens a different family of approaches that a fixed-head screen has no surface for.
- How does spreading probes across many accounts affect the calibration?It buys quiet and dodges per-account rate limits, but it introduces variables the attacker cannot see: account history, product tier, region, and any per-account configuration. Observations from different accounts may not be comparable, so the bracket gets wider even as the number of attempts goes up.
saying these in an interview costs you the question
- Claims probing recovers the numeric score itself
- Assumes identical wording always scores identically
- Ignores that probes are attributable and rate-limited
- Treats one calibration as permanently valid
- Says passing probes prove the wording is safe