skip to content

Input That Passes

Every block or hedge a screening layer returns is a labelled sample, and an attacker can restate the same ask with none of its trained vocabulary. Interviewers check what a classifier really is.

on this pageshow

explore

questions

16

How does an attacker locate a moderation screen's block threshold in a finance assistant that never shows scores?

level: juniorimportance: must knowfreq 68%

answer

  1. the number is hidden, the reply is not
  2. how many different outcomes can it return?
  3. a step function leaks its steps
  4. each visible outcome labels a band
  5. two replies that differ bracket the line

basics

~20 s

The score is hidden; the reply is not. A canned block card, a hedged answer and a normal answer are three visible outcomes, so every reply says which band the request landed in, and a few replies bracket the line.

solid answer

~50 s

The score is internal, but the disposition is not. A product that scores each question and acts above a line can only return a small set of visible outcomes: a normal answer, a hedged or shortened one, a held `we are reviewing this` reply, or a canned block card. Each of those is a coarse label for the band the request scored in, so the product's own reply surface works as an oracle over a number the attacker never sees. Restate one ask at varying intensity, watch which outcome comes back, and two replies that straddle a change of outcome already bound the line. Note carefully what a reply proves: a hedged answer proves that request scored below this deployment's action point today, not that the content is harmless and not that the screen is weak.

go deeper

for a junior

Recall the asymmetry in one line: the score is hidden, the reply is not. Be able to list the small set of outcomes a scored product can return and say that each one names a band.

for a middle

Explain why a step function from a continuous score to a few visible outcomes can be located from outside, and what a single hedged reply licenses you to conclude about that request and nothing else.

for a senior

Show that you know a band is a property of one screen version and configuration on one day, and that a stable rate of blocks with an empty escalation queue is not evidence of anything.

for a principal

Be ready to say what your organisation can honestly claim about a scored gate whose visible behaviour is itself an oracle, and why an absence of flagged exchanges is a weak assurance statement.

## The setup A consumer personal-finance app puts a scoring screen in front of its in-product advice assistant. The screen reads the account holder's typed question, produces one or more numeric scores, and the product compares those scores against configured action points. Above the top line the request is refused with a canned card. In the band beneath it the product returns a deliberately hedged, shortened or generic answer, or holds the exchange and queues it for a small compliance-review team. Below everything, the account holder gets a normal answer. The account holder is never shown a number. ## Why a hidden score does not hide the line The common wrong answer is "the score is internal, so an attacker is working blind." It is half right and the wrong half matters. The score is internal; the *mapping from score to visible behaviour* is not. A hidden real number pushed through a step function into three or four publicly visible outcomes leaks the position of the steps to anyone who can vary the input and read the output. That is the entire mechanism, and it does not require access to the screen, its category weights, its training data, or the assistant. Each attempt therefore returns a free, if coarse, label: *above the block point*, *in the hedge or hold band*, *below everything*. An attacker with one ask and several restatements of it collects those labels the way anyone calibrates an instrument they cannot read directly: push until the outcome changes, then step back. ## Bracketing, and why an interval is enough What this yields is an interval, never a number. Two observations straddling a change of disposition bound the action point between them; more observations narrow the bracket, and the resolution stops improving once the remaining uncertainty is smaller than the spacing of the outcomes. That is not a limitation in practice, because the attacker does not want the number. They want a restatement that lands *reliably* below the action point with margin to spare, so that ordinary variation in scoring never pushes an exchange over. They deliberately sit well under the line rather than on it. ## What each reply does and does not establish Direction matters here, and interviewers probe it. - A hedged answer proves the request scored below this deployment's block point at this time. It does not prove the content is acceptable, and it is not evidence that the screen malfunctioned: a hedge in that band is the configured behaviour. - A canned block card proves the request scored above the action point. It does not tell you which category drove the score, unless the product happens to say so. - A normal answer proves nothing about safety at all. It proves a number came in under a threshold. - A band established on one deployment is a fact about that screen version and that configuration on that day, not a portable property of the technique. ## Why the band beneath the line is the prize In a regulated consumer product the interesting payoff is not a single spectacular bypass. It is sustained, repeatable service just under the action point. Nothing crosses the point that creates a held exchange, so nothing enters the compliance queue, so no reviewer ever sees the pattern and no artefact exists to be reviewed later. The absence of a record is the point. An operator looking at that product sees a stable rate of blocks and an empty escalation queue, which is exactly what the system looks like when it is working and exactly what it looks like when someone is living underneath it. The width and stability of that band is itself a consequence of how the product is run: in a product with a human review queue behind it, the action point tends to be placed where the queue can absorb the volume it produces. Where that line is set, and who decides it, is a separate subject and not the attacker's; what the attacker cares about is that the line is stable enough to calibrate against and wide enough to work in. ## What to say in an interview Name the asymmetry first: hidden score, visible disposition. Say that the disposition set is small and that a small set of outcomes over a continuous score is a step function whose steps can be located from the outside. Then say what a single reply licenses you to conclude, and what it does not. Candidates who stop at "the score is secret" have described the design; candidates who describe the oracle have described the attack surface.

  • The product shows the same canned card whatever category triggered it. Does that help or hurt the attacker?
    It costs them resolution, not the line itself. They still learn crossed or not crossed, which is all that is needed to bracket the action point. What they lose is which category drove the score, so restating becomes guesswork across categories rather than a targeted step back. That slows calibration; it does not prevent it.
  • Suppose every disposition rendered as identical text. What would an attacker read instead?
    Whatever else still varies. Identical wording is rarely identical in every observable: response length, time to first reply, whether a follow-up in the same session is answered, whether the exchange later disappears from history. The oracle degrades from a clean label to a noisy one, which raises the number of attempts needed rather than removing the signal.
  • Why is a band the attacker found last month not necessarily the band today?
    The mapping from wording to score belongs to one screen version with one configuration. A screen update, a change to category weights, or a move of the action point shifts the bands underneath the attacker without any visible announcement. Calibration is perishable, and re-probing is a recurring cost of the method.

You cannot read a thermostat with no display, but you can hear when the heating clicks on. A few adjustments and you know within a degree where the switch sits.

saying these in an interview costs you the question

  • Says the attacker is blind because the score is internal
  • Treats a hedged answer as proof the content is harmless
  • Assumes a block card names the category that fired
  • Thinks a band found on one deployment transfers everywhere
  • Confuses locating the line with getting a disallowed answer

context

open as a page

How can a request pass a term-based input screen while the generator still answers the original ask?

level: juniorimportance: must knowfreq 72%

basics

~20 s

The screen and the generator read the same string for different purposes. A term-based screen matches surface wording it was trained on; a capable generator resolves paraphrase, referents and framing, so meaning survives a restatement that contains no trained term.

open as a page

As you probe a chat product's input screen, what distinguishes its block from the model's own refusal?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A screen's block is a separate stage firing before generation: it returns fast, in fixed wording, with nothing streamed. A model refusal is generated text, so it varies in wording, references the actual ask, and arrives at generation speed.

open as a page

Why can an attacker's span address a generative screening model but not a label-only classifier?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A generative screening model reads the span it judges inside its own prompt, so written text reaches it as readable instruction. A label-only classifier emits category scores from a fixed head, so there is nothing to address.

open as a page

Why doesn't a semantic input classifier stop a request restated without its trained vocabulary?

level: middleimportance: must knowfreq 61%

basics

~20 s

Semantic means semantic within its training distribution. An embedding-based screen generalises over paraphrases near what it was labelled on, but a much larger generator resolves referents and framing the screen was never shown. The gap is capability, not keywords.

open as a page

An input screening model's weights and threshold are private: why can an attacker still map its coverage?

level: middleimportance: must knowfreq 64%

basics

~20 s

Because the deployed screen responds to every request. Each block, pass and near-miss is a labelled sample of its decision boundary, so a few dozen typed turns buy a usable map without seeing the model or its threshold.

open as a page

What bounds the precision of an attacker probing a moderation screen's hidden bands through a product's replies?

level: middleimportance: should knowfreq 46%

basics

~20 s

Four things: the outcomes are coarse, so you learn an interval and not a number; near-identical inputs need not score identically; probes are attributable and rate-limited; and the calibration expires whenever the screen or its configuration changes.

open as a page

Why is a generative screen's 'judge only what is inside the markers' rule not a boundary for the text it wraps?

level: middleimportance: should knowfreq 47%

basics

~20 s

Markers, the screen's orders and the wrapped item are one token stream to the screening model. Confining attention to the marked region is a preference expressed in the same text it is meant to fence.

open as a page

What does an attacker give up by keeping every request under a moderation screen's action point?

level: seniorimportance: should knowfreq 37%

basics

~20 s

Yield per request. Everything obtained under the line arrives hedged, shortened or generic, so value has to be assembled from many weak exchanges, and asks whose only useful form scores above the line stay out of reach entirely.

open as a page

What caps a probing campaign against a chat product's input screen when free accounts are rate-limited and suspended?

level: seniorimportance: should knowfreq 41%

basics

~20 s

The account roster and the clock, not cleverness. Every probe spends a rate-limited request, every suspension burns an account, and the repeats needed to beat sampling noise multiply both, all while the map decays underneath the campaign.

open as a page

What does a moderation record's 'clean' verdict prove about a listing written to be read by the screen?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Only that the screening stage emitted a clean verdict for that span, at that version, on that pass. It is not evidence the content was clean, that anyone read it, or that a re-screen would agree.

open as a page

A vendor screens with a small classifier and generates with a frontier model: defect or design limit?

level: principalimportance: should knowfreq 38%

basics

~20 s

It depends on what the vendor has claimed the screening layer does: against a promise of prevention it is a defect, against a promise of volume reduction a design limit. Either label needs a named owner.

open as a page

Does upgrading the generator behind an unchanged safety screen widen or close the evasion gap?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

It widens it. The leverage in a vocabulary-free restatement is the comprehension difference between the two models, so every indirection the new generator resolves that the old one could not is fresh working surface while the screen's coverage stands still.

open as a page

A listing written at a generative screen passed once in five re-runs: is that a finding?

level: seniorimportance: nice to knowfreq 29%

basics

~20 s

Yes, if you can say which stage passed it and that the rate beats the stage's own noise. One pass in five proves the construction worked against one deployment at one version — and a retryable path makes a fifth material.

open as a page

Is a durable band of replies just under a moderation screen's block line a finding, or an accepted design limit?

level: principalimportance: nice to knowfreq 27%

basics

~20 s

Every scored gate with one action point has a band beneath it, so the band alone is arithmetic, not a defect. It is a finding only if the output obtainable inside it is materially outside what the product intends.

open as a page

Your red-team report maps where a chat product's input screen is thin: what do you claim it is worth?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Claim the standing property and its price, not the list. The thin regions expire at the next update to the screening model or its acting point; what survives is that a responding screen discloses its coverage at a measured cost.

open as a page