An input screening model's weights and threshold are private: why can an attacker still map its coverage?
answer
- the secret is the config, not the answers
- it responds every single time
- a block is a labelled sample
- an oracle you can query cheaply
- the price is requests, not insight
basics
~20 sBecause the deployed screen responds to every request. Each block, pass and near-miss is a labelled sample of its decision boundary, so a few dozen typed turns buy a usable map without seeing the model or its threshold.
solid answer
~50 sKeeping the configuration secret protects the *shortcut*, not the *signal*. A screening stage in a live product is queryable: someone types a turn, and the product returns a decision. That decision is a label, and labels are the only thing needed to fit a boundary. Vary one property of an ask at a time and the responses sort the space into stopped and not stopped; the turns that only just tip either way locate the boundary far more cheaply than turns deep inside a region whose answer was already predictable. Coverage is uneven because the screen was trained on some phrasings of a subject far more heavily than others, so thin regions exist in every trained screen. The cost of the map is requests and time, not insight into the model. Secrecy raises that price; it does not remove the oracle.
go deeper
Be ready to say that a live screening stage answers every request, and that its answers are themselves information about it. You are not expected to reason about boundaries, only to see that secrecy of a configuration is not silence.
Explain the mechanics: each response is a labelled sample, labels are all that is needed to fit a boundary, and coverage is uneven because training examples cluster. Be able to say what a stop, a pass and an answer each prove and what they do not.
Demonstrate the economics. Argue why probes near an uncertain outcome buy resolution that high volume does not, why repeats are mandatory near the boundary, and why a map fits one deployment at one moment and starts decaying immediately.
Own the framing an owner will hear. Say clearly what confidentiality of a screening model does buy, which is price and time, and refuse to let it be reported as a boundary, because a programme that believes it is one will misjudge every later finding about coverage.
## The wrong answer this corrects The common claim is: *our classifier is proprietary, the threshold is not published, nobody outside knows what model it is, so nobody can evade it.* Every clause is true and the conclusion does not follow. Nobody needs to see a screen that answers. ## A deployed screen is a queryable oracle In a self-serve chat product with no tools, the only thing anybody can do is type a turn and read what comes back. That is enough. Each submitted turn produces one of a small number of outcomes: the turn was stopped before generation, the turn reached the assistant and was answered, or the turn reached the assistant and was declined. Attributing correctly between the last two matters, but the point stands: the system returns a value for every input it is given. A value returned for an input is a **labelled sample**. Fitting a boundary from labelled samples is exactly the problem the screening model itself was built to solve; the only difference is that the person probing pays per label in requests instead of collecting a dataset up front. Nothing about weights, architecture, training corpus or the numeric acting point is required to do it, because none of those are what is being learned. What is being learned is where the decision changes. ## Why coverage is uneven in the first place A screening model learns from examples. The examples cluster: a subject appears many times in the wordings that were common in the data used to build it, and rarely or never in wordings that were not. Anything phrased far from the training distribution scores low, not because the screen decided it was acceptable but because it does not resemble what the screen learned to recognise. The result is a boundary with thin regions, and this is a property of trained classifiers generally rather than a defect in a particular one. That is why the map is worth building: the unevenness is real, and it is not visible from the configuration. ## Information per probe, not probes per hour Not every submitted turn teaches the same amount. A turn deep in a region that is obviously going to be stopped returns a label that was already predicted, and predicted labels carry almost no information. A turn whose outcome you genuinely cannot call in advance splits the space, and one that only just tips over or only just stays under locates the boundary within a much narrower band. A campaign that submits high volume without regard to where the boundary probably is spends far more requests for a coarser map than one that spends probes where the answer is uncertain. This is also why the number is small. A map good enough to be useful is not thousands of samples; a few dozen well-placed ones already tell you which framings the screen recognises and which regions it does not cover. ## What the responses do and do not prove Direction matters, and this is the part that gets marked in interviews. | Observation | What it proves | What it does not prove | | --- | --- | --- | | The turn was stopped | The text scored above the acting point | That the assistant would have refused it | | The turn passed | The text scored below the acting point at that moment | That the content is harmless, or that the assistant answers | | The turn was answered | The assistant produced a reply | That the screen ever fired, or that any stage judged the content | And because the outcome near the boundary is not deterministic, one observation of any of these is a sample, not a label. Points that will be built on need repeating, which multiplies the request cost of everything above. ## What secrecy actually buys It is worth saying plainly, because it is the honest half of the answer. Not publishing the configuration means the first probes are spent discovering things a leaked configuration would have handed over for free: whether there is a separate stage at all, roughly what it reacts to, and how sharply. That is a real price, paid in requests, accounts and elapsed time. It is a cost, not a boundary. The signal is emitted by the running system, and the only way to stop emitting it is to stop answering. ## The half-life A map fits one deployment at one moment. Retraining the screening model, moving the acting point, adding a normalising step before it or changing the assistant behind it all invalidate parts of the map without announcing which parts. So the map is perishable, while the property that produced it — a screen that responds discloses its coverage — is not.
- So what does keeping the screening model private actually buy?Price and time. Probes have to discover what a leaked configuration would have given away at once: whether a separate stage exists, roughly what it reacts to and how sharply. That cost is real and worth having, but it is a toll on the exercise rather than a barrier to it. The signal comes from the running system.
- Why does a turn that only just tips the decision teach more than one that is clearly stopped?Because information is surprise. A turn deep in a region whose outcome you could already predict confirms what you knew. A turn whose outcome you could not call splits the space, and one that only just changes the outcome pins the boundary inside a narrow band, so far fewer requests are needed for the same resolution.
- Does an uneven boundary mean the screen was built badly?No. A trained screen learns from clustered examples, so it recognises the phrasings that were common in its data and not those that were not. Uneven coverage is a general property of trained classifiers, which is exactly why the unevenness is worth mapping rather than a one-off flaw in one product.
You do not need the design of a lock to map it if you may try the handle as often as you like and it always tells you whether the door opened.
saying these in an interview costs you the question
- Claims an unseen classifier cannot be evaded
- Thinks evasion requires the weights or the numeric threshold
- Treats a pass as proof the text was judged harmless
- Assumes secrecy of the configuration is a boundary
- Says mapping needs thousands of requests to be useful