skip to content

In a stack where an input classifier screens user text before the model and an output classifier screens the reply, your attack corpus is blocked at the input stage, so the output classifier is never exercised. What are your options for testing the output classifier alone, and what does each option distort?

level: middleimportance: should knowfreq 45%

answer

  1. enter at the boundary under test
  2. record-only upstream vs direct component call
  3. record-only must not rewrite the text
  4. model refusal is not a classifier catch
  5. state the injection point beside every number

basics

~20 s

Two ways. Put the upstream classifier into a mode where it records a verdict but does not stop the request, so payloads reach the model and the reply is screened normally. Or drive the output classifier directly with fixture texts that stand in for replies. The first keeps the real path; the second skips generation entirely.

solid answer

~60 s

**Pass-through upstream.** Configure the input stage to evaluate and log but not stop, so the request runs, a real reply is generated, and the output stage screens it. This is the higher-fidelity option: the text the output classifier sees is text your model actually produced. Its distortions are that the model may simply refuse — in which case the output stage is still not exercised, only the model's own alignment is — and that the reply differs between runs, so a case can pass one day and fail the next. **Direct injection at the boundary.** Call the output classifier with crafted texts in the position a reply would occupy. This is cheap, deterministic and gives you full control of what the classifier sees, so it measures the classifier's decision surface well. What it does not measure is whether your model would ever emit such text, so its numbers are a property of the classifier, not of the deployed system, and must be reported that way. Most engagements need both: direct injection to characterise the layer, pass-through to check that the layer sits where you think it does.

go deeper

for a junior

Should know that if an upstream layer blocks everything, the downstream layer is untested, and that you have to let the payload through or call the layer directly.

for a middle

Names both entry points and the tradeoff: real path but noisy and possibly refused, versus deterministic and controllable but not evidence the text ever occurs.

for a senior

Adds the traps — a non-blocking stage that still rewrites text, refusals miscounted as catches, generation variance, and labelling every number with its injection point.

for a principal

Weighs whether a full stack-minus-one environment is worth maintaining at all, given drift, versus accepting component-level numbers with stated caveats.

### The general move Decide which boundary the payload has to *enter at*, and enter there. Everything upstream of that boundary stops being the thing under test and becomes scaffolding for this run — its job is to deliver your case to the layer you are measuring, unchanged. For an output-side classifier (a screen that reads the model's generated reply and may suppress it) there are two entry points, and a third option that looks attractive and usually is not. ### Entry point 1 — pass-through upstream Put every upstream stage into a posture where it evaluates and records a verdict but does not stop the request (record-only, shadow, audit, `detect`-not-`block`). The case then reaches the model, a real reply is generated, and the output classifier screens a text your own system actually produced. This is the higher-fidelity option, and it is the only one that shows the layer sitting where you think it sits. Its distortions are specific. **The model may simply refuse.** Then the output classifier fires on nothing, and what you have measured is the model's own alignment, not the layer — an easy result to misread as "the output classifier saw a clean corpus". **Generation is sampled**, so the same case yields different replies run to run; one sample per case is a coin flip wearing a percentage sign. **A record-only stage may still rewrite.** If it redacts, normalises or wraps the user text while declining to block, the model receives a different prompt than production would send, and the reply — hence the downstream layer's input — comes from a path that does not exist. ### Entry point 2 — direct injection at the classifier boundary Call the output classifier as a component, handing it texts you control in the position a reply would occupy. This is cheap, deterministic and fully controllable: it is the only way to characterise the layer near its decision boundary, to exercise harm categories your model rarely produces, and to get a number stable enough to regression-track across guard versions. Replicate the real call shape — the role or turn slot, the system instructions, the category taxonomy, the threshold — because a guard prompted with a reply in the *user* slot is answering a different question than one prompted with it in the *assistant* slot. What it does not measure is reachability: whether your model would ever emit such a text. The resulting number is a property of the classifier, and must be reported as one. ### Entry point 3 — a stack-minus-one environment Rebuild the deployment without the layer, in a test environment. It is the truest simulation of "what if we dropped this", and it doubles the environments someone has to maintain; drift between test and production then quietly invalidates every number the environment produces. Worth it when a removal decision with real money attached is genuinely in flight, rarely otherwise. ### What each option costs Direct injection costs classifier calls only: hundreds of calls, cents to low dollars, minutes of wall-clock, no generation spend at all. Pass-through costs one generation per case *per sample* — and because you need k samples to see variance, a 500-case corpus at k=5 is 2,500 generations plus 2,500 classifier calls, hours of serialized wall-clock against a rate-limited production tier. Then there is the part that does not appear on any invoice: getting a client to put a production safety layer into record-only mode is a change request with an approval chain, and it is frequently refused outright, which is why component-level numbers are so common in real reports. ### Where the numbers mislead That cost asymmetry is the trap. Direct injection is a hundred times cheaper, so it is what gets run, and its output is a headline percentage that reads exactly like end-to-end coverage. A "96% block rate" obtained by direct injection has a denominator you chose: texts you authored, at a difficulty you selected, with no natural distribution behind them. Move the fixtures toward the boundary and the same layer scores 60%; make them blatant and it scores 100%. The number is only as adversarial as the fixtures, and it says nothing about whether any of those texts is reachable. The second misreading is crediting a pass-through run's model refusals to the classifier, which inflates the layer's apparent contribution with work the model did on its own. ### What you would check Verify record-only mode really passes text through byte-identical, with a case you know that stage blocks. Log model refusals in their own column, never merged with classifier blocks. Report a variance estimate, not a single sample, for anything downstream of generation. Replicate the production call shape for direct injection, and diff one direct-injection verdict against a pass-through verdict on the same text to prove the component you are calling is the one deployed. Above all, record the **injection point in the results table itself**, on every row — rows get filtered, sorted and pasted into other decks, and a component-level rate that loses its label becomes an end-to-end claim in someone else's slide.

  • During a pass-through run the model refuses most of the corpus, so the output classifier fires on almost nothing. What does that result mean?
    It means the model's own behaviour is carrying that corpus, and the output classifier is untested by this run. Report it as such, and characterise the classifier by direct injection instead of inferring it is weak.
  • Why record the injection point in the results table rather than in the report's prose?
    Because rows get filtered, sorted and copied into other decks. A component-level block rate that loses its 'injected directly at the classifier' label will be read as end-to-end coverage by whoever sees it next.

saying these in an interview costs you the question

  • Reporting an output-classifier block rate without saying where the text was injected.
  • Counting a model's own refusal as an output-classifier catch.
  • Assuming a non-blocking upstream stage leaves the text untouched when it redacts or rewrites.
  • One generation sample per case, then treating the result as stable.
  • Only ever testing the classifier as a component and presenting the number as end-to-end coverage.

context