skip to content

A deployment screens each user message with an input classifier and then screens the model's reply with a separate output classifier. You run an attack corpus end to end against that live path and every case is blocked. Why does that green result not tell you which of the two classifiers earned the block, and what test design does?

level: juniorimportance: must knowfreq 55%

answer

  1. stack short-circuits at first block
  2. one run per boundary, fixed corpus
  3. per-case per-layer verdict matrix
  4. not-reached is not the same as not-blocked
  5. N layers, N+1 passes

basics

~20 s

End to end you only see the stack's final verdict, so the first layer that blocks hides every layer behind it. To attribute, run the corpus once per layer with the others put into log-only mode, and record per case which layer fired. Then a green run says which classifier actually earned it.

solid answer

~50 s

An end-to-end run measures the **stack**, and a stack short-circuits. If the input classifier refuses a case, the model is never called and the output classifier never sees anything, so its verdict for that case simply does not exist. A pass rate of 100% across the stack is therefore compatible with one layer doing all the work and the other contributing nothing. The design that attributes is one run per boundary: keep the corpus fixed, and for each layer in turn arrange for that layer to be the only one that can stop a request — the others record a verdict but do not block. You end up with a per-case, per-layer verdict table instead of a single pass/fail column. The cost is real and should be stated up front: N layers means roughly N+1 passes over the corpus, N+1 times the inference spend and wall-clock, and a harness that can put each layer into a non-blocking mode without changing what the next layer receives.

go deeper

for a junior

Should say that the stack stops at the first block, so an all-green end-to-end run cannot tell you which component blocked, and that you need a run per layer.

for a middle

Adds the mechanics: fixed corpus, one layer blocking at a time, a per-case matrix, and the distinction between 'not blocked' and 'never reached'.

for a senior

Adds cost and fidelity — the extra inference passes, keeping the benign half in every run, and checking that non-blocking mode does not change what downstream layers receive.

for a principal

Frames it as what the defence owner can decide with: which layers earn their latency and spend, and how much isolation work the engagement should buy at all.

### What the green run actually licenses An **end-to-end run** means pushing every case of a fixed attack corpus — a hand-labelled set of prompts you expect the product to refuse — through the deployment the way a user would, and recording only what a user could see: did the request complete, or was it stopped. That observation licenses exactly one claim: *this corpus, against this whole configuration, on this day, produced no visible failure.* It licenses none of the three claims a defence owner actually wants: which layer stopped each case, what dropping a layer would cost, and whether some layer is contributing nothing at all. ### The mechanism: a screening stack short-circuits A layered defence is an OR over block decisions evaluated in a fixed order, and evaluation stops at the first block. Make it concrete with the two layers in the question. The **input classifier** scores the user's message before the LLM is called — a hosted moderation endpoint, a small classifier model, or a rule set. The **output classifier** applies the same idea to the generated reply. When the input classifier refuses a case, the LLM is never invoked, no reply is ever produced, and the output classifier is never called. Its verdict for that case is not `pass` and not `fail`; it is a **missing cell**. So a 100% end-to-end block rate is fully compatible with layer one doing all the work while layer two is a no-op — misconfigured, pointed at a stale endpoint, or disabled by a flag nobody has read in six months. It is also compatible with the reverse. The run cannot distinguish these worlds, because the evidence that would distinguish them was never generated. Order matters for the same reason: reverse the two layers and the same corpus can produce a completely different attribution, with identical end-to-end output. ### The design that attributes: one run per boundary Hold the corpus byte-identical across runs, then: 1. Run the baseline end to end. This is the only observation of the real deployed path and everything else has to reconcile with it. 2. For each layer in turn, arrange for that layer to be the only one permitted to stop a request. Every other layer still evaluates and records its verdict but does not block — vendors call this record-only, log-only, shadow, audit, or `detect` rather than `block`. 3. Store a matrix: rows are corpus cases, columns are layers, and each cell takes one of **three** values — `blocked`, `not blocked`, `not reached`. From that matrix you can compute per-layer catch, per-layer *unique* catch (cases only that layer stopped), and the residual set nothing caught — which is the finding, and which an end-to-end run also hides whenever a later layer happens to cover it. ### What it costs N layers means roughly N+1 passes over the corpus. Each pass with a live target is one generation call per case plus one call per classifier per case. A 500-case corpus against a two-layer stack is ~1,500 target generations across three passes, plus classifier calls; if either classifier is a metered hosted service, that line multiplies too. Wall-clock usually hurts more than spend: a serialized pass at a couple of seconds per case is tens of minutes, and production rate limits often cap the concurrency that would fix it. The largest line item is normally engineer time — the harness ability to put one layer in record-only *without altering what the next layer receives* is real work, and on a client's production stack it is a config change, which means a change ticket, not an afternoon. ### Where the number misleads Three readings to refuse. First, **"the guardrails held"**: the *stack* held, on this corpus; any per-layer claim drawn from an end-to-end run is invented. Second, **recording an unreached layer as a pass** — the natural data shape in a home-grown harness is one boolean ("was the request blocked?") broadcast to every column, which silently pushes every short-circuited layer's measured catch toward 100%. The mirror error, recording it as "not blocked", deflates the layer instead; both come from having two values where the measurement needs three. Third, even a clean isolated block rate is not that layer's contribution to the deployed system: two layers at 90% each may be catching the same cases, so neither is worth ten points on removal. ### What you would check before believing it Send a case you know a layer blocks and confirm that in record-only mode it actually reaches the next stage — quiet is not the same as non-blocking. Confirm no record-only layer redacts, normalises or wraps the text, because a rewriting layer changes the prompt and you are now measuring a path that does not exist in production. Keep the benign half of the corpus in every run, since isolating a layer also changes its false-positive count. Count model self-refusals separately from classifier blocks. Finally, recompose the per-layer columns with an OR and compare against the baseline: if they disagree, your isolation changed the system's behaviour and the matrix is describing something other than the deployment.

  • Your matrix has a case where the input classifier blocked and the output classifier's cell is empty. What do you write in the report?
    That the input classifier blocked it and the output classifier is untested for that case — an empty cell is missing data, not a second success. If output-side coverage matters for it, it needs its own run where the payload is allowed through.
  • Why keep the end-to-end baseline run at all once you have per-layer runs?
    It is the only observation of the real deployed path. The isolated runs each alter the configuration, so the baseline is what you check the reconstruction against — if the per-layer results imply a block the baseline did not show, your isolation changed behaviour.

It is like testing a building by lighting one fire in the lobby: the lobby door holds, so nobody ever learns whether the fire doors upstairs would have. Their logbooks are empty, not clean.

saying these in an interview costs you the question

  • Treating a 100% end-to-end block rate as evidence that every layer works.
  • Reporting 'the guardrails held' without naming which component produced each block.
  • Recording a layer that was never reached as a pass for that layer.
  • Changing the corpus between the isolated runs, making the columns incomparable.
  • Assuming isolation is free — not budgeting the extra inference passes or saying so in the plan.

context