skip to content

A defence you are engaged to test has four screening stages: an input classifier, a rules layer, a system-prompt instruction, and an output classifier. One isolation run per stage multiplies your inference bill and your engagement hours. How do you decide how far to decompose, and what do you hand the owner at the end?

level: principalimportance: should knowfreq 25%

answer

  1. buy a run only to change a decision
  2. group the stages nobody is questioning
  3. baseline plus benign half in every run
  4. table: unique catch, false positives, cost, how measured
  5. redundancy value is not in the matrix

basics

~20 s

Decompose only where a decision hangs on the answer: a stage someone wants to remove, or one that costs latency or money. Group the rest. Budget one end-to-end baseline plus an isolation run per stage you are actually questioning, and hand back unique-catch per stage against its cost, not four separate pass rates.

solid answer

~60 s

Decomposition is not free and it is not equally valuable per stage. Start from the decisions in flight: someone is proposing to drop a stage, someone is paying per call for one, one adds latency the product team resents, one was added after an incident and nobody has re-checked it. Those stages get their own run. Stages nobody is questioning can be measured as a group — a single 'everything else' column is a perfectly good matrix entry. Budget shape: one end-to-end baseline (non-negotiable, it is the only observation of the real path), plus one run per stage under question, plus the benign half in every one of them. Buy back runs where layers read the same fixed input by scoring them offline; spend real runs only at boundaries where the system generates the next input. The deliverable is not four pass rates. It is, per stage: unique catch on this corpus, false positives on the benign half, and its cost — latency, spend, operational burden. That is the table a removal decision is actually made from, and it should say explicitly which cells were measured and which were grouped.

go deeper

for a junior

Should recognise that testing every stage separately costs several full passes over the corpus, so the number of runs has to be planned rather than discovered.

for a middle

Argues for a baseline plus targeted per-stage runs, keeps the benign half in each, and reports per-stage results rather than one aggregate.

for a senior

Chooses the decomposition from the decisions in flight, groups correlated stages, and buys back runs with offline scoring where layers share a fixed input.

for a principal

Owns the tradeoff end to end: run budget fixed in the statement of work, a deliverable framed as unique catch against cost, redundancy value stated separately, and every number labelled corpus-relative.

### The framing that stops this becoming a completeness exercise **An isolation run is bought to change a decision.** If nobody's action depends on knowing a stage's separate number, the run produces expensive documentation. Full decomposition of a four-stage defence is five passes over the corpus — a baseline plus one per stage — and the reflex to "measure everything" is how these engagements overrun. ### Choosing the decomposition *Start from the decisions in flight.* A stage under cost pressure because it is metered per call. A stage the product team wants removed for latency. A stage added after an incident that nobody has re-checked since. A stage whose ownership is unclear. Each of those is a live question with a named audience, and each earns its own run. *Group by kind, not by count.* Stages that share a detection principle correlate, so measuring them separately buys little — a single "everything else" column is a perfectly good matrix entry. Complementarity lives between stages that differ in kind: one reading intent in the request, one reading the realised content of the reply, one that is a static instruction rather than a classifier at all. That is where separation is worth paying for. *Treat the system-prompt instruction as its own case.* It is the cheapest stage to delete and the hardest to attribute, because it emits no verdict. Its effect appears as a shift in what the model generates, entangled with the model's own behaviour. Isolating it means running the corpus with and without the instruction and comparing outcome distributions, and because generation is sampled you need repeated samples per case to see a shift at all — a k of 5 turns one pass into five. Plan for that number explicitly or leave the stage grouped. *Cap it in the statement of work.* Say up front how many isolation runs the budget buys and which stages they cover. ### What the decomposition costs Concretely, for a 500-case attack half plus a 300-case benign half: one pass is 800 target interactions, and full decomposition is five passes, so ~4,000 generations plus per-stage classifier calls. Serialized at a couple of seconds each that is hours per pass, and production rate limits usually cap the concurrency that would fix it. Offline scoring buys back the passes for stages that read the same fixed user text, but it does **not** buy back generation: any boundary where the system produces the next input needs a real run. The system-prompt comparison multiplies whatever remains by k. Engineer hours are the largest line item and the least visible: each stage that must be put into record-only mode without altering what the next stage receives is harness work plus, on a client's production stack, a change ticket. ### What you hand the owner One table, one row per stage or stage *group*, with four columns: | column | what it is | |---|---| | unique catch | cases only this stage blocked, on this corpus, with an interval | | benign impact | benign cases this stage stopped — the stage's real cost to users | | cost | latency, spend, operational burden | | how measured | real run / offline scoring / grouped | Alongside it, two things a table cannot carry. The **residual set** — the cases nothing stopped — which is the finding, and which should lead the report. And a plain statement of **redundancy value**: a stage with zero unique catch still covers its neighbour's bad model update, a misconfiguration during a deploy, or a provider outage. A report that recommends deleting a stage on unique catch alone has given the owner half the picture. ### Where the numbers mislead Four readings to head off. **Per-stage pass rates presented as contributions** — an isolated stage blocking 90% may contribute nothing marginal, and the aggregate invites exactly the wrong removal decision. **Grouped columns read as single stages**: if three stages share a column, no row in that column is about any one of them, which is why the "how measured" field is not optional. **Zero unique catch read as failure**: it usually means high overlap, not weakness. **Corpus relativity**: every number answers "what would removal have cost against the attacks we sourced", never "against attacks in general", and the gap between those two sentences is where the next engagement starts. ### What you would check That the benign half rode along in every run, or the false-positive column is empty and the cost side of every removal argument is missing. That the baseline was preserved and the per-stage results reconcile with it. That the corpus was byte-identical across passes. And that the single number the owner will inevitably ask for — "how effective is the stack" — is answered with the end-to-end baseline plus a one-line description of the corpus and the residual set, not by averaging the per-stage table, since compressing to one number is what made the stack unattributable in the first place.

  • The owner asks for a single number: how effective is the stack. What do you give them?
    The end-to-end baseline against the stated corpus, with the corpus described in one line, plus the residual set. Refuse to compress the per-stage table into it — a single number is what made the stack unattributable in the first place.
  • Which stage in that list is hardest to isolate, and why?
    The system-prompt instruction. It produces no discrete verdict to read; its effect shows up as a shift in what the model generates, so isolating it means comparing outcome distributions with and without it rather than counting blocks.

saying these in an interview costs you the question

  • Decomposing every stage as a matter of completeness without asking what decision it serves.
  • Discovering the run count only after the engagement has started.
  • Delivering per-stage pass rates instead of unique catch against cost.
  • Recommending stage removal on unique catch alone, ignoring redundancy under failure.
  • Omitting the benign half, so the false-positive cost of each stage is unknown.
  • Presenting corpus-relative numbers as general properties of the defence.

context