skip to content

To avoid paying for one full corpus run per layer, an engineer proposes scoring the attack corpus against each moderation layer offline and OR-ing the verdicts to predict what the deployed stack would do. Where does that reconstruction hold, and where does it break?

level: seniorimportance: nice to knowfreq 30%

answer

  1. what is each layer a function of
  2. same fixed input, OR is valid
  3. generated input breaks it
  4. existence, non-determinism, coupling
  5. hybrid: one real pass plus offline scoring

basics

~20 s

It matches for layers that read the same fixed user text, like two input-side classifiers: OR-ing their independent verdicts predicts the stack. It breaks for anything downstream of generation, because the output classifier reads a reply that only exists if the request ran, varies between runs, and changes if an upstream layer rewrote the input.

solid answer

~50 s

The reconstruction assumes every layer is a pure function of the same fixed input. That assumption holds for the pre-model layers: several classifiers all reading the user's message can each be scored offline once and combined, and the OR is exactly what the deployed stack computes, at one pass over the corpus instead of N. It fails the moment a layer's input is produced by the system rather than supplied by you. An output-side classifier reads a generated reply, and that reply exists only if the request ran, is not stable across samples, and is a function of whatever the upstream stages left of the prompt. Scoring it offline against a stand-in reply measures the classifier, not the stack. So the practical rule: reconstruct across layers that share one fixed input, and pay for real runs at every boundary where the system generates the next input. That usually collapses N+1 runs to two or three, which is where the actual saving is.

go deeper

for a junior

Should recognise that you cannot score an output-side check without a reply to score, so the shortcut cannot cover that layer.

for a middle

Separates layers by what they read: fixed user text can be scored offline and OR-ed; anything reading a generated reply cannot.

for a senior

Gives the three failure reasons — existence, sampling variance, prompt coupling — and proposes the hybrid that keeps most of the saving.

for a principal

Decides how much fidelity to trade for run budget, and requires every matrix cell to be labelled with how it was obtained so later readers do not mix the two.

### Frame it as: what is each layer a function of? **Offline scoring** here means calling a screening layer directly, outside the serving path, on stored texts, and keeping its verdict. The proposal is to do that once per layer and then combine the columns with a boolean OR, on the theory that the deployed stack blocks a case exactly when *some* layer blocks it. That theory is correct only when every layer is a **pure function of the same fixed input** — same text in, same verdict out, no dependence on anything the system produced along the way. ### Where it holds: layers upstream of generation Several classifiers all reading the user's message satisfy the assumption. The text is yours, it is byte-identical in every run, the layers do not see each other, and their verdicts do not depend on order. Score each once over the corpus, keep the per-case verdicts, and the OR of the columns *is* the pre-model decision the deployed stack computes. The saving is real: instead of N+1 full passes you make one classifier call per (case, layer) with no generation at all. Two checks keep it honest. No upstream layer may mutate what later ones read — a stage that redacts or normalises makes the layers an ordered composition rather than independent functions, and the OR silently models the wrong pipeline. And no layer may carry conversation state: a defence that scores a turn in the context of the conversation is not a function of a single message, so a per-message reconstruction does not describe it. ### Where it breaks: layers downstream of generation An output-side classifier reads a reply the system produced. Three independent reasons the offline OR is wrong there: - **Existence.** For every case an upstream layer blocks, production never generates a reply at all. An offline verdict on a hypothetical reply predicts a decision the deployed system never makes, so the reconstructed column claims coverage of events that cannot occur. - **Non-determinism.** Sampled generation gives a different reply per run. Scoring one stored artefact fixes a single draw and reports it as the answer; the deployed behaviour is a distribution whose variance the reconstruction cannot show. - **Coupling.** If any upstream stage rewrites, redacts or prepends to the prompt, the reply distribution shifts. Your stored reply came from a prompt production would not have sent. There is a fourth, purely mechanical trap: guards that take a role-tagged conversation — Llama Guard is the canonical example — classify a text differently depending on whether it occupies the user turn or the assistant turn of the template, because the taxonomy and the question differ by role. Score candidate replies in the user slot and the whole column is a different measurement wearing the right label. ### The honest design, and what it saves A hybrid, stated in the plan: **one** real pass that generates replies through the actual path with upstream layers in record-only mode, plus offline scoring for every layer that reads the fixed user text. You keep the per-case matrix and buy back most of the passes. The saving is concentrated where the money is — generation calls dominate the bill; classifier calls are often an order of magnitude cheaper per case, and a local guard model costs GPU minutes rather than API dollars. Collapsing five passes to two on a 500-case corpus removes roughly 1,500 generations and the hours of rate-limited wall-clock that come with them, at the price of a matrix whose cells were obtained by different methods. Which is the last cost item: every cell must be **labelled with how it was obtained** — real run, offline scoring, grouped — or a later reader mixes reconstructed and observed numbers into one table and the provenance is gone. ### Where the number misleads The reconstruction predicts *block verdicts*, and people read it as predicting the system. It does not carry latency, it does not carry per-request cost, it does not carry ordering effects, and for anything after generation it does not carry reachability. A reconstructed "the stack would block 482/500" is a claim about a pipeline that exists only in the spreadsheet. The seductive confirmation is agreement with the baseline. If the reconstruction exactly matches the end-to-end run, that is weak evidence at best: the two methods can only differ on cases where the layers *disagree*, so if one layer decides nearly everything, agreement is guaranteed regardless of whether the method is sound. Validate on the disagreement subset specifically, and report the size of that subset — a reconstruction validated on twelve cases is a reconstruction validated on twelve cases. ### What you would check Diff a sample of offline verdicts against verdicts observed in the real path for the same texts, including the call shape: role slot, system instructions, category taxonomy, threshold and model version. Confirm no upstream layer mutates text or holds state. Confirm the guard version used offline is the version deployed — a guard upgrade between the two invalidates every reconstructed cell. And keep the end-to-end baseline: it is the only evidence that the pipeline you modelled is the pipeline that runs.

  • Your reconstruction exactly matches the end-to-end baseline. Does that validate the method?
    Only weakly. The two can only differ on cases where the layers disagree; if one layer decides nearly everything, agreement is guaranteed regardless of method. Check the disagreement subset specifically.
  • What breaks the reconstruction even for two input-side layers?
    An upstream layer that mutates the text the next one reads, or any layer that carries conversation state. Both make the layers non-independent functions of the single fixed message, so the OR no longer models the ordered path.

saying these in an interview costs you the question

  • OR-ing offline verdicts across a boundary where the system generates the next input.
  • Scoring an output-side classifier against one sampled reply and calling it the stack's behaviour.
  • Ignoring that an upstream stage which rewrites text destroys the independence the OR assumes.
  • Applying single-message reconstruction to a defence that keeps conversation state.
  • Never checking the reconstruction against the end-to-end baseline.

context