skip to content

You replay a published red-team benchmark's prompt set against your deployed assistant endpoint instead of against the bare model API. What is one mechanism that pushes the attack-success number DOWN compared with the published run, and one that pushes it UP?

level: middleimportance: must knowfreq 60%

answer

  1. wrapper cuts both ways
  2. filters and scope prompt push down
  3. retrieval and tools push up
  4. log where each item terminated
  5. delta is not the wrapper's contribution

basics

~20 s

Down: your system prompt narrows what the assistant will discuss and an input or output classifier sits in the path, so items that landed on the bare model get deflected. Up: retrieval carries text you did not write, tools turn a compliant answer into a real action, and a long preamble dilutes plain refusal behaviour.

solid answer

~60 s

Both directions are real, which is why the published figure is not a bound. **Downward pressure** comes from everything the deployment adds in front of the model: a scoped system prompt that gives the assistant a job unrelated to the item, a classifier on the input, a classifier or redaction pass on the output, and short output limits that truncate anything long. Items whose success depended purely on the model's own tuning now die before or after the model. **Upward pressure** comes from everything the deployment adds around the model: retrieved documents and tool results enter the same context as trusted-looking text, tools make a text-level hit into an action with a consequence, per-tenant data raises what a hit is worth, and a large fixed preamble shifts the model away from the distribution its safety tuning was measured on. The practical consequence: report the replayed slice as its own measurement with its own denominator, and never subtract it from the published number as if the difference were 'the wrapper's contribution'.

go deeper

for a junior

Gives one plausible mechanism in each direction, typically the system prompt for down and tools for up.

for a middle

Orders the downward mechanisms by where they act in the request path and names retrieval as a distinct inbound channel on the upward side.

for a senior

Insists on per-item termination logging and item-level flip analysis so the aggregate is decomposable, and refuses to read the delta as the wrapper's contribution.

for a principal

Sets the standard that any application-boundary run ships with a stage breakdown and a realistic tool sandbox, otherwise it is not accepted as evidence.

**Why both directions must be named.** A candidate who lists only the downward mechanisms believes the application layer is a safety net. A candidate who lists only the upward ones believes it is pure liability. Both halves are true at the same time, and that simultaneity is exactly what makes the difference between the published number and your replayed number uninterpretable. The published run drove a bare chat-completions endpoint with one message in the array; your run drives a product. Several variables moved at once. **Downward mechanisms, in the order they act on a request.** 1. **Input classifier.** A separate small model or hosted moderation call returns a verdict before your orchestrator ever calls the model. Items it rejects never reached the model at all, so that share of your "safety" is a filter configuration, changeable by one pull request. 2. **Scoped system prompt.** The assistant has a job — invoice questions, product support — unrelated to most published items, so off-domain requests get a stock deflection regardless of the model's own tuning. 3. **Output classifier or redaction.** A post-generation pass rewrites or blocks the completion. What you then measure is partly the sanitiser. 4. **Shape constraints.** Maximum output length and a required response schema truncate or box in exactly the long, discursive completions many items need to count as a hit. **Upward mechanisms.** - **A second inbound text channel.** Retrieved documents and tool results land in the same context window as the user's turn. The published run had no equivalent, because the model API accepts one channel. - **Consequence.** When the app can call tools, identical response text is worth more: a hit stops being a sentence and becomes an action with a blast radius. - **The system prompt is itself an object worth extracting**, since it carries scope rules, tenant identifiers and connector names that did not exist during the published run. - **Distribution shift.** A long fixed preamble plus conversation history moves the model away from the short single-turn shape both the suite and much safety tuning were measured on. - **State.** A session gives multi-turn build-up that a single-turn corpus cannot express, so your surface is strictly larger than the one that was scored. **What measuring both directions costs.** Honestly separating them takes two arms, not one: a production-shaped arm with every layer on, and a staging arm with app-side filters disabled but everything else identical. That doubles both the application calls and the judge calls for the slice — a 300-item slice at 3 samples becomes 1,800 application calls and 1,800 judge calls — and it requires a staging deployment that is config-identical apart from the filters. Keeping those two configurations in step is where the engineer-days actually go. Then the tool sandbox: no-op stubs are nearly free and make every upward mechanism invisible to the run, while a sandbox faithful enough for an action to register is usually the single most expensive component of the harness. **Where the number misleads.** - *Published minus replayed equals what our safety stack bought.* A confound. Interface, context length, output shape, available actions and possibly judge behaviour all changed together; no arithmetic separates them. - *A low replayed rate means the model is safe.* It may mean the input classifier absorbed the items. You would be publishing a filter's current configuration as though it were a property of the model. - *An unchanged rate means the wrapper is neutral.* An aggregate is a net. Equal numbers of items can have flipped hit-to-miss and miss-to-hit, which is a completely different system from one where nothing moved. - *Silent denominators.* Items rejected pre-model, rate-limited, truncated or errored must appear in a stated denominator. Dropping them computes a rate over an undefined set. **What you check afterwards.** Per-item termination stage, so the aggregate is decomposable into filter, model and output pass. Item-level flip lists in both directions, not just the total. A judge-agreement figure measured on the application arm's response format. Counts of retries, 4xx/5xx and truncations, so infrastructure noise is not silently read as refusal. And the filter configuration recorded with the result — ideally a config hash — because the number is only true for that configuration on that day.

  • Your replay shows a much lower rate than the published run. What single log field tells you whether that is the model or the input filter?
    A per-item termination stage - whether the request was rejected pre-model, answered by the model, or rewritten post-generation. Without it you cannot tell a safe model from a strict filter.
  • Why is the difference between the two numbers not a valid measurement of your safety layers?
    The two runs differ in interface, context length, output shape and available actions all at once, and the judge may behave differently on each. Attributing the whole delta to one layer is a confound.

saying these in an interview costs you the question

  • Naming only downward mechanisms and concluding the deployment is strictly safer than the model.
  • Interpreting the difference between the published number and the replay as a measurement of the wrapper.
  • Running the replay with tools stubbed to no-ops and still reporting it as an application-boundary result.
  • Not logging which layer stopped each item, leaving the aggregate undecomposable.

context