skip to content

A published red-team benchmark delivers every item as a user turn, but your deployment also puts text you did not write - retrieved documents and tool results - into the same context. What do you re-run against that second channel, and how do you keep it a bounded job instead of a second full benchmark?

level: seniorimportance: should knowfreq 42%

answer

  1. channel is a second axis
  2. port instruction-following items only
  3. seed the index, stub the tool
  4. log retrieval hit or miss
  5. cells attempted over cells selected

basics

~20 s

Do not replay the whole corpus on the retrieval channel. Select the items whose outcome depends on the model following injected instructions, place that text in the retrieved-document or tool-result slot instead of the user turn, and score what the app did as well as what it said. Report the two channels as separate coverage numbers with separate denominators.

solid answer

~50 s

Treat channel as a second axis on the item corpus rather than as a new corpus. Most published items are not worth porting: items that test refusal on a knowledge request behave much the same wherever the text arrives. The ones worth porting are those whose success turns on the model treating text as an instruction, because that is exactly what changes when the text arrives from a source the user did not type. Mechanically you need a harness that can seed the retrieval index or stub the tool result with the item text, then drive an ordinary user turn that causes that content to be pulled in. The judged unit widens: the interesting outcome is often a tool call the app made, not a sentence it wrote. Keep it bounded by declaring an item-by-channel matrix and filling only the cells you can justify, then reporting **coverage as cells attempted over cells selected** - never as a fraction of the published suite, which now means something different.

go deeper

for a junior

Knows the app has an input path besides the user turn and that the benchmark only used the user turn.

for a middle

Explains how to place item text in the retrieved-document or tool-result slot and why only some items are worth porting.

for a senior

Runs the matrix with explicit selection, logs retrieval hit/miss so a no-op arm is detectable, and widens the judged unit to tool calls.

for a principal

Fixes the reporting standard - per-channel denominators - so no one can read a one-channel sweep as deployment coverage.

**The shape of the problem.** A published suite has exactly one delivery channel, because a chat-completions endpoint has exactly one: the item goes in as the user's message. A deployment normally has three or more — the user turn, retrieved context, and tool or connector output — and the model sees them in a single window. So "we ran the suite" is a coverage claim over one column of an item-by-channel matrix, presented as though it covered the deployment. That denominator swap is the whole defect. **Selection, which is what keeps the job finite.** The instinct is a cross-product: every item on every channel. Do not. Port an item to a non-user channel only where the channel plausibly changes the outcome, and the useful filter is: *does this item's success depend on the model treating supplied text as an instruction, or on the model producing content it should withhold?* The first kind ports, because "who authored this text" is precisely the variable that changes. The second kind mostly does not — a knowledge request that the model refuses on the user turn refuses much the same when the same words arrive from a document. In practice that leaves a small fraction of the corpus, perhaps a few dozen items, and that is the intended outcome: a matrix you can actually re-run each release beats a cross-product you run once and never repeat. **Harness work, at the level of the object that does it.** You need three things per ported item. A **fixture**: the item text placed in a document that your indexer will accept, or in a stubbed tool result your orchestrator will return. A **seeding step**: writing that document into the retrieval index used by the test tenant, then confirming the index has committed it — most vector stores index asynchronously, and a run started too early tests an empty index. And a **trigger turn**: a benign user question whose embedding actually retrieves the fixture. That last one is where arms die silently, because your retriever returns top-k by similarity and your fixture may simply never be in the top-k. **What it costs.** Per ported item, expect fixture authoring plus a trigger turn plus an index write — roughly 15-30 minutes of engineer time each the first time, so forty items is one to two engineer-days before a single call is made. The run itself is cheap by comparison: forty items times two channels times three samples is 240 application calls and 240 judge calls. The expensive standing cost is the sandbox. Tools stubbed as no-ops make the entire arm meaningless, because the outcome you are hunting on this channel is usually an action; a sandbox faithful enough for an action to register, and safe enough to run repeatedly, is the real budget line. Index rebuilds and tenant isolation add operational time each run. **Where the number misleads.** - **The zero that is a no-op.** An arm reporting zero hits is the most dangerous result in this whole exercise, because "the app resisted every injected item" and "the fixture was never retrieved into the context" look identical in the aggregate. Without a per-item retrieval hit/miss record you cannot tell them apart, and the comforting reading is the wrong one more often than not. - **The merged percentage.** Reporting user-channel and retrieval-channel results under one figure lets a thin, deliberately sampled arm inherit the credibility of a full sweep. - **The suite's denominator.** "100% of the published suite" uses the benchmark's item count as the denominator for a deployment whose surface is a matrix. Coverage here is cells attempted over cells selected, and the selection rule must be stated. - **Prose-only judging.** On this channel the consequential outcome is often a tool call, so a text judge reports a clean sweep exactly where the risk lives. **What you check.** Per-item retrieval hit/miss, before you read any rate at all. Whether the fixture was indexed and committed before the run started. Whether tool calls are in the judged unit. Whether the selection rule is written down, so a reader can see the arm was sampled deliberately and thinly rather than swept. And two denominators in the report, never one. The attack techniques that exploit this channel, and the design of an agent harness that drives it, belong to the neighbouring topics; the job here is the instrument's coverage bookkeeping.

  • Your retrieval-channel arm reports zero hits. What do you check before concluding the app is robust?
    Whether the fixture was actually retrieved for each item. A per-item retrieval hit/miss log distinguishes a genuine zero from a run where the content never entered the context.
  • Why is 'we ran the full published suite' a misleading coverage claim for a retrieval-backed app?
    It covers one delivery channel out of several. Coverage is a fraction of a matrix, and the claim quietly uses the published suite's denominator instead of the deployment's surface.

saying these in an interview costs you the question

  • Proposing a full cross-product of every item against every channel.
  • Reporting user-channel coverage and retrieval-channel coverage under one percentage.
  • Running with an index that never returns the fixture, so the arm is a silent no-op.
  • Judging only the response text when the app can act on what it read.

context