Your harness substitutes adversarial content into an agent's tool results and the agent never misbehaves. Before reporting the agent as resistant, how do you verify the content actually reached the model?
answer
- delivered is not the same as seen
- read the rendered context, not the shim log
- canary token, then quote-back probe
- truncation, summarisation, caching, escaping
- unverified null = invalid run, not resistance
basics
~20 sLog what the model actually received, not what you handed the shim. Frameworks truncate, summarise, re-serialise or drop tool results before they become context, so your text may never arrive intact. Without that check a clean run is not evidence of resistance; it may only prove your content was cut.
solid answer
~50 sA null result from an interception harness has two very different explanations: the agent resisted, or the payload never got in front of the model. Separating them requires reading the **rendered context** -- the exact messages sent to the model on the turn after the intercepted call -- and confirming the injected text is present and intact. Things that eat a payload between substitution and context: length truncation of long tool results, summarisation of results before insertion, schema validation that rejects an unexpected field and substitutes an error, escaping that mangles the text, and framework caching that serves a previous result instead of yours. The cheap diagnostic is a **canary**: include a distinctive, harmless marker string in every substituted result and assert it appears in the rendered context, and ideally probe whether the model can repeat it back. If the canary is missing, the run is invalid rather than negative -- fix delivery and re-run before recording anything.
go deeper
Should know that a clean run might mean the payload never arrived, and that you can check by looking at what the model was actually sent.
Should name concrete ways a tool result is altered before it becomes context -- truncation, summarisation, escaping, caching -- and use a canary marker to detect it.
Should treat an unverified null as an invalid run, verify with a canary plus a quote-back probe, and report accidental truncation-based mitigation for what it is.
Insists that any resistance claim in a report is backed by delivery verification, since defenders will use such claims to close remediation work.
Negative results from an interception harness are the ones most often wrong, because the failure mode is silent. You handed a shim some text, the shim returned it, the run finished clean. Nothing in that chain establishes that the model ever read a single word of it, and no component raised an error to tell you otherwise. ### Where content disappears between the shim and the context window The gap you have to close is between *the value your shim returned* and *the rendered context* -- the exact message list sent to the model on the turn after the intercepted call. Several ordinary scaffold behaviours destroy a payload in that gap: - **Length truncation.** Most agent scaffolds cap tool-result length, often at a few thousand characters. A payload placed late in a long result is simply gone, with no error. - **Summarisation.** Some scaffolds compress tool output before insertion, often with a cheaper model. An instruction survives as a paraphrase, or not at all. - **Serialisation and escaping.** A result forced into a schema may be re-encoded; newlines, markup or delimiters that carried the payload's structure are neutralised. - **Validation and error substitution.** An unexpected shape can be replaced wholesale by a generic error string, and that string is what the model then reads. - **Caching and dedup.** A repeated identical call may be served from a result cache, so your later substitution never reaches the model. - **Wrong intercept point.** The agent used a different tool, or a different code path, than the one you wrapped -- and a harness that logs "substituted" for a call the agent never made looks identical to one that worked. ### The verification loop Emit a **canary**: a distinctive, harmless marker string inside every substituted result. Then assert three things in order. (1) The canary appears in the rendered messages for the next model turn. (2) The payload text appears alongside it, unmodified. (3) For the strongest check, run a probe turn asking the model to quote back what the tool returned -- this catches the case where the head of a long result survived truncation and the tail, carrying your instruction, did not. Also record *where* the delivered text landed: a payload sitting at the very end of a long context is a materially weaker test than the same text delivered where the workflow would naturally have returned it. ### What verification costs The canary itself is free -- a dozen tokens per result and a substring assertion. Capturing rendered context costs storage and a redaction pass, because those transcripts contain the full tool payloads. The quote-back probe is the real line item: it adds one extra model call per case, roughly a 10-15% inference increase on an eight-turn task, and it perturbs the run, so it belongs in a separate verification pass rather than inside the scored run. Wiring the context dump into a framework that does not expose the final message list is often half a day, and re-running the cases that fail delivery costs a full extra pass. Against that, the cost of skipping it is a resistance claim the defending team uses to close remediation work. ### Where the number misleads "0% attack success at the tool boundary" is the most dangerous number a harness can emit, because it is consumed as evidence of a defensive property. If the true cause is that the scaffold truncates every tool result at 2,000 characters, the honest finding is different and more useful: an *incidental* mitigation, unowned and undocumented, that disappears the day someone raises the cap or adds a long-context tool. The second misreading is aggregate: mixing delivery-verified nulls with unverified ones inflates the apparent coverage of the run. Ten verified nulls and ninety cases whose payload never arrived is not a 100-case resistance result; it is a 10-case result and a delivery bug. Report the two denominators separately -- cases attempted, and cases where delivery was confirmed. ### What you would check Before recording any null, confirm the canary in the rendered context, confirm the payload's tail with a quote-back probe, and confirm the intercepted call is the one the agent actually made. Treat an unverified null as an **invalid run**, not a negative one: fix delivery and re-run before it touches a coverage table. When delivery consistently fails for a structural reason, stop and write that up on its own -- the truncation limit, the summariser, the cache -- because that is a real property of the target that nobody on the owning team has necessarily noticed.
- The canary appears in the rendered context but the payload's tail does not. What happened, and what do you do?The result was truncated. Shorten the substituted content or move the instruction to the head, re-run, and note the scaffold's truncation as an incidental mitigation.
- Why should delivery-verified and unverified null results be recorded separately?Only verified nulls support a resistance claim. Mixing them inflates the apparent coverage of the run and lets a delivery bug masquerade as a defensive property.
A canary token is a return receipt. Without one, "I sent the letter" and "they read the letter" are the same sentence in your report, and only the second one supports the claim you are making.
saying these in an interview costs you the question
- Reports resistance from a clean run with no evidence the payload reached the model.
- Trusts the shim's own log as proof of delivery.
- Never considers truncation, summarisation or caching of tool results.
- Counts unverified nulls in a resistance rate alongside verified ones.