An injected supplier email took ten drafts, mostly learning the buyer's own routing vocabulary — what follows?
answer
- count what the drafts were spent on
- reconnaissance cost, not model cost
- target-specific, not model-specific
- who else has already paid that cost
basics
~20 sThe cost sat in reconnaissance, not in model behaviour, so the construction is target-specific rather than model-specific. It is unlikely to port to another buyer, more likely to survive a model change, and by design it reads as ordinary correspondence.
solid answer
~50 sCount what the drafts were spent on. If they went into matching the buyer's vocabulary, the state of the thread and what the workflow is for, the construction exploits the application's task semantics rather than a quirk of one model's refusal surface. Three things follow. It will not transfer to a different buyer without redoing the reconnaissance, so a report claiming a general technique overstates it. It is more likely than a phrasing-tuned span to survive a model version change, because nothing in it depends on a specific refusal boundary — though adherence is probabilistic and that has to be re-measured, not assumed. And the cost profile is itself the finding: the strongest span reads exactly like ordinary correspondence, so a reviewer skimming the thread has nothing anomalous to see, and an attacker already inside a supplier relationship pays almost none of that cost.
go deeper
Know that an injected span can be expensive to write and that most of the cost may go into knowing the target, not into anything about the model.
Be able to split the drafting effort into model-facing probing and target-facing reconnaissance, and to say what each half predicts about where the construction still works.
Show the triage judgment: what the cost split says about transferability and durability, why a rate over runs beats a single success, and why the same property that made it costly also makes it unremarkable to a reader.
Own the claim in the writeup. Decide what the organisation is being told it has — a class demonstrated once, or a technique — and make sure severity is argued from what it costs the population who would use it.
## Why the cost breakdown is the interesting artefact When a red-teamer files this kind of finding, the question that separates a useful report from an anecdote is not 'did it work' but 'what did it cost, and who else can pay that'. Ten drafts is a number; where the ten went is the evidence. Broadly the effort in an injected span goes into one of two buckets. **Model-facing cost** is drafting spent probing what a particular model will and will not do: which phrasings are resisted, where a refusal boundary sits, what a screening layer scores highly. **Target-facing cost** is reconnaissance: the buyer's routing vocabulary, the shape of an open purchase-order thread, what the unattended run is actually for, what an ordinary supplier reply on that thread looks like. ## What a target-facing cost profile implies **Transferability is low, and the report must say so.** A span built from one buyer's vocabulary is not a technique that ports; it is an instance. The transferable thing is the class — that a run whose job is to read correspondence as fact can be handed a premise from which a different task follows — and the class is what belongs in the writeup. Presenting the instance as a general capability inflates the claim and will not survive the first attempt to reproduce it elsewhere. **Durability across model changes is comparatively high.** The construction does not depend on a phrasing that a refusal boundary happens to sit next to, so a model update that changes that boundary need not touch it. That is a reasonable prior, not a guarantee: adherence is a probabilistic property of a model reading a context, and a version change can move it for reasons nobody predicted. The honest statement is that it should be re-measured on a schedule, and that a re-measurement means a rate over many runs, not one confirming pass. **Reconnaissance cost is unevenly distributed, and that is the severity argument.** The cost that dominated the build is precisely the cost an attacker with an existing counterparty relationship has already paid — a compromised or merely cooperative supplier address on an open thread starts with the vocabulary, the thread history and a legitimate reason to be writing. Severity is not a function of what the exercise cost the exercise; it is a function of what it costs the population who would do it. ## The second-order consequence: nothing looks wrong The same property that made the span expensive to write makes it hard to notice. A span that reads as an ordinary supplier assertion, on a thread the workflow already corresponds with, is not anomalous text. There is no marker in prose separating a counterparty's assertion of fact from an attempt to redirect a task, because the run's entire job is to read such assertions as fact. A reviewer skimming the correspondence sees correspondence. And on the routine path nobody is skimming. The payoff here is not a leaked secret but the filed verdict itself — a routing or approval decision written into the system of record by an unattended run, which every later reader treats as already checked because a machine signed it. The provenance record proves which run filed it and from which message, not that anyone chose the verdict on the merits. ## What to write in the finding State the class, not the instance. State the obstacle the span got past — the standing rules restated at the top of the run, and the run's own stated task — because that is what makes the result meaningful. State the measured rate across runs and the number of runs behind it; one success against one deployment is a data point. State the cost split, because the person deciding what this is worth needs to know whether they are looking at a technique or an instance. And keep the span itself out of the report body: the mechanism is the transferable content, and a quotable string is the part with the shortest half-life.
- Does target-specificity lower the severity?Not by itself. Severity depends on what the construction costs the people who would actually build it, and the reconnaissance that dominated your build is already sunk for anyone with a counterparty relationship on that thread. What target-specificity does change is the claim: you report a class demonstrated against one deployment, not a portable technique.
- Would you expect it to survive a model version change?More readily than a phrasing tuned to a refusal boundary, because it turns on the application's task semantics rather than on what a model declines to say. But adherence is probabilistic and version changes move it unpredictably, so treat durability as a hypothesis to re-measure as a rate across runs, not as a property you can assert in the report.
- How many runs before you would call the result a finding?Enough to state a rate with the number of trials behind it. One success is a data point about one deployment and one context; a construction that lands two times in ten is still a finding, but it is a different finding from one that lands nine times in ten, and the report should say which one it is.
saying these in an interview costs you the question
- Reports an instance as a general, portable technique
- Treats one successful run as a reliable result
- Assumes low transferability means low severity
- Puts the span text in the report instead of the mechanism
- Assumes durability across model versions without re-measuring