A goal-substitution finding reproduces once in five runs and the step trace is green — how do you adjudicate it?
answer
- two claims, two kinds of evidence
- the outcome is a rate, the structure is not
- a green trace is what success looks like
- instance patch versus the property
- route it to whoever owns the state model
basics
~20 sAdjudicate the structural property, not the outcome. That the objective is run-writable and the step check reads it reproduces every time, while the end-to-end result is one in five. Report the rate with its trials, and file it against the architecture owner.
solid answer
~50 sSplit the finding into two claims with different evidence. The outcome claim — this run pursued a substituted goal — is probabilistic, so report it as successes over trials and say plainly that five runs support almost no confidence interval. The structural claim is not probabilistic: the objective is re-read from a record some step writes, and the per-step verifier dereferences that record, so a green trace is exactly what a successful substitution produces. That reproduces every time and is the finding worth filing. On ownership, this is a defect neither in the injected span nor in the verifier, so filing it against either buys a patch the next construction walks around. Whether the standing objective stays run-writable belongs to whoever owns the agent's architecture, and a lead should say so rather than accept "we verify every step" as closure.
go deeper
Know that a probabilistic system can produce a real finding that does not reproduce every time, and that a green trace is not proof nothing happened.
Be able to separate what varies run to run from what is a fixed fact about the deployment, and to state which parts of a report rest on each.
Show how you would write the report: the rate with its trials, the inspectable structural facts, and why a green trace cannot discriminate between a clean and a substituted run.
Own the routing and the standard. Decide whether this is filed as an instance or a property, name the cost asymmetry between the cheap and the relevant change, and set the programme norm for reporting low-rate findings before the argument happens.
## Why this is a judgment call and not a technical one A one-in-five reproduction against a probabilistic system arrives at an organisation as a dispute. The person who filed it believes something real. The owner sees a green step-by-step trace, an objective revision written by the agent's own consolidation step in a legitimate session, and a demonstration that failed four times out of five, and reasonably asks what they are supposed to do. Somebody has to decide what the finding is, what it is worth, and who owns it — and could defensibly refuse it. That decision is the answer. ## Separate the two claims The single most useful move is refusing to let one report carry two kinds of evidence. **The outcome claim** is that a particular run ended up pursuing a substituted objective. This is inherently probabilistic: sampling, retrieval ordering, whether the row fell in scope, and whether the consolidation step kept the content all vary. Report it the way an empirical result is reported — successes over trials, with the conditions held fixed — and state plainly that five trials support essentially no confidence interval. Inflating a 1-in-5 into "the agent can be redirected" costs credibility the next time; deflating it into "unreproducible, closing" throws away the part that is certain. **The structural claim** is not probabilistic. Two facts about the deployment either hold or do not: the standing objective is re-read from a record that a step in the run writes, and the per-step verifier dereferences that same record. Both are inspectable, both reproduce every time, and together they entail that a green trace is precisely what a successful substitution looks like. That is the durable finding. It says the assurance the trace was supposed to provide does not cover this class at all — not that it is weak against it. ## What the green trace is worth as evidence Nothing either way, and saying so is half the adjudication. A fully green step-verification trace attests that every action matched the objective record at the time it was proposed. It does not attest to the record's authorship or its correspondence to the request. A run with a substituted goal produces exactly the same trace as a clean one, which means the trace cannot be cited as evidence against the finding. The same holds for the objective record's own writer field: it truthfully names the consolidation step, and that is equally true of every legitimate refinement. ## Bug or design limit Both framings are available and they lead to different places. Filing it as a bug points at the instance — this row, this span, this revision — and invites a fix at the instance: purge the row, add a pattern to whatever screens the retrieved text. That closes the ticket and leaves the property, so the next construction that reaches the same writer succeeds identically. Filing it as a design limit says the deployment has a run-writable premise that its own verification consumes as ground truth, and that this is a consequence of a real architectural choice — usually made for good reasons, because a multi-hour objective cannot be carried in context and is expected to gain detail as work proceeds. The instance then becomes a demonstration attached to the property, not the subject of the report. That is the shape a lead should insist on. ## Ownership and cost asymmetry Be explicit about the asymmetry, because it is what makes routing contentious. Making the standing objective something the run cannot rewrite touches how the agent is built: how long-running work is carried, how the goal is refined, what the loop reads at the top of each step. Tightening whatever screens the incoming text is comparatively cheap. Cheap and relevant are different questions, and the cheap change does not address a property that lives in the state model. Route the finding to whoever owns that state model, name the expensive option honestly, and let them decide with the cost visible — including the legitimate decision to accept the exposure for this deployment. A finding routed to the cheap owner because it will be accepted there is a finding you have closed rather than resolved. ## What a programme should conclude Two things generalise. First, on a probabilistic system, a red team's most defensible product is often a property with a demonstration rather than a reproducible exploit — build the reporting norms around that in advance, or every low-rate finding turns into the same argument. Second, treat controls that dereference run-mutable state as unable to attest to that state, and audit which assurance claims currently rest on them. That audit is a lead's job because it changes what the organisation is allowed to say it knows. ## The trap to avoid Do not accept, and do not offer, "we verify every step against the plan" as the resolution. It is the wrong answer restated as a remedy, and the whole point of the finding is that this particular control confirms the attack rather than catching it.
- How would you report a one-in-five outcome without over- or under-claiming?As successes over trials with the conditions fixed and stated, alongside the structural claim that reproduces every time. Say explicitly that five trials support almost no confidence interval, and that the value of the report is the property, demonstrated once, rather than a reliable exploit. That framing survives scrutiny; a headline rate from five runs does not.
- Why is filing this as a bug against the specific injected content the wrong shape?Because it names the instance rather than the property. A fix aimed at that row or that phrasing removes the demonstration and leaves the write path and the dereferencing check exactly as they were, so the next construction to reach the same writer succeeds identically. The ticket closes and nothing measurable changed.
- The owner points at a fully green step-verification trace as counter-evidence. What do you say?That the trace cannot discriminate. It attests that each action matched the objective record when proposed, which is equally true of a clean run and a substituted one, so it is consistent with the finding rather than against it. Ask instead whether the record is writable by the run and whether the verifier reads it — both answerable by inspection.
- When is accepting this exposure a defensible decision?When the owner understands the property rather than the anecdote, has priced the architectural change, and can state what the deployment's blast radius actually is. Acceptance made with the cost visible is a legitimate outcome of adjudication. What is not defensible is acceptance obtained by treating a low reproduction rate as absence of the property.
saying these in an interview costs you the question
- Closes a low-rate finding as unreproducible
- Reports one in five as if it were a reliable exploit
- Cites a green step trace as counter-evidence
- Files against the injected content rather than the property
- Routes the finding to the cheapest owner to get it accepted