A red-team run made only granted, ordinary-looking calls. Do you file that as a bug or a design limit?
answer
- nothing malfunctioned, so nothing to file against
- name the claim the run falsified
- the owner follows the claim, not the code
- accepted with an owner beats closed by-design
- the roster's shape outlives the specific run
basics
~10 sNeither label fits until you name the claim the run falsified. Nothing malfunctioned, so there is no component to file against; what broke is the belief that a roster granted once bounds a run.
solid answer
~60 sStart by admitting what did not happen: no component misbehaved. The operation was granted, the orchestrator executed it as designed, the action summary rendered it accurately, and the reviewer triaged it exactly as a fast triage does. Filing that as a bug against the assistant gets it closed as working-as-intended, and the finding never comes back. The write-up has to identify the assumption the run falsified instead — typically that a capability roster assembled in one onboarding decision bounds the effects any single run can have, and that a human skim of an action summary sorts by consequence when it actually sorts by how ordinary each call looks. Say what would have to become untrue for the behaviour to stop, name the owner of that decision — the person who granted the roster, not the team that wrote the assistant — and be honest that the likely outcome is an accepted limit with a named owner and a review date. An accepted risk with an owner is a real result; a bug closed as by-design is not.
go deeper
Know that a red-team result is not automatically a bug, and that a run in which every call was permitted still needs someone to decide what it means.
Be able to distinguish what a run record shows from what it implies, and to state a finding as a claim that failed rather than as a sequence of calls that occurred.
Show you can write the report so it survives triage: name the falsified assumption, separate the durable part from the specific run, and be honest about how many attempts it took.
Own the adjudication and the routing. Decide whether the limit is accepted, name the owner of the decision that would change it, price that change honestly, and refuse the default outcome where a ticket is acknowledged and nothing is owned.
## Why this adjudication is hard The usual triage question — is this a defect or intended behaviour — assumes a component that either met its specification or did not. Here every component met it. The operation was on the granted roster. The orchestrator ran what the model requested, which is its job. The action summary listed the call truthfully. The reviewer glanced at the line and moved on, which is what a skim of a list is. There is no misbehaving part to point at, and a report that points at one will be dismissed correctly. That dismissal is the real risk to the finding. A report filed as a bug against the assistant comes back marked working-as-intended, and the organisation has now spent its attention on the finding and learned nothing. ## Adjudicate by naming the falsified claim The productive move is to state, in one sentence, the belief that the run showed to be false. Common candidates: - That a roster granted during onboarding, and never re-read since, bounds what any individual run can do. - That the operations somebody put an approval in front of are the operations whose effects matter most. - That a human triage of a run's calls sorts by consequence. It sorts by ordinariness, and the two orderings disagree. - That an assistant's coherent, task-shaped sequence of calls is evidence the sequence was intended. Each of these is a claim somebody is relying on, usually implicitly. A finding that names one and shows a run in which it did not hold is a finding, whatever label it ends up with. A finding that shows a call happening is not. ## Who owns it The owner follows from the claim, not from the code. If the falsified claim is about the roster's shape, the decision belongs to whoever granted it — often a platform or IT function that made a single onboarding call, not the team that maintains the assistant. If it is about what an after-the-fact review of runs buys, it belongs to whoever staffed that review and told a stakeholder it was sufficient. Getting the owner wrong is the second-most-common way these reports die: the receiving team agrees the behaviour is real, agrees it is not theirs, and the ticket sits. This is why a red-team write-up in this area is partly an organisational document. It has to say: here is the claim, here is who is relying on it, here is who could make it true, and here is roughly what that would cost. Note that the cost is genuinely asymmetric — re-deriving forty grants against actual per-run need is expensive and touches everyone who depends on the assistant working, while the alternatives are cheap and do not address the claim. ## Accepting a limit is a legitimate outcome A lead should be comfortable landing on "design limit, accepted, owned by X, reviewed in six months". That is a decision, it is recorded, and it survives staff turnover. What a lead should not accept is the third option that gets chosen by default: the report is acknowledged, no claim is named, no owner is assigned, and the behaviour is regarded as understood because everyone read the ticket. That state is indistinguishable from having done nothing, and it will read as a surprise the next time. One more honesty requirement. A construction like this reproduces against a deployment, not against a model. It worked in this roster, with these tasks, at this point in a run. Say that plainly, including how many attempts it took, because a probabilistic system that obeyed once has not been shown to obey reliably, and a report that overclaims reliability invites a rebuttal that discredits the part that was solid. The durable part of the finding is the shape of the roster, which does not change when a prompt is edited or a model version rolls; the fragile part is the specific run. Distinguishing them is what makes the report survive its first review meeting. ## The one-line version Bug and design limit are not the axis. The axis is: which claim did this falsify, who is relying on it, and what would have to change for it to hold. Answer those three and the label follows, and it usually turns out not to matter.
- The team offers to add an approval step in front of the operation the run used. What do you say?That it answers the specific call and not the claim. The finding was that the ordering used to decide what gets watched does not track what has reach, so moving one operation across that line leaves the ordering unexamined. Accept it as a reasonable local response, but do not let it close the finding, and record which claim remains unaddressed.
- How do you report reliability when the construction landed on some runs and not others?State the attempt count and the conditions plainly, and separate the durable claim from the fragile demonstration. The roster's shape is stable and is what the finding is about; whether a given run obeyed is probabilistic. Overclaiming reliability invites a rebuttal that discredits the solid part along with the weak one.
- What makes this finding different from a normal excessive-agency write-up?It is not that the assistant could do too much. It is that the operation selected was the unremarkable one, which means the finding is as much about the review path and the ordering it uses as about the grants. A write-up that only lists over-broad permissions misses half of what the run demonstrated.
saying these in an interview costs you the question
- Files it as a bug against the assistant with no claim named
- Assigns it to the team that built the assistant rather than the grant owner
- Treats accepting a limit as a failure of the report
- Presents a single successful run as a reliable reproduction
- Lets one added approval step close the finding