An agent's call log shows an approved operation with valid arguments — how do you triage a report that it should not have run?
answer
- ask what the record can actually prove
- which call ran, not who chose it
- completeness is not provenance
- diff the argument against the internal record
basics
~20 sStop asking whether the call was allowed and ask who chose each value. The log settles which operation ran and that it validated; it carries no provenance, so compare each argument against the record meant to determine it.
solid answer
~50 sThe report and the log are answering different questions, which is why the triage stalls. The log establishes that an approved operation ran, in character, with values that passed every check — and it is complete in the sense its designers meant. What it does not carry is where any single value came from, so 'we log every call, we would have seen it' is not the reassurance it sounds like. The move is to reconstruct provenance by hand: for each argument, identify what was supposed to determine it, and compare. If the beneficiary on a disbursement should have been fixed by the purchase order and instead matches a field on the submitted invoice, you have the answer, and it is visible nowhere in the record itself. Then be honest about verdict and reproduction: one success against one deployment is one success, and a construction that lands sometimes is still a real class.
code
json · 15 lines{
"event": "tool_call",
"tool": "disbursement.create",
"actor": "invoice-agent",
"status": "ok",
"validation": { "schema": "pass", "format": "pass", "amount_limit": "pass" },
"arguments": {
"invoice_id": "INV-4471",
"po_ref": "PO-2210",
"amount_minor": 918400,
"currency": "EUR",
"beneficiary_account": "<well-formed account identifier>"
},
"...": "no provenance field exists on this record"
}go deeper
Learn the habit of asking what a piece of evidence proves before arguing about it. A call record proves which operation ran and that its values were well formed, and nothing beyond that.
Be able to reconstruct provenance by walking the pipeline: for each argument, name the source that was supposed to determine it and compare with what actually did.
Show that you can run the adjudication: state whose expectation was violated, separate the faithful implementation from the false trust assumption, and be precise about what one reproduction supports.
Own the framing question of which assurance the organisation thought its logging bought, and be prepared to say plainly that it covers operations rather than the origin of the values they act on.
## Triaging a call that was allowed, valid, and wrong You have inherited a report. It says an autonomous procurement agent 'made a call it should not have made'. The attached evidence is a call record: an approved operation, on a real invoice, an amount inside policy, every argument passing validation. Nothing in it is anomalous. Somebody upstream has already suggested closing it as working-as-designed. ### First, name what the artefacts actually prove This is the discipline that unblocks the triage. Take each artefact and state its claim precisely: - **The call record proves which operation ran**, when, under which identity, and that its arguments satisfied the declared rules. That is all it was built to prove. - **A validation pass proves the values were of admissible shape.** Not that they were correct, and not that anyone with authority selected them. - **The absence of an alert proves nothing was of a kind anybody watches.** Kind, rate and sequence were all normal here, because the operation was the one the workflow performs daily. None of those three touches the question in the report, which is about a *value*, not an operation. ### The wrong answer to confront early 'We log every tool call, so we would have seen it.' It is a reasonable-sounding claim and it is the one that closes these tickets prematurely. The log is complete along the axis of *which call*; the dispute is along the axis of *where each value came from*, and that axis is not recorded at all. Completeness on one dimension is not coverage of another. Nothing in an ordinary record diffs argument provenance, and almost nobody is doing that comparison by hand either — which is exactly why the class survives in systems with excellent logging. ### The reconstruction that settles it Provenance is recoverable even when it is not recorded, because the pipeline is deterministic enough to walk. For each argument on the disputed call, ask: what was supposed to determine this value, and what actually did? 1. Take the arguments one at a time and identify the intended source — internal state such as the matched purchase order, or the ingested document. 2. Fetch the internal record and compare. If the beneficiary was meant to be fixed by the purchase order and instead corresponds to a field on the submitted invoice, the comparison is the finding. 3. Check whether that mapping is stable across other runs. If every invoice contributes that field, the exposure is structural rather than an incident. At the end of this you can state the claim in one sentence with evidence behind it: this argument was determined by content an outside party authored, and the workflow's owners believed it was determined by an internal record. ### Against whose expectation was the call wrong? This is the adjudication, and it is where triagers disagree honestly. Two positions are available: - **A design limit.** The workflow was specified to take that field from the document. The system did what it was written to do, and the report is a disagreement with the specification. - **A vulnerability.** The specification was written without anyone realising that the field's author is an outside party, and every downstream consumer — reviewers, approvers, the finance team reading the record — reads the value as though the organisation chose it. The honest verdict is usually that both are true, and that the second is the reportable half. A specification that grants an outside party the choice of what a payment lands on is a trust decision; if nobody knew they were making it, the fact that it was implemented faithfully does not settle anything. Write the finding against the trust assumption rather than against the code, and say which stakeholders were relying on the assumption. ### Reproduction, and what a single success proves Expect this to reproduce unevenly. Extraction and matching may involve a model, submissions vary, and the same document can land differently on different runs. A construction that lands one time in five is not a flaky test to be dismissed; it is a class that works whenever the mapping fires, on a queue that processes every submission. But be precise about the claim: a successful reproduction demonstrates that this construction worked once, against this deployment, in this configuration. It is not a rate, and generalising from it to 'this always works' or 'this is unexploitable' are both unsupported. ### What to hand back A good triage output states: the operation was approved and in character; the arguments were valid; one argument's value originated in content authored outside the organisation; here is the internal record it should have matched and did not; here is how often the mapping fires; and here is the question the existing artefacts cannot answer, which is who chose the value. That last sentence is the substance of the report, and it is worth writing down explicitly, because the natural instinct of everyone reading a clean log is to conclude there was nothing to find.
- The finding reproduces once in five attempts. Does that change your verdict?It changes the wording, not the verdict. One reproduction proves the construction worked once against this deployment in this configuration; it is evidence, not a rate. A path that fires intermittently on a queue that processes every submission still fires. What I would not do is report a measured success rate off five attempts, or let intermittency be read as unexploitable.
- The owner says the log is complete, so nothing was missed. How do you answer?By separating the axes. The log is complete about which operation ran and that its values validated; the dispute is about where one value came from, and no field on the record addresses that. Completeness along one dimension is not coverage of another. I would show the argument beside the internal record that should have determined it and let the mismatch make the point.
- How do you decide whether to file this as a design limit or as a vulnerability?Ask who was relying on the opposite assumption. If the specification knowingly sources that argument from the document and everyone downstream understands the value is an outside party's choice, it is a design position. If reviewers, approvers and finance all read it as an internally determined value, then the trust assumption is false and that is the reportable half — even though the implementation is faithful.
saying these in an interview costs you the question
- Closes the ticket because the log shows nothing anomalous
- Says a complete call log would have surfaced the problem
- Treats validation pass as evidence the value was intended
- Dismisses an intermittent reproduction as not a real finding
- Reports a success rate from a handful of attempts
- Files it against the code rather than the trust assumption