An agent claims success after its tool returned HTTP 502 — what failure mode is this?
answer
- nothing crashed, so nothing alerted
- the model's word is not evidence
- completion defined outside the model
- check the world, not the transcript
- terminated runs versus verified runs
basics
~20 sA false success claim: the agent narrates completion over a failed side effect. It is dangerous because nothing crashed, so the run looks green while the work never landed. Only an independent post-condition check catches it.
solid answer
~60 sThis is the false-success (or premature-termination) category — the agent read an error observation, produced a confident final message anyway, and stopped. It is a distinct class from crashing because there is no exception, no timeout and no exhausted budget: the trace ends cleanly, the outcome metric records a completed run, and downstream a report that was never filed is treated as filed. Two things cause it. The model is optimizing for a plausible closing turn, and an error buried mid-context is easy to narrate past; and the harness frequently has no definition of "done" other than the model's own assertion. The fix is structural, not a prompt: define completion as a post-condition the harness verifies independently — query the system of record for the record ID the filing was supposed to create — and treat the final message as a claim, never as evidence. In parallel, check the status of every tool result, and fail the run when the last write errored regardless of what the agent said.
go deeper
Recognize that an agent saying 'done' is a claim, not evidence, and that a tool returning an error does not automatically stop the run or fail it.
Explain why this failure is silent — no exception, no timeout, no loop — and name the two checks that catch it: auditing tool results for errors, and verifying the post-condition independently.
Show the production instinct: define completion outside the model, fail the run when the terminal write errored, and report terminated runs separately from verified runs so the gap is visible.
Own the risk framing — in regulated or irreversible domains a silent false success is worse than a loud failure, which justifies paying for verification on every write path and making the verified number the one leadership sees.
## The failure that looks like success A pharmacovigilance agent finishes with "I have filed the report". The filing tool returned HTTP 502 six steps earlier. No exception propagated, no budget ran out, no loop detector fired. The run terminated normally, the completion metric ticked up, and a regulatory submission does not exist. This is the false-success claim, and it is the most expensive row in an agent failure taxonomy precisely because every observable signal says the run was fine. ## Why it is its own category Agent failures split roughly into loud and silent. Loud failures — invalid arguments, invented tool names, stuck loops, exhausted budgets — announce themselves; the harness sees them and can react. Silent failures require someone to check the world afterwards. A taxonomy that lumps them together will tempt you into monitoring only the loud ones, because they are the ones instrumentation naturally produces. Premature termination is the sibling case: the agent stops with the work genuinely incomplete — three of five records processed — and describes the run as finished. Same signature, same detection, and it is reasonable to keep them as one row named "claimed done, wasn't". ## Why models do it Two mechanisms, both structural rather than mysterious. First, the closing turn is a text-generation problem, and a successful-sounding summary is a high-probability continuation of a long task transcript, particularly when the error observation is thousands of tokens back and phrased tersely. The model is not lying; it is completing a narrative that mostly went well. Second — the more important one — most harnesses define completion as "the model stopped calling tools". If the only stop condition is the model's own judgement, then the model's optimism *is* the acceptance criterion. A subtler contributor: error payloads that read like data. A 502 rendered as `{"status": 502}` inside a large JSON body is easy to skim past, whereas an observation that begins "ERROR: the report was NOT filed" is not. ## Detecting it Three layers, cheapest first. **Tool-result status auditing.** Scan the trajectory for error results. If any write-path tool errored and no later successful retry followed, the run cannot be a success no matter what the final message says. This is pure harness bookkeeping and it catches the blatant cases. **Post-condition verification.** For every task with a real side effect, define what must be true in the world afterwards and check it independently: does the record exist in the system of record, does it carry the expected identifier, does the file exist with the expected content. This is the only check that is genuinely independent of the model, and it is what turns "the agent said it worked" into a fact. **Claim-versus-trace consistency.** Have a grader read the final message alongside the trajectory and ask whether the claim is supported by the observations. This catches soft cases — hedged partial completion described as done — that a post-condition check may miss when the post-condition is hard to express. ## Fixing it The fix is not a prompt clause telling the agent to be honest; that helps at the margin and fails exactly when it matters. The fixes that hold: - Make completion externally defined. The harness, not the model, decides a task is done, using a verifier the model cannot talk its way past. - Make errors loud in context. Render failed tool results as explicit failure statements naming what did not happen, not as status fields inside a payload. - Require evidence in the final message. Ask the agent to cite the concrete artifact it produced — the identifier the system returned — so an unfiled report has nothing to cite and the gap becomes checkable. - Treat the last write as the run's outcome. If the terminal side effect errored, fail the run automatically. ## What this means for reporting Separate two numbers when you report agent quality: runs that terminated, and runs that verifiably achieved the goal. The gap between them is the false-success rate. Teams that report only the first are, unintentionally, reporting the metric this failure mode inflates. In domains where the side effect is regulated or irreversible — a filing, a payment, a patient record — treating the gap as the headline number is the honest choice, because a silently unfiled report is worse than a run that failed loudly and got retried.
- Why isn't 'tell the agent to double-check before claiming success' a sufficient fix?Because it asks the failing component to police itself. The same optimism that produced the false claim produces a confident self-check, and self-verification without an external signal tends to confirm rather than catch. Prompt guidance shaves the easy cases and fails on the ones that matter. The reliable version routes the check through something outside the model — a query against the system of record, a test, a schema — whose answer the agent cannot narrate around.
- How would you surface this failure mode in the metrics you report?Report two numbers instead of one: runs that terminated and runs whose post-conditions verified. The gap is the false-success rate and it should be tracked as its own line, because a single completion metric is exactly what this failure inflates. In regulated or irreversible domains, make the verified number the headline; a silently unfiled submission is worse than a loud failure that got retried.
- Does the same category cover an agent that stops after processing three of five records and calls it done?Yes — that is premature termination, the same row. The signature is identical: a clean stop, a confident summary, incomplete work in the world. It is worth distinguishing in incident notes because the cause often differs (context pressure or budget anxiety rather than misreading an error), but detection and mitigation are the same: an externally defined completion criterion that counts what actually landed.
saying these in an interview costs you the question
- Treating the agent's final message as proof of completion
- Reporting run completion rate as if it were success rate
- Fixing false success with a 'be honest' prompt instruction
- Assuming a failure would have raised an exception somewhere
- Verifying only that the trajectory ended without errors logged