Why do LLM tool calls arrive with invented ids or out-of-enum values, and how do you catch them?
answer
- the model completes, it does not look up
- values appearing nowhere in the conversation
- categories the schema never defined
- parsing is not the same as true
- reject, never quietly coerce
basics
~20 sTool arguments are generated as text, so a model fills in a plausible order id, category or amount rather than leaving a field blank. Validate every call at the tool boundary — types, enum membership, units, and whether the referenced record actually exists — before executing anything.
solid answer
~50 sA tool call is not a lookup; it is a continuation. The model emits a tool name and a JSON object token by token, so when the context does not contain the order id, the most probable completion is a well-formed id rather than a refusal. Four failures show up in production: **fabricated identifiers**, **enum drift** (a refund `reason` of `EVENT_POSTPONED` when the schema allows only `EVENT_CANCELLED`, `VENUE_CHANGE` and `GOODWILL`), **unit or scale errors** (120 in a field that means cents), and **missing required fields**. Treat the model as an untrusted client: validate the payload against the schema, existence-check ids against the system of record, and range-check amounts before any side effect runs. Never silently coerce an unknown enum to a default — that converts a loud, catchable error into a quietly wrong action. Track invalid-argument rate per tool so drift after a schema change is visible.
code
python · 24 linesALLOWED_REASONS = ("EVENT_CANCELLED", "VENUE_CHANGE", "GOODWILL")
def validate_issue_refund(args, orders):
problems = []
order = orders.get(args.get("order_id"))
if order is None:
problems.append(f"unknown order_id {args.get('order_id')!r}")
if args.get("reason") not in ALLOWED_REASONS:
problems.append(
f"reason {args.get('reason')!r} is not allowed; "
f"use one of {list(ALLOWED_REASONS)}"
)
amount = args.get("amount_cents")
if not isinstance(amount, int) or amount <= 0:
problems.append("amount_cents must be a positive whole number of cents")
elif order is not None and amount > order["paid_cents"]:
problems.append(
f"amount_cents {amount} exceeds the {order['paid_cents']} paid on this order"
)
return problemsgo deeper
Be ready to say that the model writes tool arguments as text, so it can fill in an id or a category that was never mentioned in the conversation, and that the tool must check the values before acting.
Explain the concrete failure modes — invented ids, out-of-enum values, wrong units, missing required fields — and say where each is caught: type and enum checks on the payload, existence and range checks against your own data.
Demonstrate that you treat the model as an untrusted client: validation before any side effect, no silent coercion of unknown values, ownership checks on destructive tools, and an invalid-argument rate tracked per tool so drift after a schema or model change is visible.
Own the argument-correctness budget across the stack: what is worth enforcing at sampling time, what belongs in the boundary gate, what is only detectable after the fact by reconciliation, and what a one-percent invalid-call rate actually costs in this particular workflow.
## Where the arguments come from When a model "calls a tool", nothing is looked up. The model produces a structured continuation — a tool name plus an arguments object — using the same next-token machinery it uses for prose. The schema you supplied biases that continuation, and constrained decoding can make it syntactically impossible to violate, but the *values* are still generated. The model has no channel through which it can discover that order 8812 does not exist, or that the customer paid 12000 cents and not 120. If a field is required and the context does not contain a value for it, the highest-probability output is a plausible value, not an admission of ignorance. This is the whole explanation for a class of bugs that teams keep re-diagnosing as "the model is bad at JSON". The JSON is usually fine. The values are the problem. ## The four argument-level failure modes **Fabricated identifiers.** The model supplies an order id, customer id or SKU that appears nowhere in the conversation, or transposes one from an earlier ticket still sitting in the context window. This is the most dangerous class because the value is syntactically perfect and often refers to a *real but wrong* record. **Enum drift.** The schema allows three refund reasons; the model emits a fourth that sounds right. Models carry a strong prior over what refund reasons exist in the world, and that prior competes with your enumeration. Drift gets worse over time: when you rename or remove an enum member, the old value keeps appearing, because it lives in the model's prior and in whatever examples circulate in your prompts. **Unit and scale errors.** A field named `amount_cents` receiving 120 for a $120 ticket. A duration in seconds filled with minutes. A percentage supplied as 0.15 where the tool wants 15. These pass every type check. **Missing or spurious fields.** A required argument omitted under a long context, or an optional filter invented because the model reasoned it "should" narrow the query. ## Where to catch each one The boundary between the model and your side effect is the only place with enough information to judge all four. Build one gate per tool that runs before execution and returns a list of problems rather than throwing on the first: - **Schema and enum checks** catch types and out-of-enum values cheaply. - **Existence checks** against your database catch fabricated ids. This is the check most teams skip, because the call parsed cleanly and looked authoritative. - **Cross-field and range checks** catch units and scale: a refund larger than what was paid, a date before the order was created, a quantity above the seats on the booking. - **Ownership checks** catch a real id that belongs to a different customer than the one in this session. Decode-time constraints and boundary validation are complements, not alternatives. Constraining the sampler can guarantee the `reason` field is one of three strings; it cannot know whether order 8812 exists. Everything semantic has to be checked against the world. ## The coercion trap The most common self-inflicted wound is the helpful fallback: an unknown enum value quietly mapped to the nearest allowed one, or to a default, so the call can proceed. It looks like robustness. What it actually does is turn a visible, fixable failure into a silent data-quality problem — a goodwill refund recorded as an event cancellation, distorting the finance report months later with no trace of what happened. Reject and report. The only safe coercion is one you would be happy to explain to whoever reads the resulting records. ## Prompting is a weak fix Adding "never invent order ids" to the system prompt reduces the rate and never eliminates it, because the pressure that produces the invention — a required field with no value available — is structural. Two structural mitigations work better. First, reduce what the model has to recall: make ids something it copies verbatim from a prior tool result rather than something it reconstructs, and force the grounding lookup that puts the real id in context. Second, make the schema hard to misread — a field named `amount_cents` invites fewer unit errors than one named `amount`. ## Measuring it Log every rejected call with the offending value, keyed by tool and argument. Invalid-argument rate per tool is the headline number; the distribution underneath it is the diagnostic. Enum drift almost always clusters on one or two invented values, and that cluster tells you the enum name or the tool description is misleading. A rate that jumps on a specific day usually maps to a schema change, a prompt edit, or a model version change — which is why the metric needs to be per tool and continuously visible, not sampled during an incident.
- The call parsed, matched the schema, and every enum was legal. What can still be wrong?Everything semantic. Schema conformance says the shape is right, not that the values are true: the order id may not exist, may belong to a different customer, or may be a real id the model transposed from earlier in the conversation. Units can be off by a factor of a hundred while remaining a valid integer. Those checks require your data, so they belong in the tool boundary, not in the decoder.
- You rename an enum member. What happens to tool-call quality, and for how long?Expect a spike in out-of-enum values using the old name, because it survives in the model's prior and in any few-shot examples or transcripts still in context. It does not decay on its own. The practical fix is to make the rejection message name the current allowed values so the model can correct within the same conversation, sweep old values out of prompts and examples, and watch the per-argument rejection metric until it flattens.
- Would you ever accept an argument the model supplied for a destructive tool without an existence check?No. An existence check is the cheapest possible guard and the one that catches the highest-severity class — a well-formed id pointing at the wrong real record. For destructive or financial tools it should also verify ownership: that the record belongs to the customer bound to this session. A call that fails either check should never reach the side effect.
Treat the model like a form submitted by an anonymous web client: you would never insert its contents without validating the fields and checking the ids exist, no matter how well-formed the submission looked.
saying these in an interview costs you the question
- Assumes a strong model does not invent argument values
- Maps an unknown enum value to a default so the call proceeds
- Trusts an id in the call without checking the record exists
- Treats JSON that parses as JSON that is correct
- Relies only on a prompt line saying not to make up ids