How should an agent serialize a tool's return value before feeding it back to the model?
answer
- the model only sees the text
- self-describing beats a raw dump
- keep the ids the next call needs
- empty is not the same as failed
- every byte is re-sent each turn
basics
~20 sSerialize it into a short, self-describing string the model can read on its own: name the fields, keep units and status explicit, include the identifiers the next call will need, and leave out everything the model cannot act on.
solid answer
~50 sThe model never sees your return object — it sees whatever string your code puts in the tool result turn, so that string is the whole perception of what happened. Make it self-describing: label the values, state units and counts, and include any identifiers (record ids, cursors, file paths) the model will need for a follow-up call. Distinguish "ran fine, found nothing" from "failed" in words, because an empty string reads as a broken tool. Trim aggressively — a raw API dump costs tokens on this turn and on every later turn it stays in the window, and irrelevant fields actively distract the model. JSON and plain prose both work; what matters is that a reader with no access to your code could act on it. Prefer returning the answer over returning the data the answer was computed from.
go deeper
Know that the model only ever sees the text your code puts in the tool result, and that this text should name its fields, state units, and say clearly when nothing was found.
Explain why result size is paid for on every later turn, and show how you filter fields at write time and return handles instead of large payloads.
Demonstrate judgment about what belongs in a result at all: push aggregation into the tool, return the answer rather than the raw feed, and make truncation and empty results unmistakable.
Own the consistency of the whole tool surface — shared conventions for counts, units, ids, empties and truncation so the model generalizes across tools, and a token budget per tool result that survives many turns.
## What the model actually receives When an agent calls a tool, three different things happen in three different places. The model emits a request to call a named tool with some arguments. Your code runs the real work — an HTTP call, a database query, a shell command — and gets back a native object: a dict, a record, a response body. Then something in your harness turns that object into text and puts it into a tool result turn in the conversation. Only that last step is visible to the model. It cannot inspect your object, cannot see a stack trace you swallowed, cannot ask what a field means. The serialized string is the model's entire perception of what your tool did. That reframes what looks like a plumbing detail: choosing a tool's return format is choosing the sentence the model reads before it decides what to do next. ## Self-describing, not raw A good result stands alone. Numbers carry their units, values carry their names, and anything ambiguous is spelled out. `{"latency": 340}` forces the model to guess milliseconds versus seconds; `{"p99_latency_ms": 340}` does not. Enumerations should use words the model can reason about rather than internal codes: `"status": "expired"` beats `"status": 7`. The same applies to scale. If the tool returned 5 of 812 matches, say both numbers and say that the list is truncated. Silent truncation is one of the more damaging serialization bugs, because the model confidently reports a partial answer as complete. ## Keep the handles the next step needs Agents work in chains. A search result that omits record ids forces the model to search again to act; a paginated result that drops the cursor strands it after page one; a file-processing tool that returns content but not the path leaves nothing to reference later. This pairs with a useful pattern: return the handle, not the payload. Instead of inlining a 200 KB document, return its path, size and a one-line description, and let the model fetch or query it when — and only if — it needs the body. The context window holds a pointer; the data stays outside it until it earns its place. ## Empty is not the same as failed Three outcomes need three distinguishable strings: success with data, success with nothing, and failure. A tool that returns `""` or `[]` for "no rows matched" invites the model to conclude the tool is broken and retry it, or to invent a plausible answer. Write it plainly: `No incidents matched severity=P1 between 2026-04-01 and 2026-04-07 (0 of 0 rows).` That sentence tells the model the query ran, the filter that was applied, and that the empty set is real — which is often exactly the finding the user wanted. ## Every byte is paid for repeatedly A tool result is not consumed and discarded. It stays in the conversation and is re-sent with every subsequent turn until something removes it, so an oversized result is charged again and again in tokens and latency. Worse, long low-signal spans degrade the model's attention to the parts that matter. Serialization is the cheapest possible place to fix this: filter fields at the moment the result is created, before it ever becomes context. A blunt but effective test: for each field you are about to serialize, name the decision the model could make differently because of it. Fields with no answer — internal trace ids, HATEOAS link blocks, empty optionals, repeated boilerplate on every array element — get dropped. ## Format is less important than clarity JSON is convenient when the result is structured and the model must extract exact values; prose is fine, sometimes better, when the result is narrative. Markdown tables read well for small tabular results. What consistently hurts is noise: raw HTML, base64 blobs, deeply nested envelopes wrapping a single value, ANSI escape codes from a shell tool. Strip those in the wrapper. Consistency across your tool suite also pays. If every tool reports counts the same way and names errors the same way, the model generalizes across tools it has seen fewer examples of. ## Shape the result to the decision The strongest framing is that a tool result is an answer, not a data feed. If the agent's job is to find which service regressed, a tool that returns the three services whose error rate rose, with the numbers, beats one that returns every service's metrics and asks the model to do arithmetic in its head. Push computation into code, where it is exact and cheap, and let the model spend its turn on judgment instead of parsing.
- When would you return a reference instead of the content itself?Whenever the payload is large, only conditionally needed, or better queried than read — documents, large query results, generated files. Return the path or id plus size and a one-line description, and expose a tool that fetches or filters it. The context window then holds a pointer, and the bytes enter it only if the agent actually needs them.
- How should a tool report that it returned only part of the data?Explicitly and in the result text: state how many items were returned out of how many matched, and give the cursor or filter needed to get more. Silent truncation is the dangerous case, because the model reports a partial answer as a complete one and has no signal that anything was cut.
- Does the result have to be JSON?No. The model reads it as text either way. JSON helps when the model must extract exact values or when downstream code parses the same shape; prose or a small markdown table can be clearer for narrative results. Consistency across your tools matters more than the format you pick.
saying these in an interview costs you the question
- Assumes the model can inspect the tool's return object directly
- Dumps the entire API response because the model will figure it out
- Returns an empty string for no results, which reads as a broken tool
- Strips the ids and cursors the next tool call needs
- Treats result size as free because the tool call already succeeded