What should a ReAct agent do when a tool returns an empty or null observation?
answer
- silence is a result, not nothing
- blank output invites invention
- empty, null and error differ
- name the literal in the prompt
- echo the arguments back
basics
~20 sRender the empty case as explicit text, never as blank output. A literal such as NO_RESULT, plus the arguments used, tells the model the tool ran and matched nothing, which is a fact to reason about rather than a gap to fill from memory.
solid answer
~50 sAn empty observation is the most dangerous shape in the loop, because a blank line reads to the model as 'nothing happened' and the strongest continuation of a thought that expected data is to supply plausible data. A geocoding tool that returns null for a valid-looking address, rendered as an empty observation, reliably produces an agent that writes coordinates out of its priors. The fix has three parts. First, distinguish three cases in the observation text: **empty** (the call succeeded, zero matches), **null field** (a record exists but the field is missing), and **error** (the call failed) — each implies a different next action: reformulate, report unknown, or retry. Second, name those cases in the system prompt with the exact literals the runtime emits and say what the agent must do with each, including that inventing a value is forbidden. Third, echo the arguments back in the observation so the next thought can see whether the query itself was wrong before it retries.
code
python · 8 linesdef render_observation(tool: str, args: dict, result, error: str | None = None) -> str:
if error is not None:
return f"Observation ({tool} {args}): TOOL_ERROR - {error}"
if result is None:
return f"Observation ({tool} {args}): NO_RESULT - ran successfully, matched nothing."
if isinstance(result, (list, dict, str)) and len(result) == 0:
return f"Observation ({tool} {args}): NO_RESULT - ran successfully, zero records."
return f"Observation ({tool} {args}): {result}"go deeper
Remember that a tool returning nothing must still produce visible text in the transcript, and that the agent must not make up a value it never saw.
Explain why a blank observation causes invention rather than a stall, and distinguish an empty result, a missing field and a failed call by the different next action each implies.
Show how you would surface this in production: grounding checks over replayed traces, injected empty results in the harness, and per-tool empty-rate monitoring.
Own the policy question of when absence is a valid deliverable, how uncertainty is carried into downstream systems, and what an agent is contractually never allowed to assert without a source.
## Why emptiness is the hard case A ReAct transcript is a sequence the model continues. When a thought says 'I will look up the coordinates for 400 Maple Avenue' and the observation that follows is blank or `null`, the model faces a continuation problem: the most fluent next token sequence is a thought that proceeds as if the lookup succeeded. Language models are not built to notice absence; they are built to continue text. Absence rendered as absence therefore invites invention, and the invention is fluent, specific, and wrong — a latitude and longitude that look exactly like a real answer. This is different from an error. An error at least produces text the model reacts to. Emptiness produces nothing, and nothing is the worst possible signal. ## Make the empty case a first-class observation The runtime that formats observations should never emit an empty string. Instead it emits a named case: - `NO_RESULT — the geocoding tool ran and matched zero records for "400 Maple Avenue, Springfield".` - `NULL_FIELD — a record was found but the `coordinates` field is not populated.` - `TOOL_ERROR — the call failed: upstream timed out after 5s.` Three properties matter here. The case is a **literal token** the prompt can refer to. The observation **echoes the arguments** actually used, so the next thought can tell an empty result apart from a badly formed query. And the observation states **what the tool did**, not just what came back, so 'ran successfully and found nothing' is not confusable with 'never executed'. ## The three cases imply different next actions They are frequently collapsed, and collapsing them is a real bug. **Empty** means the query was well-formed and the world does not contain a match. The right move is usually to reformulate once — broaden the filter, drop a constraint, try an alternate spelling — and if the second attempt is also empty, to report the absence as the finding. Absence is often the answer: 'no open incidents match that service' is a correct, useful result. **Null field** means the record exists but the attribute is missing. Reformulating the same query will not help; the agent should either try a different source for that attribute or carry 'unknown' forward explicitly into the final answer. **Error** means the call did not complete. This is the only case where a retry of the identical call makes sense, and only for transient classes, with a bound on attempts. A prompt that says merely 'if the tool returns nothing, try again' produces an agent that retries a deterministic empty query until its step budget is gone. ## Say it in the prompt, not just in the runtime Formatting the case correctly is necessary but not sufficient; the model has to be told what the literals mean and what it may not do. A workable instruction set: the exact literals and their meanings; a rule that a value never present in an observation must never appear in a thought or an answer; a rule that after a bounded number of empty results the agent reports the absence instead of continuing; and one worked exemplar showing a thought reasoning *from* an empty observation to a decision. The exemplar carries more weight than the rule text, because it demonstrates the continuation you want in the format the model is continuing. ## Detecting the failure in practice The symptom is quiet: outputs are complete, well-formed and wrong on the subset of tasks where a lookup came back empty. Ways to surface it: - **Grounding checks** on replayed traces — assert that every concrete value in the final answer appears verbatim in some observation. Values with no observational source are fabrications. - **Empty-result injection** in the eval harness — force a tool to return the empty case and assert the agent reports absence rather than a value. - **Rate monitoring** — track how often each tool returns empty in production. A tool that is empty 30% of the time is both a hallucination risk and a signal that the agent is calling it with bad arguments. ## Related judgement calls Should the runtime hide the empty result and retry transparently? Usually no: silent retries remove the model's ability to reason about why the result was empty, and can hide a systematic argument-formatting problem. Should an empty result terminate the task? Only when the prompt has named absence as a legitimate final answer — otherwise the agent will treat 'no data' as failure and thrash. The best behaviour to demonstrate in an interview is an agent that can end with 'I could not find this, here is what I tried', which is far more useful than a confident fabrication.
- How would you prove an agent is inventing values after empty observations?Replay traces with a grounding assertion: every concrete value in the final answer — number, coordinate, id, date — must appear verbatim in at least one observation. Anything unsourced is fabricated. Then add an injection test that forces a tool to return the empty case on tasks that need it, and assert the agent reports absence instead of producing a value.
- Should the runtime retry an empty result before showing it to the model?Generally no. An empty result is a fact about the world, and hiding it removes the model's ability to reason about why the query matched nothing — often a malformed argument it could fix. Transparent retries also burn latency invisibly. Retry silently only for transient errors, where the model has nothing useful to contribute.
- When is an empty observation the correct final answer?Whenever absence is informative: no open incidents for a service, no matching transaction in the window, no prior record for a patient. The prompt has to authorise it explicitly, otherwise the agent treats absence as failure and keeps searching. The strongest form reports what was searched and with what arguments, so a human can judge whether the search was adequate.
saying these in an interview costs you the question
- Treating an empty result as a tool error and retrying it
- Emitting a blank observation and letting the model fill the gap
- Collapsing empty, missing-field and failure into one case
- Assuming the agent will notice absence on its own
- Never allowing 'not found' as a final answer