How do you keep a ReAct thought grounded in what the observation actually said?
answer
- what did it actually say?
- quote before you conclude
- cite observation ids, don't paraphrase
- the dropped line is the deciding one
- keep the exception header, not just frames
basics
~20 sRequire the thought to quote a short span or cite the observation id before it draws a conclusion, and make sure your reduction preserved the line the conclusion depends on. Ungrounded thoughts drift into plausible paraphrase within two or three steps.
solid answer
~60 sGrounding is a joint property of the prompt and the reduction policy. On the prompt side, ask the thought to point at its evidence — a short verbatim quote or `per observation 3` — before it asserts anything, because a paraphrase is where drift enters and each subsequent step paraphrases the paraphrase. On the reduction side, make sure the deciding detail survived. A stack trace truncated to its last N lines is the classic failure: the frames are kept and the exception header at the top is cut, so the next thought guesses `NullPointerException` from the shape of the frames and the agent debugs the wrong fault for five steps. Reduce format-aware — always keep the exception line, the column headers, the status code — and mark what was dropped. Then verify: on replayed traces, assert that every concrete value in the final answer appears verbatim in some observation. Values with no observational source are fabrications, and that check finds drift far more reliably than reading traces by hand.
code
python · 9 linesdef reduce_stack_trace(trace: str, max_frames: int = 8) -> str:
lines = trace.splitlines()
header, frames = lines[0], lines[1:]
kept = frames[:max_frames]
dropped = len(frames) - len(kept)
out = [header, *kept]
if dropped:
out.append(f"...[{dropped} frames omitted]")
return "\n".join(out)go deeper
Know that a thought should point at the observation text it relies on rather than restating it loosely from memory.
Explain how repeated paraphrasing drifts across steps, and why a reduction that drops an exception header makes the next thought guess rather than stall.
Demonstrate the pairing: format-aware reduction that protects deciding spans, citation discipline shown in exemplars, and automated value-grounding checks over replayed traces.
Own the evidentiary standard for the system — what a final answer must cite, how source provenance is retained, and how grounded-but-wrong outputs from stale tools are caught.
## What grounding means inside the loop A ReAct agent alternates model text and tool text. Grounding is the property that each thought's claims are traceable to observation content rather than to the model's priors or to an earlier paraphrase of its own. It is not a binary: an agent can be grounded at step 2 and thoroughly ungrounded at step 8 because each step restated the previous one slightly more confidently and slightly less accurately. Two distinct mechanisms break it, and they need distinct fixes. ## Mechanism 1: the deciding detail was reduced away Observations are shortened, and the shortening decides what the model can know. Take a Java stack trace of 400 lines cut down to fit a step budget by keeping the last 40 lines. The frames survive; the header line naming the exception type and message — the single most load-bearing line in the whole payload — is gone. The next thought does not stall. It reads the frames, recognises a familiar call path, and produces `Thought: this looks like a null dereference in the mapper`. That is a guess wearing the register of a conclusion, and the following five steps are spent investigating a fault that never occurred. The fix is format-aware reduction: for every payload type you handle, name the spans that must never be dropped. Exception type and message for a trace. Status code and error body for an HTTP response. Column headers for tabular output. Total-count and page metadata for a paginated result. Keep those, then spend the remaining budget positionally or by relevance. The general rule — reduce by relevance, not position — applies here with a stronger constraint: some content is structurally required regardless of relevance scoring. ## Mechanism 2: the thought paraphrases instead of citing Even with a perfect observation in context, a thought that restates content in its own words introduces error. `Observation 2: balance = 41.37 USD, as of 2026-08-11` becomes `Thought 3: the account has about forty dollars` becomes `Thought 6: the account is nearly empty` becomes a final answer that asserts a low balance without a figure or a date. Nothing hallucinated a fact outright; each step was a small, reasonable compression, and the composition is wrong. The countermeasure is a citation discipline in the prompt and, more importantly, in the exemplars: a thought that draws a conclusion from an observation quotes the deciding span or names the observation id. `Per observation 2, balance = 41.37 USD as of 2026-08-11, so the payment will fail.` This costs a handful of tokens and buys three things — the value stops drifting, the final answer is auditable claim by claim, and a human reviewing the trace can spot the exact step where reasoning left the evidence. Exemplars carry more weight than instructions here. The model is continuing a pattern; show it two thoughts that cite, and it cites. ## Verification, not vibes Reading traces by hand does not scale and misses precisely the fluent failures. Three automated checks work well: **Value grounding.** Extract concrete tokens from the final answer — numbers, dates, identifiers, names, error types — and assert each appears verbatim in at least one observation from that run. Unsourced values are the fabrications you care about. The check has false positives (legitimately derived arithmetic) so allow a small annotated exemption list rather than weakening it. **Citation coverage.** Measure the fraction of conclusion-bearing thoughts that reference an observation id or quote a span. A run whose coverage collapses after step 5 is drifting. **Reduction retention.** Replay traces against the full payloads and check whether the reduced observation still contained the span the reference answer depends on. This separates 'the model reasoned badly' from 'the model was never shown the answer', which are different bugs with different owners and are constantly confused in postmortems. ## What grounding does not fix Grounding constrains the model to what the observation said; it says nothing about whether the observation was true. A tool returning stale or wrong data produces a perfectly grounded and perfectly wrong answer. That is why observations carry provenance — source and arguments and timestamp — and why a final answer that cites its observations is more useful than one that merely happens to be right: the citation lets a reviewer challenge the source. ## In an interview The strong answer connects the two mechanisms. Weak candidates propose 'tell the model to be accurate'. Strong candidates say: guarantee the deciding span survives reduction, require citation in the exemplars, then measure value grounding on replayed traces and treat a drop as a regression.
- Which spans would you protect from reduction in an HTTP error response?The status code and the reason phrase, any structured error code or correlation id, and the first part of the error body where providers put the human-readable message. Response headers and long HTML error pages can go. The principle generalises: for each payload type, enumerate the spans a next-step decision turns on and exempt them from the budget before positional or relevance cutting runs.
- Doesn't forcing quotes make thoughts longer and more expensive?Slightly, and it is usually worth it. A cited span is a few tokens; a step spent re-fetching because a value drifted costs a full tool round trip, and a wrong final answer costs far more. The pragmatic rule is to require citation only on conclusion-bearing thoughts — the ones that assert a fact or choose the next action — not on every line of reasoning.
saying these in an interview costs you the question
- Assuming the model will notice a missing detail
- Truncating a stack trace from the top
- Letting each thought paraphrase the previous one
- Reviewing traces by hand instead of asserting grounding
- Believing a grounded answer is necessarily a correct one