skip to content

An assistant attributed a fabricated statistic to a named institute: how do you tell an invented citation from a fetched page's own metadata?

level: seniorimportance: should knowfreq 40%

answer

  1. start from the turn, not from a retry
  2. fetched set, declared fields, rendered strip
  3. was the name ever in the retrieved text?
  4. a fetch log is not a copy of the page
  5. the artefact decides it, not the hit rate

basics

~20 s

Work from the turn's fetch record, not from a retry. If a fetched page declared that publisher and the chip reproduces it, the attribution was supplied by that page; if no fetched text carried the name, the generation produced it.

solid answer

~50 s

Triage this from the artefacts of the original turn, because a retry answers a different question. Pull what was fetched for that turn, what each page declared about itself, and what the interface actually rendered. If a fetched page declared that publisher and byline and the chip reproduces them, the attribution came from the page, and you are looking at a planted-provenance finding rather than a generation defect. If nothing was fetched, or nothing fetched carried the name, the model produced the attribution itself, which is a different failure with a different owner. Two traps. A fetch log proves a URL was requested and a status returned; it does not prove what that page contains now, and a page written for this purpose gets changed or pulled once noticed. And failing to reproduce proves the same page was not fetched again, not that the finding was unreal.

code

json · 19 lines
json
{
  "turn_id": "...",
  "fetched": [
    { "url": "https://<host elided>/2026/03/quarterly-note", "status": 200, "at": "2026-03-11T09:14:02Z" }
  ],
  "page_declared": {
    "title": "<title elided>",
    "byline": "<person name elided>",
    "publisher": "<institution name elided>",
    "date": "2026-03-04"
  },
  "chip_rendered": {
    "publisher": "<institution name elided>",
    "byline": "<person name elided>",
    "date": "2026-03-04"
  },
  "page_text_retained": false,
  "...": "..."
}

go deeper

for a junior

Know which two possibilities you are separating: the assistant produced the attribution itself, or a fetched page supplied it. Recall that the answer lives in what was fetched during that turn.

for a middle

Explain the inference across the fetched set, the page's declared fields and the rendered strip, and why a match between the last two confirms faithful copying rather than corroborating anything.

for a senior

Show you have triaged one. Lead with the original turn instead of a retry, state that a fetch record is not a copy of the page, and route the finding away from both generation quality and store compromise.

for a principal

Own the call on evidence retention as a tradeoff rather than a checklist item, and be able to say what your organisation can and cannot substantiate about a finding once the page it depended on has gone.

## Why the retry is the wrong first move The instinct on a report like this is to re-ask the question and see what happens. It is close to useless here. A research assistant's fetch set changes with ranking, freshness and availability, and a page built to be cited can be edited or withdrawn as soon as anyone notices it. Re-running samples today's world; the finding is about a turn that already happened. Start with the record of that turn. ## What the record can settle For the reported turn you want three things, and their relationship is the whole triage: 1. **What was fetched.** The URLs requested for that turn, with statuses and timestamps. 2. **What each page declared about itself.** The title, byline, publisher and date fields extracted at fetch time. 3. **What was rendered.** The strings the source chip actually displayed beside the answer. If (2) contains the institute's name and (3) reproduces it, the attribution was **supplied**. Someone wrote a page that declared itself published by that institute, the pipeline read the declaration, and the interface displayed it. Nothing malfunctioned, which is exactly why nothing raised an alarm. If (1) is empty for that turn, or (2) contains no such name anywhere, the attribution was **produced by the generation**. That is a different defect with a different owner and a different remedy, and conflating the two sends the whole investigation sideways. ## What the record does not settle Be precise about the direction of each claim, because this is where a triage goes wrong: - A fetch entry proves a URL was requested and a status returned. It does **not** prove what that URL contains now, and it does not prove what it contained then unless the extracted text or bytes were retained. - A rendered chip proves what was displayed. It does **not** prove any reader looked at it, clicked it, or believed it. - A match between the page's declared fields and the chip proves the pipeline copied faithfully. It does **not** corroborate anything, because both values have one source. - Absence of a retained page snapshot is the single thing that decides whether this finding can still be examined at all. If nothing kept the fetched text, and the page has since changed, the artefact is gone and you are left with a screenshot and an argument. ## The reproduction question On a probabilistic system, one success in five attempts is a familiar triage problem, and this class has an unusual answer. The uncertainty is not mainly in the model - it is in whether the same page is reached again. So the evidence that decides the finding is the **artefact**, not the hit rate: a retained copy of a page that declared itself published by an institution and carried a claim that institution never made settles the question, whether or not the assistant cites it a second time. Conversely, a clean re-run proves the page was not fetched again, which is compatible with it having been pulled an hour ago. ## Routing it Once you know the attribution was supplied, several natural conclusions are wrong and should be said out loud in the interview: - It is not a model-quality issue, so tuning generation will not touch it. - It is not evidence that any store or index was compromised; nothing was breached, a page was published. - It is not prompt injection, because no directive exists anywhere in the artefact. Anything watching for instruction-shaped text saw a normal page and was right to. What you actually have is a finding about a third party's name appearing on a claim they never made, in a surface readers treat as verification - and the person carrying the harm is not your user and not your operator, but the institution named in the chip. ## How to answer Describe the three artefacts and the inference between them, then name the two traps: a fetch record is not a copy of the page, and a failed re-run is not a disproof. Finish by routing the finding away from generation quality and away from store compromise, which is the part that shows you have actually triaged one of these.

  • The page 404s when you check it a day later. What does that change about the finding?
    It removes the evidence, not the event. A 404 today is equally consistent with the page having been withdrawn once it was noticed. If the fetched text was retained, the finding stands on that copy; if it was not, you have a rendered chip and a timestamp and no way to show what the page declared, which is a materially weaker position to take to anyone.
  • The team wants to log this as a hallucination and move on. What is your objection?
    The value was in the retrieved text, so the generation did what it was supposed to do. Filing it as a hallucination assigns it to model quality, where no amount of work will touch it, and it loses the actual fact: a third party's name was displayed in a position readers read as verification, on the strength of a string that party never wrote.
  • Does a clean re-run let you close the report?
    No. A clean re-run shows the same page was not fetched this time, which the page being changed, deprioritised or removed all explain. For this class the deciding evidence is the artefact rather than a success rate, so closing on a failed reproduction is closing on the wrong measurement.

saying these in an interview costs you the question

  • Re-runs the query first and closes on a clean result
  • Files it as a hallucination when the value was in the fetched text
  • Reads a fetch log as proof of what the page contains
  • Concludes the index or store must have been poisoned
  • Hunts for a hidden instruction that the artefact never contained

context