skip to content

An assistant repeats a false metric definition and triage closed it as a hallucination. What reopens it?

level: seniorimportance: should knowfreq 41%

answer

  1. does it come back the same way
  2. clean sessions, new phrasings, a model change
  3. find the sentence in the chunk, then the row
  4. the log shows the row, never the intent
  5. faithful to its source is the tell

basics

~20 s

Determinism plus provenance reopens it. The same wording returns across fresh sessions, phrasings and model changes, and that wording is present verbatim in a retrieved chunk and in the stored row behind it - a decoding accident reproduces neither.

solid answer

~50 s

Argue it in two moves. First, reproducibility: run the question again in clean sessions, in several phrasings, under different users, and after any model change - a decoding accident drifts, while a retrieved artefact returns the same claim in the same words. Second, provenance: show the sentence in the retrieval record for that query, then in the underlying catalog row, with its edit timestamp; then show that answers changed when the row changed. At that point the answer is demonstrably faithful to its source, which is the opposite of a hallucination and relocates the defect upstream of the model. Be equally explicit about what the record does not show: a log proves which row supplied the text and when it changed, never who meant what by it. A careless edit and a deliberate plant leave identical traces, so write it up as authored with intent unknown.

code

json · 13 lines
json
{
  "query": "how is weekly active accounts defined",
  "results": [
    { "chunk_id": "metrics.weekly_active_accounts#description",
      "score": 0.86,
      "source_row": "catalog.metrics/weekly_active_accounts",
      "field": "description",
      "updated_at": "2026-07-14T09:12Z",
      "updated_by": "<account id elided>",
      "text": "[the false definition, quoted verbatim in the finding]" },
    { "chunk_id": "metrics.weekly_active_accounts#owner_notes", "score": 0.61, "...": "..." }
  ]
}

go deeper

for a junior

Know the first question to ask: does the same wrong claim come back in a fresh session with the same wording? Repetition is what separates a stored artefact from a one-off.

for a middle

Be able to trace the chain from the answer to the retrieved chunk to the stored row, and explain why an answer faithful to a false source is the opposite of a hallucination.

for a senior

Show the discipline of the write-up: a reproduction matrix with an honest rate, a provenance chain someone else can re-walk, and a clear statement of what the record does not establish.

for a principal

Own the classification consequence - a finding closed as model quality leaves the artefact in place, so the triage decision itself is what makes this construction durable.

## The chair you are sitting in Somebody has already decided. The finding is filed against the model, labelled a hallucination, and the owner is not hostile - they are simply working from the strongest prior in the industry, which is that a confident wrong answer came from the model. Your job is to change that classification in writing, using evidence that survives someone who disagrees with you. This matters beyond one ticket. A misdiagnosis here is not neutral: it closes the ticket and leaves the artefact in the corpus, so the same answer returns and the durability of the plant is purchased with the triage decision itself. ## Move one: determinism A decoding accident is produced at generation time and is sensitive to everything that touches generation. A retrieved artefact is not. So the first exhibit is a reproduction matrix, and the point of each axis is to remove an explanation: - **Fresh sessions**, so nobody can say prior turns steered it. - **Several phrasings** of the same question, so it is not one unlucky prompt. - **Different users**, so it is not one account's history or personalisation. - **Before and after a model change**, if one is available - the strongest single exhibit, because a claim that survives a different model was not stored in any model. A hallucination usually varies across those axes; the exact wording moves even when the gist does not. A plant returns the same sentence, because the same sentence is being handed to the model each time. One honest caveat, and it is the one that separates a careful finder from an enthusiastic one: reproduction is probabilistic at the edges. If the passage only sometimes ranks high enough to be retrieved, the claim only sometimes appears, and a partial reproduction rate is a fact to report rather than to hide. Report the rate. ## Move two: provenance Determinism says the claim is not being invented. Provenance says where it comes from. The chain is short and each link is checkable by somebody else: 1. The retrieval record for the query, showing the candidate chunks that were returned. 2. The offending sentence present in one of those chunks, verbatim. 3. The stored row that chunk was built from - in a catalog assistant, typically a description field - containing the same sentence. 4. The row's modification timestamp, and the change in answers dated after it. That chain converts a claim about the model into a claim about a stored artefact, and it is the reason the write-up should lead with the record rather than with a transcript. ## What the record proves, and what it does not Be scrupulous here, because overclaiming is how a reopened finding gets closed again. - A retrieval record proves **which chunk was returned** for that query, not that it was returned every time, and not that it was ranked first for any reason other than distance. - A row's edit metadata proves **which account wrote the current text and when**. It does not prove intent. A domain expert who genuinely misunderstands a metric and a person planting a false definition produce byte-identical rows. - The answer being faithful to the source proves **the model grounded correctly**. It says nothing about whether anyone reviewed the source. So the defensible conclusion is: the claim is authored, not generated; it entered through a write path that renders into answering context; intent is unestablished. That is a stronger finding than an accusation, and it is much harder to argue with. ## Reframing the class before you hand it over The last paragraph of the write-up is the one that matters most. One wrong sentence is a content defect. The class is that a field editable by anyone, which nobody re-reads, is rendered verbatim into answers that people quote as operational fact - and that a failure of that class is by construction indistinguishable, at the point of triage, from a model error. State the class, name the write path, and note that the previous classification would have closed it silently. That is what moves the finding from a ticket queue that closes to an owner who has to decide something.

  • It only reproduces in three runs out of five. Is it still a finding?
    Yes, and the rate is part of the finding rather than a reason to withhold it. Partial reproduction usually means the passage ranks high enough only for some phrasings, which is a fact about retrieval, not evidence that the sentence was imagined. Report the rate, the phrasings that trigger it, and the provenance chain - the stored artefact is the finding, and it is fully deterministic even when its retrieval is not.
  • Triage says a model change would fix it. What do you answer?
    That it tests the hypothesis rather than fixing anything. If the claim survives the change, the sentence is stored outside the model and the model was never the defect. If it stops appearing, that is far more likely to be a shift in how strongly retrieved content is weighted than a repair, and the artefact is still sitting in the row waiting for the next phrasing to reach it.
  • The row's editor is a domain expert who says it was a mistake. Does the finding survive?
    The security finding survives; the attribution never existed to begin with. The evidence only ever supported authored rather than generated, and a mistaken edit demonstrates the class as neatly as a deliberate one: a write nobody reviewed became an answer people quote. Drop any implication of intent and keep the class.

saying these in an interview costs you the question

  • Reports a single reproduction as proof of reliability
  • Reads an edit log as evidence of intent
  • Accepts the hallucination label without reproducing anything
  • Argues from a transcript with no retrieval record
  • Claims a model change would remove a stored artefact
  • Overclaims a breach when ordinary write access explains it

context