skip to content

What makes a plainly false claim planted in a catalog description durable rather than a one-off?

level: middleimportance: should knowfreq 48%

answer

  1. data, not a decoding accident
  2. it re-enters every future answer
  3. written in the asker's own vocabulary
  4. must not contradict its neighbours
  5. no wording for an update to invalidate

basics

~20 s

Durability comes from the plant being a stored text artefact rather than a decoding accident. It is retrieved again for every future asking of the same question, so it reproduces exactly where a hallucination does not.

solid answer

~50 s

Three properties do the work. First, persistence: the sentence lives in a row, so it re-enters the answering context every time somebody asks the question it was written for, and it survives new sessions, re-embedding and model changes because it was never inside the model. Second, on-topic phrasing: it has to be worded in the vocabulary of the question people actually ask, or it is never the passage that comes back. Third, quiet plausibility: it must not contradict its neighbours loudly, because a visible contradiction is what makes somebody open the source and check. What it costs the attacker is write access to one field that is rendered into prompts, plus knowing which question is routine. What it does not cost is any wording that a screening layer or a model update can invalidate, because there is no directive to become obsolete.

go deeper

for a junior

Recall the core contrast: a hallucination is produced fresh each time and moves when you rephrase; a planted sentence is stored, so the same wrong answer comes back the same way.

for a middle

Explain the three properties - it sits in a field that is rendered into prompts, it is phrased like the routine question, and it stays plausible beside its neighbours - and why none of them involves hiding anything.

for a senior

Demonstrate the cost comparison an interviewer is listening for: directive-shaped work is perishable against retraining and model updates, while a false declarative sentence has no wording that can be invalidated.

for a principal

Be ready to reason about corpus lifetime: material that anyone can write and nobody re-reads accumulates, and a programme that only samples answers will never find the passage that produced them.

## The difference between a wrong answer and a durable one An assistant that answers from a data catalog gets a question, retrieves the passages that look closest to it, and writes an answer from them. Two very different things can make that answer wrong. One is a **decoding accident**: the model produces a plausible-sounding claim that appears nowhere in what it was given. Rephrase the question, open a fresh session, change sampling, and the claim usually moves or vanishes. It leaves no artefact behind, because there was never one to leave. The other is a **stored false claim**. Somebody wrote a sentence into a field - a metric's description, a column's meaning, an owner note - that the application renders into the prompt as prose. That sentence is data. It comes back for the same question tomorrow, next month, and after the corpus is re-embedded, because none of those operations is an opinion about whether the sentence is true. This is why the construction is interesting to an attacker and irritating to whoever inherits the corpus: the effect is not probabilistic. It is as reliable as retrieval is. ## The three properties that make it land **Persistence with re-entry.** The point is not that the text exists somewhere; it is that it exists in the specific field the application pulls into answering context. A false sentence in a document nobody indexes is inert. A false sentence in the description field of a metric people ask about weekly is in the answer path every week. **On-topic phrasing.** A retrieval stage returns passages by closeness to the query, so a plant that uses different vocabulary from the question it is meant to answer will simply lose to the passages that do. In practice this means the plant is written in the same words the asker would use - which is also, conveniently for the attacker, exactly what makes it look native to the corpus. **Quiet plausibility.** The binding constraint is not evading a control; it is not being noticed by a human. A claim that flatly contradicts the passage sitting beside it invites somebody to open both and reconcile them, and that is the moment the construction dies. So the strongest version is a small, defensible-sounding shift: a definition that excludes a case it should include, a caveat that was never true, a stated unit or window that is subtly wrong. It reads like the corpus because it was written to. ## What the attacker does not have to solve Set against the directive-shaped alternative, the accounting is stark. A construction built around an instruction has to be worded a particular way, and that wording is in a race with whatever inspects text. It can stop working when a screening layer is retrained, when a model update shifts how strongly the assistant weights retrieved content, or when the exact string is published and trained out. A false declarative sentence has no wording to lose. It is not evading an instruction screen - it never triggered one. It is not evading a comparison between the answer and the retrieved documents either, because the answer genuinely follows from the retrieved document; the falsehood is a property of the source, and no stage of the pipeline evaluates sources against the world. The construction is stable across the exact axes that make directive-shaped work perishable. ## Where it does stop working It stops at humans and at arithmetic. - Somebody who opens the source and knows the domain disagrees with the sentence, and the row's edit history becomes a question. - Somebody recomputes the number the sentence describes and gets a different one. A claim that can be checked against something other than text is a weak plant; a claim about a *definition* is a strong one, because definitions have no independent referent to check against. - The passage stops being retrieved - the question people ask changes, the field stops being rendered into prompts, or a fresher passage on the same subject outranks it. Notice what is absent from that list: any moment where a control identified an attack. The failure ends when a person disagrees with a fact, which is a completely different event from a detector firing, and it is the reason findings of this class are so often written up as content quality rather than as security work. ## The interview version of this Asked what makes a plant durable, a weak answer talks about hiding the text. A strong answer says the opposite: nothing is hidden, and that is the mechanism. Persistence in a retrievable field, phrasing that matches the routine question, and plausibility that survives a reader's glance - and the cost is one write to one row, paid once.

  • Why does re-embedding the corpus not remove it?
    Re-embedding recomputes vectors from the same stored text. It changes how passages are compared, not what they say, so a false sentence survives intact and remains retrievable for the question it was written for. The same is true of a model change: the plant never lived in the model, so nothing about a new one removes it.
  • Would a plant that contradicts the surrounding corpus be stronger or weaker?
    Weaker. Contradiction is the cheapest thing for a human to notice, and one reader reconciling two passages is enough to end it. The construction is optimised for a reader shrugging and moving on, not for defeating a scorer, so agreeing with its neighbours everywhere except the one claim that matters is the point.
  • Is a claim about a definition stronger than a claim about a number?
    Usually, yes. A number has an independent referent - somebody can recompute it and get a different answer, and the disagreement is immediate and objective. A definition is text about text; disagreeing with it means arguing about intent, which people are far more willing to concede to an official-looking catalog entry.

saying these in an interview costs you the question

  • Says durability comes from hiding or encoding the text
  • Believes re-embedding or a model change removes stored text
  • Confuses a decoding accident with a retrieved artefact
  • Thinks the plant must outrank everything to work
  • Assumes a groundedness comparison catches a false source

context