A display name an outsider chose appears verbatim in an assistant's summary - what does that prove?
answer
- three artefacts, not one
- what can a template actually do?
- copying versus rewriting
- echo is not obedience
basics
~20 sAn echoed string proves only that the value travelled from the record to the rendered output. The panel may have templated it around the answer, and even if the model read it, echo is not obedience.
solid answer
~50 sTwo separate claims get collapsed here, and the whole answer is keeping them apart. First, an exact copy of the string in the output does not establish that the model ever saw it - the application may print the record header itself, or a client-side template may stitch the name around a generated paragraph. What separates those is a transformation only a generative model performs: the value coming back paraphrased, translated, case-normalised, pluralised, or woven into a sentence about a different field. Second, even a confirmed model read only shows the span was in the context. It does not show that a directive placed there would change the task, which is a different observation - output content the surrounding data cannot account for. So an echo is evidence about the assembly path, not about instruction-following, and one observation of a probabilistic system is not yet a measurement.
go deeper
Know that seeing your own text come back is not the same as the model doing what your text said. Be able to say what else could have put that string on the screen.
Explain the assembly chain - stored row, assembled prompt, rendered output - and name the one transformation that only a generative model can perform. That transformation is your discriminator.
Demonstrate that you convert an observation into a claim with a stated confidence: what you saw, how many times, what else fits, and which one extra probe would settle it.
Own the standard your team applies to evidence: what upgrades an inference to a finding, and what a screenshot alone is allowed to say in a report that other people will act on.
## Three artefacts, not one When an outsider sets a value on a record and later sees it in an assistant panel, there are three distinct artefacts in play and only one of them is visible from outside: 1. **The row** - the stored field value, which the attacker wrote. 2. **The assembled prompt** - whatever the application actually put in front of the model: some fields, in some order, wrapped in some framing wording, within some budget. Nobody outside sees this. 3. **The rendered panel** - the text on screen, which may be model output, application template, or a mix of both. Every inference the outside observer makes runs from artefact 3 backwards, and the discipline is refusing to name a claim stronger than the evidence supports. ## What each observation licenses | What you observe | What it licenses | What it does not license | | --- | --- | --- | | The exact string in the panel | The value reached the rendered output | That the model read it, or that it was in the prompt at all | | The value paraphrased, translated or summarised | A generative model read the value | That a directive placed there would be followed | | A sentence about a different field visibly shaped by yours | Your text influenced generation | That the effect is reliable, or that it survives another version | | Nothing at all | Nothing | That the field is not ingested | The second row is the useful one, and it is cheap. A template can copy, truncate and re-case; it cannot summarise. So the single most efficient probe is not a clever one - it is a value whose *rewriting* is unmistakable when it comes back. ## The false positives worth naming - **The record header.** Panels routinely print the record's own name above the generated text. That is the application quoting its own database and says nothing about the prompt. - **Client-side rendering.** The name may be stitched into the page by the front end after the response arrives. - **A cached summary.** The panel may be showing a summary generated before your edit, so what you are reading is a stale artefact and your correlation is with the wrong event. - **Coincidence.** A generated paragraph varies between runs. One changed sentence after one edit is a hypothesis, not a result; the change has to track your edits across repeats. ## Echo versus obedience This is the distinction that separates a middle answer from a junior one. Presence in the context and effect on the task are different facts. A model quoting a hostile span back is behaving exactly as a summariser should - it is describing its input. Only output the surrounding data cannot account for - a task the panel was not asked to perform, content the record does not contain, a refusal where none was warranted - is evidence about instruction-following. A report that shows the echo and calls it injection has demonstrated the delivery path and stopped one step short of the claim it makes. ## Why the direction of the claim matters here Each of these mistakes points the arrow the wrong way, and each one is expensive later: - *The string appeared, so the model has it* - it may have come from the template. - *The model has it, so it obeys it* - presence is not precedence. - *It worked once, so it works* - a probabilistic system needs a rate, not an anecdote. - *It did not appear, so it is not ingested* - absence has many parents, including truncation, relevance, caching and a stale index. ## What a good answer sounds like Say explicitly what you would claim on the evidence in front of you and what you would not, then name the one cheap extra observation that upgrades the claim. Interviewers on this topic are largely testing whether you can resist the temptation to write *injectable* in a report on the strength of a screenshot.
- What single observation most cheaply separates a templated echo from a model that read the field?The value coming back changed - paraphrased, translated, shortened, or mentioned alongside another field's content. Copying is what a template does; rewriting is what a generative model does. It costs one probe and one look, and it does not require the value to be directive at all, so it is the first thing to establish before anything else.
- The summary changes when you edit a field, but your text never appears. What have you learned?That something you control influenced generation without being quoted, which is consistent with the field reaching the context. It is weaker than a paraphrase because generated text varies run to run, so a single correlation proves nothing. Repeat the edit both ways and see whether the change tracks; if it does not, you have a coincidence.
- Why does a cached panel break this inference?Because the text on screen may have been generated before your edit landed, so you are correlating your write with an older artefact. Any result carries a time coordinate: what the panel showed, and when relative to the write. Re-checking after a known regeneration is what separates a real echo from a stale one.
A photocopier can reproduce a page perfectly and has read nothing. Someone who hands you a shorter version in their own words has read it.
saying these in an interview costs you the question
- Treats an echoed string as proof the model obeyed it
- Cannot distinguish the application's template from generated text
- Calls a single run of a probabilistic system a result
- Assumes everything in the panel is model output
- Reports the delivery path as a confirmed injection