An error log line is byte-identical when written and when an LLM triage assistant reads it — what makes it directive only on re-read?
answer
- authority is positional, not intrinsic
- who is reading, and what can they do
- the logger name is the only attribution left
- the formatter had no reason to carry provenance
- same sentence, two different rooms
basics
~20 sThree things change around the unchanged text: the reader now interprets its whole context, the text arrives as the system's own operational evidence rather than a stranger's input, and that reader holds capabilities the write path lacked.
solid answer
~50 sAuthority here is positional, not intrinsic. The same characters sit in a validation log at write time, where the only consumer is a formatter and a disk, and in a model's context at read time, where every span in the window is text the model weighs when deciding what to do. Two things travel with the span on the second read only. First, framing: it is retrieved as a line from the application's own error log, the account a system gives of itself, so nothing marks it as attacker-authored — the field's origin was dropped by formatting code that had no reason to carry it. Second, reach: the reader can change alert routing and write a shift summary humans act on. Instruction-following is a trained preference weighed over the whole window, so a span that looks like operational evidence competes with the app's own framing rather than sitting inertly beneath it.
go deeper
Know that identical text can be inert in one place and directive in another, and that the difference is who reads it and what that reader can do.
Be ready to walk the stages and say, for each, what it knew about the characters it handled. The loss of origin at the formatting step is the detail to land.
Demonstrate that you can separate model behaviour from pipeline shape, and state which of the two generalises. Expect to be asked where the effect weakens as quotation gives way to extraction.
Own the framing that this is a property of how produced artefacts are consumed, not a defect in one component, and be able to say which pipelines have the shape and which do not.
## The text did not change; the setting around it did Candidates often reach for a property of the string — it was obfuscated, it was encoded, it looked like a system message. In the deferred case none of that is required. The bytes at write time and read time can be identical. What differs is everything around them, and it is worth naming the differences precisely because interviewers push on exactly this. ### 1. The reader changed from a formatter to an interpreter At write time the consumers of the string are a formatting call and a storage layer. Neither interprets; they concatenate and persist. At read time the consumer is a model, and a model does not have a channel that means "this is inert data". It receives one window of text and produces a continuation conditioned on all of it. Any separation between the application's instructions and retrieved content is a preference the model was trained to have, not a boundary the runtime enforces — which is why a span that reads like authoritative operational context competes for influence rather than sitting quietly below the app's own framing. ### 2. The framing changed from stranger's input to the system's own account This is the part that does the most work and the part most often missed. | Stage | What the stage knows about the characters | |---|---| | Submission | A field value from an unauthenticated request that failed validation | | Log write | A substring to interpolate into a message template | | Log store | A line of text with a timestamp and a logger name | | Re-read | Operational evidence the system recorded about itself | Provenance is not stripped by an adversary. It falls away at the log-write step because the code doing the formatting had no notion that a consumer might one day act on what it wrote. By the time the assistant retrieves the line, the only attribution attached is the logger's name — which is the application, not the person who supplied the substring. ### 3. The reach changed At submission the span had no reachable capability: the request was rejected. At re-read the assistant that consumes the log can adjust alert routing on the low-severity path and can emit a shift summary that the next on-call engineer reads and acts on. The identical sentence goes from having no consequence available to having two. ## Why this is not the same as "the model was fooled" A useful way to answer is to separate the two questions an interviewer is really asking. - *Why did the model follow it?* Because the model weighs its whole context and the span arrived in a position associated with reliable operational fact. Following it was not an error state; it was the normal behaviour of a component asked to act on evidence. - *Why was the span there to be followed?* Because a stage that could not act carried a stranger's characters into a store that a stage which can act reads as its own. The first question is about model behaviour. The second is about pipeline shape, and it is the one that generalises: the same mechanism runs through a summary regenerated for a different reader, an exported record re-ingested by a downstream job, or a metadata field rendered into a console another component scrapes. Each is a place where the artefact one stage produced becomes the input another stage trusts. ## Getting the direction of the claims right Several claims here are easy to state backwards, and doing so encodes the misconception: - A model acting on a retrieved line proves the line reached the context. It does not prove the store was tampered with, nor that the writer had privileges. - The line's logger name proves which component emitted the record. It does not prove who authored the characters inside it. - The span looking like operational evidence proves that it occupies a position the reader treats as evidence. It does not mean the model verified anything. ## Where the effect weakens The further the second reader is from raw quotation, the weaker this gets. A stage that summarises log volume into counts, that extracts fields into a structured record, or that renders text through something which paraphrases, may never place the original characters in front of the assistant at all. Conversely, the closer the pipeline is to "paste the recent error strings into the window", the more completely the write-time framing is replaced by the read-time one. That gradient — how much of the original text survives into the acting component's window — is the honest way to describe when this construction is strong and when it is not.
- If provenance had travelled with the characters, would the model still have followed them?Possibly, but the question changes. Carrying origin forward means the acting component is at least in a position to weigh it; without it, the component has nothing to weigh at all. The honest framing is that the span currently arrives with no marker distinguishing it from text the system authored, so nothing in the pipeline is even able to ask the question.
- Which part of this generalises beyond logs?Any artefact one stage produces that a later stage consumes as its own account: a regenerated summary read by a different audience, an exported record re-ingested by a downstream job, a metadata field rendered into a console that another component scrapes. The pattern is a produced artefact being read as evidence, with no step re-deriving who supplied the text inside it.
- How would you tell this apart from the assistant simply hallucinating the route change?Look for the span. A fabricated action leaves no matching text in the context that was assembled; a followed one leaves a retrieved line whose content maps onto the argument values chosen. If you cannot reconstruct the window the run was given, you cannot distinguish them, and that inability is itself the finding.
saying these in an interview costs you the question
- Says the text must be obfuscated or encoded to work on re-read
- Treats the instruction hierarchy as an enforced runtime boundary
- Claims the log store must have been tampered with
- Says the logger name attributes the characters to the application
- Assumes retrieved text is safer than typed text