skip to content

An invoice workflow restates its rules each run — why doesn't a later position give an injected span leverage?

level: middleimportance: should knowfreq 46%

answer

  1. leverage is semantic, not positional
  2. recency is weak and length-dependent
  3. position decides survival, not adherence
  4. ask what it claims, not where it sits

basics

~20 s

Position is not what decides adherence. Recency is weak and length-dependent; leverage comes from what a span claims — a creditable source, a reason the task changed, a finishable job — not from where it sits.

solid answer

~40 s

Sitting later buys almost nothing on its own. Recency effects are real but weak, vary with context length and model, and are swamped by the fact that the run's restated rules arrive in the turn a model treats as most authoritative. The leverage in this construction is semantic. A span outweighs the earlier orders by supplying a source the model can credit, a reason the situation changed, and an alternative task coherent enough to complete — so no contest between instruction sets is ever posed. Position does matter, but for a different obstacle: whether the span survives a chunker or a truncation budget and reaches the context at all. That is a question about being read, not about being followed, and an attacker who conflates the two spends their drafts on the wrong axis.

go deeper

for a junior

Know that where an injected span sits is not why a model follows it. The claim the span makes is what matters; position is about whether the text arrives at all.

for a middle

Be ready to describe recency as a weak, length-dependent tendency and then name the three things a span supplies instead — a creditable source, a reason the task changed, and a completable job.

for a senior

Demonstrate measurement discipline: ablate components, count rates rather than outcomes, and resist concluding from a single successful run that you have identified the mechanism.

for a principal

Be able to say what a report may claim. 'The span was followed despite restated rules' supports a statement about coherence, not about recency, and not about anything upstream being absent.

## What the run looks like An unattended invoice-and-purchase-order workflow opens each execution by restating the buyer's standing routing and approval rules, then states the run's task. Somewhere below that sits the supplier email it must read. An attacker controls the email body and signature block, and therefore controls text that is, structurally, later than the rules. The intuition that later text wins is one of the most common wrong answers in this area, and it survives because it is occasionally coincidentally true. ## What recency actually is Attention over a long context is not uniform, and material near the end of a context is often weighted differently from material in the middle. That is a real effect, but it is weak, it varies with context length and with the model, and it is a tendency, not a rule. It is nowhere near strong enough to overturn an operator turn the model was trained to prefer. Treating it as the mechanism predicts that padding a span toward the end of the context would raise adherence, which is not what measurement shows. ## The three things that do carry **A source worth crediting.** The span arrives on a thread the workflow already corresponds with. The model has no notion of who wrote a passage or whether they were entitled to say it, but it does read the surrounding correspondence, and a passage consistent with that correspondence is one it has no reason to discount. **A reason the task changed.** An assertion about the world, from which a different task follows — not a demand that the standing rules be abandoned. The rules stay true, and the situation they govern is claimed to be different. This is the component that keeps the span out of the head-on comparison it would lose. **A job coherent on its own.** The alternative reading must run to completion and produce something that satisfies the run's stated task, or the model falls back to the original reading. A half-specified alternative is the most common reason a span that reads well still does nothing. A span with none of these does not get better by being moved later; a span with all three rarely needs to be. ## Isolating which component did the work The measurement that settles this is ablation: build the variant that removes one component and leave the rest intact, then run each variant enough times to see a rate rather than an outcome. With a probabilistic model a single pass tells you nothing — the same span can be acted on and then ignored. The variant whose rate collapses names the load-bearing component, and in practice it is usually the reason, not the position and not the emphasis. ## Where position does matter Position governs a different obstacle entirely: survival. Whether a span reaches the model at all depends on the chunker, on truncation budgets, and on which part of a source a pipeline actually forwards. A span that never reaches the context cannot be followed regardless of what it claims. That calculus deserves its own analysis and is not this question — but the distinction is exactly the one an interviewer is probing. Reaching the model and outweighing what the model was told are two separate problems with two separate answers. ## The direction of the evidence Restating the rules at the top of every run is a common thing to find in a deployment, and it is worth naming as the obstacle the span had to get past. What it is not is evidence about the mechanism: a span acted on despite restated rules proves that the model found an alternative reading more coherent, not that the restatement was ignored, and not that recency won. Confusing those two encodes the misconception the whole leaf exists to correct.

  • How would you find out which part of the span actually did the work?
    Ablate one component at a time — the creditable source, the stated reason, the self-sufficient task — and run each variant enough times to see a rate rather than a single outcome. Adherence is probabilistic, so one pass proves nothing either way. The variant whose rate collapses tells you what was load-bearing; if nothing collapses, you were probably not measuring what you thought.
  • Does position matter at all, then?
    Yes, for a different obstacle. Where a span sits inside a source decides whether a chunker or a truncation budget drops it before the model ever sees it. That is a question about reaching the context; this one is about outweighing what the model was already told. Answering the second with the first is the mistake.

saying these in an interview costs you the question

  • Claims the last instruction in the context always wins
  • Says the fix is putting the rules closer to the model
  • Cannot separate reaching the context from being followed
  • Judges a span from one successful run rather than a rate

context