skip to content

In a long-running LLM agent, what must an obeyed span reach to change the standing objective rather than one step?

level: seniorimportance: should knowfreq 40%

answer

  1. one step versus every step
  2. who actually writes the objective record
  3. a consolidation step is the writer
  4. must read as task detail, not an order
  5. restated each step leaves nothing to edit

basics

~20 s

It must reach the record the loop re-reads, through a step that writes into it — usually one consolidating the standing goal. The agent performs that write, so the span must read as task detail, not an instruction.

solid answer

~50 s

Two conditions have to hold. First, the standing objective must be re-read from a mutable record rather than restated verbatim from a fixed source each step; otherwise nothing an obeyed span does can outlive its turn. Second, some step in the loop must feed that record — a consolidation, refinement or running-summary step is the usual candidate, and in a data-analysis assistant holding a multi-hour reconciliation goal that step exists precisely because the task is expected to sharpen as the data is understood. The attacker does not perform the write; the agent does. That constrains the span heavily: it has to survive the summarising step's selection of what to keep, which means reading as scope or task detail belonging to the work product, not as an imperative addressed to a model. The per-step verifier then dereferences the new revision and signs off.

code

text · 11 lines
text
step 13  reconcile_rows(source=ledger_a, target=table_b)
           read row 88.notes -> "[span elided - reads as scope detail about the data]"

step 14  consolidate_objective
           in   : objective_v3 + step 13 results
           out  : objective_v4  "...unchanged text... ; [added clause elided]"
           writer: consolidation step, session 4471

step 15  verify_action(action=..., against=objective_v4) -> pass
step 16  verify_action(action=..., against=objective_v4) -> pass
...

go deeper

for a junior

Know the basic requirement: the goal has to live in a record the run itself can rewrite, and something in the loop has to do the rewriting. Without both, an obeyed span affects one step only.

for a middle

Explain which step performs the write and why that matters — the revision arrives in the agent's register, with the run's own provenance, in the shape of an ordinary refinement.

for a senior

Show the reasoning you would apply to a real deployment: establish whether the objective is re-read or restated, find which step feeds the record, and say honestly what the construction costs and where it does not apply.

for a principal

Recognise that whether the standing objective is a run-writable artefact is an architecture decision somebody made, usually for context-length reasons, and be able to name what that decision bought and what it exposed.

## The question behind the question "Where would that span have to sit?" is the practical form of this leaf. An obeyed span in a retrieved row is common and usually cheap; a substituted standing objective is rare and expensive. What separates them is a **write path** from the step that consumed the span into the state the loop consults. ## Condition one: the objective must be re-read, not restated Some agents receive their goal fresh on every step from a source the run cannot modify — the operator's message replayed verbatim, a task definition loaded from outside the loop. Others hold a standing objective in a record the run itself maintains, because the task runs for hours, the context window cannot carry the whole history, and the goal is expected to be refined as work proceeds. A data-analysis assistant asked to keep a reconciliation between a ledger and a warehouse table current is the second kind: nobody restates the objective every few minutes, and the objective legitimately gains detail as the shape of the data is discovered. Only the second kind has anything to edit. This is why the same construction is worthless against one deployment and decisive against another, and why an interviewer asks about the architecture before the payload. ## Condition two: a step that writes into the record The attacker does not get to write the objective. Something inside the run does — the step that consolidates progress, rewrites the goal with what has been learned, or produces the running summary that the next step reads as its brief. The construction rides that step. The span is consumed as ordinary input by a step whose job is to produce a revised objective, and the revision it produces carries the change. This is the property worth articulating in an interview, because three consequences fall out of it: - **The wording is not the attacker's.** The revision is written in the agent's own register, from the agent's own summarisation. Searching for the injected string in the record finds nothing resembling it. - **The provenance is the run's.** The record correctly shows that the consolidation step wrote the revision during a legitimate session. That is true, and it is not evidence of anything, because a provenance field attests to which step wrote a value, never to who meant it. - **The revision looks like refinement.** Objectives on this kind of task are supposed to gain scope detail over hours. A revision that adds detail is indistinguishable in shape from the dozens of legitimate ones around it. ## What the span has to look like, described as a property Not as wording — this is a class of construction, not a string. The property is that the span must read as **content belonging to the work product**: task scope, a qualification about which rows or which period matter, a note about how something should be treated in the reconciliation. A summarising step keeps task-shaped content and drops conversational noise, so an imperative addressed to a model is the shape most likely to be discarded, and the shape most likely to be noticed if anyone reads the revision. The channel reinforces this. A free-text column of the very data being reconciled arrives as part of the material under analysis, not as a message, so content that reads like a note about that data is exactly what the pipeline is built to carry forward. ## What it costs More than a single-step abuse, and the honest answer says so. The attacker needs write access to a field that the reconciliation will actually pull in — the row has to fall in scope. They cannot choose the wording that lands or the moment it lands, since a consolidation step fires on the agent's schedule, not theirs. Several passes may be needed before anything persists, and each pass is a separate opportunity for the run to simply not keep it. There is no confirmation signal: the attacker generally cannot see the record. ## Where it stops It stops where either condition fails. If the standing goal is restated on every step from a source the run cannot write, the span steers one step. If the record is append-only and later steps read the first revision rather than the latest, the write happens and changes nothing. If the consolidation step's input is limited to the run's own action results rather than raw retrieved content, the span never reaches the writer. ## Answering well Name the two conditions, then say who performs the write and what follows from that — agent register, run provenance, refinement shape. Finish with cost and with the architectures the construction simply does not apply to. Candidates who jump to what the span says have answered a different, worse question.

  • Why must the span read as a refinement rather than as an instruction?
    Because the writer is the agent's own consolidation step, which keeps task-shaped content and discards conversational material. An imperative addressed to a model is the shape most likely to be dropped, and the shape most likely to stand out to anyone who reads the resulting revision. Content that looks like scope detail about the data under analysis is what the pipeline is built to carry forward.
  • What does it cost the attacker that the agent, not they, performs the write?
    Loss of control and loss of feedback. They choose neither the wording that lands nor the moment, since the consolidation fires on the run's schedule. Several attempts may be needed, each one free to be discarded. And they usually cannot observe the record, so there is no confirmation that anything stuck. The compensation is that the revision then carries the agent's register and the run's own provenance.
  • If the standing goal is restated verbatim from a fixed source on every step, what changes?
    The construction does not apply. An obeyed span can still steer the step that consumed it, but there is no operand for it to change, so nothing carries into the next step. This is why the first thing to establish about a deployment is whether the objective is re-read from a run-writable record or handed in fresh.
  • The record shows the consolidation step wrote the revision in a legitimate session. What does that establish?
    Only which step performed the write and when. A provenance field attests to authorship by a component, never to intent by a person. A revision written by the run's own summariser is exactly what both a legitimate refinement and a substituted objective look like, which is why the field cannot separate them.

saying these in an interview costs you the question

  • Assumes the attacker writes the objective record directly
  • Thinks any obeyed span automatically persists into later steps
  • Ignores that a consolidation step selects what is kept
  • Reads a provenance field as evidence of intent
  • Answers with what the span says rather than where it must reach

context