skip to content

In an LLM agent, how does editing the standing objective differ from getting one tool call obeyed?

level: juniorimportance: should knowfreq 58%

answer

  1. one turn versus every turn
  2. the agent re-reads something each step
  3. the injected text need not stay
  4. later actions are the agent's own
  5. corrupt the goal, not the action

basics

~20 s

One obeyed call ends when that step ends. An edited objective persists: the agent re-reads it at the top of every later step and plans from it, so it keeps pursuing the substituted goal on its own initiative with no further injected text.

solid answer

~40 s

Getting one call obeyed means untrusted text sitting in this turn's context was read as directive, and one action ran with argument values the caller never chose. When that text scrolls out of the window, the effect stops; a second action needs the text read again. Editing the standing objective attacks the state the loop consults, not the action. A long-running agent typically re-reads a recorded goal at the top of each step, so a single successful edit is re-consumed indefinitely, and every later action is the model's own generation planned from the substituted goal. That makes it durable, self-driving, and awkward to spot: each action is individually defensible against the goal now on file, and only the aggregate is wrong.

go deeper

for a junior

Be ready to state the difference in one sentence: one obeyed call is bounded by that step, an edited objective is re-read at every step afterwards. Know that the injected text does not have to persist for the effect to persist.

for a middle

Explain the mechanism: the loop consults a recorded goal before choosing each action, so a write into that record changes the premise of every later decision, and the later actions are the model's own generations.

for a senior

Show why this is hard to triage in real runs — each action is locally consistent with the goal on file, so the wrongness only appears in the aggregate, and searching later steps for the injected span turns up nothing.

for a principal

Own the framing consequence: any assurance story that treats the plan as the trusted fixed point collapses here, so decide deliberately whether the standing goal is a run-writable artefact and who owns that choice.

## The loop that makes the distinction matter A tool-using agent runs a loop. It reads state — a recorded objective or task record, some recent transcript, the last tool result — decides on an action, calls a tool, reads the return value, and repeats. Untrusted text can reach that loop from many channels: a retrieved chunk, a fetched page, an inbound message, a tool's return value, or a database row rendered into the prompt as part of the work product. Two quite different things can be attacked in that loop, and interviewers separate candidates on whether they keep them apart. ## Getting one call obeyed Here the untrusted span is present in the context for this step, the model reads it as directive rather than as data, and one action runs that the operator never chose — a call to a capability the agent already had, usually with argument values selected by the span rather than by the user. What is important is the scope. The effect is bounded by the step. When the span falls out of the context window, or the sub-task that fetched it ends, nothing carries. To get a second action the attacker needs the span read a second time. Note also what an obeyed call proves and does not prove: it proves the span reached the model's context and was treated as instruction. It does not prove any store was breached, any credential was stolen, or any boundary was crossed other than the instruction/data one that was never enforced in the first place. ## Getting the objective edited Here the target is not the action but the state the loop consults before choosing an action. In a long-running agent — one holding a standing objective across hours, such as keeping a reconciliation between two data sources current — the goal is generally not restated by a person on every step. It lives in a record the run itself re-reads, and often refines, as work proceeds. If an obeyed span causes that record to be rewritten, the attacker has moved from influencing one action to influencing the premise of all of them. Three properties follow, and they are the whole answer: - **Durability.** The injected span does not have to stay anywhere. Once the recorded goal carries the change, the record is what gets read; the original text can be gone from every context window in the run. - **Initiative.** The subsequent actions are not injected. They are the model's own plans, generated from the goal it now believes it has. The attacker supplies no wording for them and does not need to anticipate them. - **Local defensibility.** Every one of those actions is consistent with the objective on file. Read one call in isolation and there is nothing wrong with it. The wrongness lives in the aggregate — in what the run as a whole accomplished — and in the record, not in any step. ## Why interviewers ask this first Because the confident wrong answer is right next door. Many engineers describe agent injection entirely in terms of the abusive call: which capability was reached, what argument values were passed, whether an approval was clicked. That framing quietly assumes the plan is the fixed part and the action is the variable part. Once the recorded goal is attacker-influenced, the assumption inverts, and controls built on it stop being obstacles and start being confirmations. ## What the construction needs It is not free. It requires an objective that is re-read from a mutable record rather than restated verbatim from a fixed source each step, and a step in the loop whose output is written back into that record — a consolidation, refinement or running-summary step is the usual one. The attacker does not perform the write; the agent's own machinery does, which is precisely why the resulting revision looks like ordinary progress. If the standing goal is handed to the model fresh from an unchanged source on every step, an obeyed span can steer that one step and nothing more. ## Saying it in an interview One sentence for the distinction — one call is bounded by the turn, an edited objective is re-consumed at every turn — then one sentence for the consequence, that later actions become the agent's own initiative and each is locally correct. That is a complete answer at this level.

  • Does the injected span have to remain in the context for the substituted goal to keep working?
    No, and that is the point. Once the change is in the record the loop re-reads, the record is what drives later steps. The span can be gone from every subsequent context window, which is also why hunting for it in the transcript of a later step finds nothing.
  • If no single call would have been refused, what exactly is the harm?
    The harm is the aggregate. Each call is judged against the objective on file and passes, because the objective is what was attacked. The run as a whole accomplishes something no operator asked for — scope quietly changed, a data set treated differently, work spent in a direction nobody chose — while every individual line of the trace looks ordinary.
  • What does an agent obeying a span in a retrieved row actually prove?
    It proves the span reached the model's context and was read as an instruction rather than as data. It does not prove the data store was compromised, that any authentication failed, or that the agent had privileges it should not have had. The row may be legitimate content that anyone with normal write access can edit.

Forging one payment request gets one payment. Editing the standing instruction the bookkeeper works from gets every payment after it, written in the bookkeeper's own hand.

saying these in an interview costs you the question

  • Treats one obeyed tool call and an edited objective as the same attack
  • Assumes the injected text must stay in context to keep working
  • Says each step being defensible means nothing bad happened
  • Thinks the attacker must author every later action
  • Claims an obeyed span proves the data store was breached

context