skip to content

An obeyed instruction reaches a coding assistant inside failed CI output. How does that change target choice?

level: seniorimportance: should knowfreq 45%

answer

  1. routine relative to the task in progress
  2. the instruction arrives already timed
  3. entailed by the visible work, not merely permitted
  4. a coherent chain reads as self-directed
  5. shared state a pipeline pulls from

basics

~20 s

It moves the bar. Whether a call looks routine depends on the task the assistant appears to be doing, and an instruction arriving in a tool's return value lands mid-run, when repository writes already fit the story.

solid answer

~60 s

Remarkability is not a fixed property of an operation; it is how the call reads against the work in progress. Untrusted text arriving in a tool's return value — a failing job's output the assistant reads back and reasons over — lands at a point where the assistant has an established, visible reason to be touching the repository. A branch, a file edit, a workflow change, a build rerun all belong to "fixing the build", so a reviewer skimming the end-of-run summary reads them as consequences of the assistant's own work rather than as anything to stop on. The same operation requested in the first turn of a session, before any task justified it, would stand out. So the attacker's ranking shifts with placement: the best target is the widest-reaching operation that is also plausible for the task the run is already committed to. A second effect compounds it — a return value is text produced by the assistant's own action, which makes the follow-on calls read as self-directed.

code

json · 10 lines
json
[
  {"step": 1, "operation": "repo.read_file", "target": "src/parser.ts"},
  {"step": 2, "operation": "ci.fetch_job_output", "target": "job/8841",
   "returned": "142 lines of runner output [directive span elided]"},
  {"step": 3, "operation": "repo.open_branch", "target": "fix/parser-8841"},
  {"step": 4, "operation": "repo.edit_file", "target": ".ci/pipeline.yml"},
  {"step": 5, "operation": "ci.rerun_job",   "target": "job/8841"}
]
// step 4 is one line among five and reads as part of the fix;
// its effect is consumed by every later run of the pipeline.

go deeper

for a junior

Know that text a tool hands back to an assistant is untrusted input just like a fetched page, even though the assistant asked for it.

for a middle

Explain that whether a call looks routine depends on the task the run is visibly on, and that an instruction arriving mid-run inherits that justification for the calls it triggers.

for a senior

Demonstrate that you have reconstructed runs from action summaries. Say what such a record proves — which operation ran with which arguments — and what it cannot show, which is why the model requested it.

for a principal

Be ready to argue what an after-the-fact review of agent runs can honestly be claimed to buy when the constructions worth worrying about are the ones that produce coherent, task-shaped call sequences.

## The setting A coding assistant holds write access on a repository host: it opens branches, edits files including workflow definitions, moves labels, reruns jobs and cuts tags. Roughly forty operations were granted in one onboarding decision. It is asked to look at a failing build. It calls the operation that fetches the job's output, and that output — a test runner's stderr, a linter's advice line — comes back into the context as text the assistant reads and reasons over. A span in that text is treated as an instruction. This channel is unusual among agent inputs: it did not arrive from a document someone handed the assistant, it arrived from the assistant's own action. Whoever wrote it had to get text into a build's output, which is its own problem, but once there it is delivered by the assistant's own tool call. ## Why placement changes the target ranking The previous idea in this area is that the productive target is the widest-reaching operation that still looks routine. "Looks routine" is doing a lot of work, and this scenario shows what it actually depends on. A reviewer looking at the run afterwards does not read each call against a general model of what assistants do. They read it against what this run was doing. The mental question is "does this call belong to the job it was on". That makes the set of unremarkable operations a function of the task, not of the roster: | when the instruction lands | calls that read as ordinary | | --- | --- | | first turn, before any task | reads; almost any write invites a second look | | mid-run, assistant is fixing a build | branch, file edit, workflow edit, job rerun, push | | mid-run, assistant is preparing a release | tag, release notes, label moves, artifact publication | An instruction that arrives in a build's output arrives with a story already attached. Everything in the middle row is now cheap, and some of it reaches far: an edit to the definition that CI itself executes is state the entire team's pipeline pulls from on every subsequent run, and it sits on the same summary line as any other file edit. ## The compounding effect of the channel There is a second, subtler shift. Text from a fetched page or an uploaded file is visibly external. A tool's return value is the result of something the assistant chose to do. A reviewer reconstructing the run sees a coherent chain — asked to fix a build, fetched the build's output, edited a file, reran the job — and the coherence is itself reassuring. The chain is real; the reason the middle step produced the fourth is not what it appears to be. This is why the attacker's target selection here favours operations that are *entailed* by the visible task rather than merely permitted by the roster. Entailment is what buys the silence. ## What the run record does and does not show Care with direction. The action summary proves which operations ran under the assistant's credential and with what arguments. It does not record what caused the model to request them. An operator's request and a span obeyed from a job's output produce identical entries. Nor does the coherence of the sequence prove intent — a plausible chain is exactly what this construction produces, and plausibility is the property that was selected for. ## What to say in an interview The answer worth giving is that target selection is contextual: the same operation is expensive at one point in a run and free at another, because the thing it has to get past is a human judgment about whether a call fits the work in progress. Then name the consequence: a construction that waits for the assistant to be doing something before it aims at anything is worth more than one that fires immediately, and inputs that arrive from the assistant's own tool calls arrive already timed. A candidate who adds that a workflow definition is a particularly attractive destination — shared, consumed automatically, and indistinguishable from an ordinary file edit on a summary line — is showing they have thought about reach and ordinariness together rather than one at a time.

  • Why is text arriving in a tool's return value different from text arriving in a fetched page?
    Both are untrusted, but the return value arrives because the assistant acted, so it lands mid-task with a justification already visible in the run, and the calls that follow read as consequences of the assistant's own work. A fetched page is visibly external and often arrives before any task has been established.
  • The end-of-run summary shows a coherent sequence of calls. What does that coherence prove?
    That the calls made sense as a chain, nothing more. Coherence is the property the construction selects for, so its presence is not evidence of intent. The record shows which operations ran with which argument values under the assistant's credential; it holds nothing about what caused the model to request them.
  • Does this mean an attacker prefers to wait rather than fire immediately?
    Usually, yes, when the input channel allows it. Waiting until the assistant has a visible task means more operations are entailed by that task, which enlarges the cheap set. It costs reliability, though: the run may end, the context may be truncated, or the task may never require the operation the attacker wanted.

saying these in an interview costs you the question

  • Treats how routine a call looks as a fixed property of the operation
  • Says a plausible sequence of calls shows the assistant acted on its own
  • Ignores that a tool's return value is untrusted text
  • Assumes an edit to a pipeline definition reads differently from any other file edit
  • Claims the action log records why a call was requested

context