skip to content

Why does an assistant's durable edit to a calendar or contact record rate higher than a wrong answer in one turn?

level: seniorimportance: nice to knowfreq 34%

answer

  1. the session ends, the entry does not
  2. the write carries the owner's name
  3. it comes back as first-party data
  4. one write, many later readers
  5. the record shows which change, not whose choice

basics

~20 s

Because the edit survives the session, carries the owner's authorship to everyone who later reads it, and re-enters the assistant's own context on later turns as the owner's data. A wrong answer stops when the conversation does.

solid answer

~40 s

Rate impact by what persists, not by what the model said. A durable write -- an added attendee, an altered event, a rewritten contact record -- has three properties a one-turn answer lacks. It outlives the session, so nothing about ending the conversation undoes it. It is attributed to the owner's account, so other people read it with the trust they extend to the owner, and it propagates to their calendars as well. And it is read back: on later turns the assistant retrieves its own store and finds that text again, now indistinguishable from the owner's own entry, so the effect can recur without anyone acting again. Reversal also needs someone to know what changed, and the audit record shows only which change ran under which account -- not who chose it.

code

json · 10 lines
json
{
  "event_id": "evt_8f21",
  "change": "attendee_added",
  "actor": "[email protected]",
  "via": "assistant/session_41c",
  "origin_field": "invite.description",
  "origin_span": "[directive span elided]",
  "recurrence": "weekly, no end date",
  "ts": "..."
}

go deeper

for a junior

Recall that an assistant acting on someone's account can change data other people later read, and that such a change is still there after the conversation ends.

for a middle

Explain the mechanics of read-back: the assistant retrieves the owner's own store on later turns, so text written there returns as first-party data rather than as external content.

for a senior

Demonstrate triage judgment -- rank by what persists, who reads it, and how it is attributed; and know that a change record proves which change ran under which account, not who chose the values.

for a principal

The call you own is what a single successful run licenses you to claim. Push severity onto durable, owner-attributed state changes rather than on how often one phrasing happened to work.

## Rate what persists, not what was said A finding that ends with "the assistant produced text it should not have" and a finding that ends with "the assistant wrote something into a store other people read" are not the same severity, even when the construction that produced them is identical. Four differences do the work. ### It outlives the session A conversation ends and its context is gone. A calendar entry, an attendee on a recurring series, or a rewritten contact record is still there tomorrow. The session boundary, which bounds most of what a chat surface can do wrong, does nothing here. If the record is a recurring series, one write reaches every future occurrence. ### It carries the owner's authorship The assistant acts on the owner's account, so the change is attributed to the owner. Anyone who later reads the entry -- a colleague seeing a new attendee, someone opening the contact record to find a phone number or an address -- applies the trust they extend to that owner. The text has been laundered of its origin not by any clever step but by the ordinary attribution of a write. This is also why the blast radius exceeds the target: an added attendee to a shared series puts the entry on other people's calendars too. ### It is read back into the same assistant The assistant reads the owner's own diary and address book to do its job. So the written text returns to the model's context on later turns, now arriving as first-party data rather than as an external invite. A single write can therefore keep producing effects with no further action from whoever composed it, and it will do so past any change to how new invites are handled. ### Reversal is not free Undoing requires knowing what changed and what it was before. The audit trail helps less than people expect. A change record establishes which change ran, at what time, under which account, and through which session; it does not establish who chose the value. Reading the `actor` field as the person who decided is the classic misreading -- the account is the owner's, and the assistant's session is the mechanism, not the author. If the originating invite has been deleted or the series edited since, the provenance link may be gone entirely. ## What this does to reproduction Durability cuts the other way when you try to demonstrate it. You cannot re-run a write against the same record cleanly: the first run mutated the target, so the second run starts from different state. Practically, that means capturing the record before and after, doing the demonstration against a scratch record you own rather than a live shared one, and being explicit that a probabilistic model may not repeat the same behaviour on the next attempt. A single successful run tells you the construction worked once against one deployment -- which is a real observation, and not the same claim as a reliable one. ## Writing it up The strongest version of the finding says three things: the state that changed and who can now read it; that the change is attributed to the owner and is therefore trusted by everyone downstream of them; and that the assistant will read the changed state back as its own. Notice that none of those depends on how the text was phrased. That matters, because it is what stops the discussion from collapsing into whether one particular sentence still works. ## The traps Two bad instincts show up in triage. The first is grading by refusal: "the model refused most of the time, so this is low" -- a refusal is the answer being declined on one attempt, and it says nothing about the write that already landed on another. The second is grading a calendar or contact change as cosmetic because no data left the building. Nothing was exfiltrated, but a durable, owner-attributed edit to shared state is exactly the kind of thing other people act on -- a meeting that now includes an attendee nobody vetted, or a contact record whose details are read as authoritative months later.

  • How does persistence change the way you reproduce and report the finding?
    You cannot cleanly re-run a write against the same record -- the first attempt mutated the state the second would start from. Capture the record before and after, demonstrate against a scratch record rather than a live shared series, and say plainly that one successful run shows the construction worked once against one deployment on one date, not that it is reliable.
  • Does the change record identify who is responsible?
    No. It shows the change ran under the owner's account through an assistant session. Provenance tells you which session performed the write, not who chose the argument values that went into it. The originating invite is the evidence for that, and only if it still exists and is still linked -- which is often the first thing lost.
  • Why is a rewritten contact record worth more than an altered one-off meeting?
    Because it is read repeatedly and rarely re-verified. A one-off event passes; a contact record answers "who is this" and "where do I send this" for months, both for the owner and for the assistant resolving names on later turns. Longer read life and lower scrutiny beat a single high-visibility change.

saying these in an interview costs you the question

  • Rates impact by whether the model refused
  • Calls a calendar or contact edit cosmetic
  • Assumes ending the session undoes the effect
  • Reads the audit actor as the person who chose the value
  • Claims reliability from one successful run

context