skip to content

The team patched the assistant's prompt to ignore invite instructions and closed the finding -- what do you tell them?

level: principalimportance: should knowfreq 40%

answer

  1. credit the measurement, bound the claim
  2. a likelihood moved; a property did not
  3. who may write it is unchanged
  4. the cheap fix lives inside one team
  5. the residual needs a name and a date

basics

~20 s

Say what the patch bought -- lower compliance with the phrasings that were tried, on one build, on one date -- and what it left untouched: who may write the field, what its validation was for, and that nobody owns re-deriving that.

solid answer

~50 s

Separate the measurement from the property. The patch measurably reduced how often the assistant complied with the constructions you actually tried, which is a real result worth recording with its scope attached. It changed nothing about the three things that made the field eligible: an outsider can still write it legitimately, its validation is still the shape check a calendar renderer needed, and no stage re-derives what the field is trusted for now that something downstream can act. Take as given that what the patch leans on is a trained preference, not an enforced rule, so a lower rate is not zero. Then make the organisational point: the unanswered question spans two owners, and closing the finding closes it with nobody's name on the residual. Accepting that is a legitimate call -- but it should be decided, not achieved by ticket hygiene.

go deeper

for a junior

Recall that a change to what an assistant is told is not the same as a change to who can write the text it reads, and that the difference is what the conversation is about.

for a middle

Be able to state what a prompt change measurably altered and what it did not, in terms of write access and what a field's validation was written for, without arguing about any particular sentence.

for a senior

Show that you would bound your claim to the constructions tested, the build and the date, and that you can explain why a non-reproduction is not evidence the class is gone.

for a principal

Own the ownership call and the cost asymmetry: the cheap remedy sits inside one team, the real decision spans two, and your job is to make sure the residual is accepted by a named person rather than closed by ticket hygiene.

## Two different claims, and only one of them was tested A prompt change and the eligibility of a carrier are claims about different things, and the conversation goes wrong when they are conflated. **What the patch bought.** After the change, the constructions the reviewer tried stopped producing the behaviour, on that deployment, on that date. That is a genuine observation and belongs in the record with its scope attached: which phrasings, which build, when. Take as given that the mechanism it leans on is a trained preference rather than an enforced rule -- so what moved is a likelihood, and a likelihood does not become a guarantee by being small. **What the patch did not touch.** Three properties made the field eligible, and all three are exactly as they were: | property | changed by the patch? | |---|---| | an outsider may write the field, legitimately, by design | no | | its validation was specified for a consumer that could not act | no | | no stage re-derives what the field is trusted for | no | That table is the whole argument, and it is worth being able to draw it from memory in a room with a feature owner. ## Why this particular remedy always wins the argument Not because anyone is careless. Because of cost asymmetry. Editing the assistant's prompt is a change one team makes in an afternoon, entirely inside its own repository, with no coordination. Anything that touches the eligibility of the carrier reaches across an ownership boundary -- to the team that runs the calendar write path, to whoever decides what the assistant is allowed to do on the owner's behalf, sometimes to a product decision about accepting external invites at all. The cheap change is available to one person; the expensive one needs a room. Predicting that in advance changes how you write the finding: if it lands as "the assistant did a bad thing", the cheap change is the natural response and the finding closes. ## What to actually say Four moves, in order. 1. **Credit the measurement, and bound it.** Confirm what you can confirm: the constructions tried no longer worked, on that build, on that date. Do not confirm that the class is closed, because you did not test the class. 2. **State the property that is unchanged**, in the language of who can write what and what it is trusted for, rather than in the language of one sentence that used to work. This is what keeps the conversation off the phrasing. 3. **Name the unowned decision.** "A model reads this field" was a change of consumer, and no team's review process fires on a change of consumer. The re-derivation is not on the calendar team's board and not on the assistant team's board. 4. **Offer the honest ending.** Either the residual gets an owner and a decision, or the organisation accepts it deliberately -- with a name attached and a date to revisit. Both are legitimate outcomes. A ticket closed as fixed, when what happened is a reduced likelihood against a fixed set of tried phrasings, is not. ## Deciding whether to re-open, and what that costs Re-opening a finding somebody has already declared closed spends credibility, so spend it on the right thing. Re-open when the claim on the ticket is stronger than the evidence -- "instructions in invites are ignored" is a claim about a class, and the evidence is a handful of runs. If the ticket instead says "reduced compliance with tested phrasings; residual accepted by <owner> until <date>", the finding is genuinely handled even though the carrier is still eligible, and pressing further is now an argument about risk appetite rather than about facts. ## The failure modes to name out loud - **Non-reproduction read as absence.** A construction that stops working proves that construction stopped working on that build. With a probabilistic system it is not proof of removal, and the same is true in the other direction: one success is not reliability. - **The assurance that quietly widens.** "We fixed the invite injection" starts as a note about one field and one test set and ends up cited as coverage for a whole channel. - **Filing against the reachable owner.** The assistant team is easy to reach, so findings land there, so remedies are prompt-shaped. That is a routing artefact, not an analysis. ## What a strong answer sounds like It does not say the patch was worthless -- it was not, and saying so loses the room. It says: here is what we now know, here is what did not change and why that is not the assistant team's fault, here is the decision nobody has made, and here is who I think should make it. That is a judgment about ownership and honest claims under real cost constraints, which is what the question is testing.

  • They ask you to confirm the fix works. What can you honestly confirm?
    That the constructions you tried no longer produced the behaviour on that build on that date. Not that unseen phrasings will not, because the change moved a preference rather than removing the field's eligibility, and not that the class is closed, because a class was never what you measured. Put the scope in the sentence, not in a footnote.
  • Who should own re-deriving what that field is trusted for?
    Neither team alone can answer it: the calendar owners control who may write the field, the assistant owners control what reads it and what it may do. It escalates to whoever owns the product surface that joined them. The concrete ask is that a named person accepts or funds the decision, with a revisit date -- not that a particular change be made.
  • Is accepting the residual ever the right call?
    Yes. An organisation may reasonably decide the exposure is tolerable for now given what the alternative costs. What makes it legitimate is that it is decided rather than assumed: named owner, stated scope, a date to revisit, and a claim in the ticket that matches the evidence. The defect is a closed ticket whose wording implies more than was tested.

saying these in an interview costs you the question

  • Accepts a prompt patch as a closed finding
  • Treats non-reproduction as proof of removal
  • Says the model was told not to, so it will not
  • Leaves residual risk with no named owner
  • Argues about the phrasing instead of the property

context