A system prompt says "always cite sources" and the user says "no citations" — what should the assistant do?
answer
- conflict has more than two outcomes
- obey, partial, refuse, or ask
- who is the constraint protecting?
- most conflicts are partial, not total
- write the resolution into the prompt
basics
~20 sIt depends on why the constraint exists. For a stylistic preference, satisfy the user's intent as far as the constraint allows and say what was kept. For a constraint protecting someone other than the user, keep it and refuse the removal — briefly, without lecturing.
solid answer
~50 sConflict resolution has four outcomes, not two: obey the higher layer outright, comply partially and disclose what you held back, refuse, or ask for clarification. Picking between them turns on what the constraint is *for*. If citations exist because the product wants a certain look, the right move is to serve the user's underlying intent — a cleaner answer — while keeping the minimum the system layer requires, and to say so in one line: "I've kept short source links because this assistant is required to cite; here is the answer without the footnote block." If citations exist because the answers are medical or legal and a reader must be able to verify them, the constraint protects a third party and should survive: state that plainly and offer what you *can* change. Silent partial compliance is the worst outcome, because the user believes their instruction landed when it did not.
go deeper
Know that a system-level constraint normally survives a conflicting user request, and that the assistant should say what it kept rather than quietly ignoring the user.
Be able to name all four outcomes — obey, partial compliance with disclosure, refuse, ask — and pick between them by asking who the constraint protects. Explain why silent partial compliance is the worst of them.
Show that most conflict quality is designed in advance: constraints marked negotiable or fixed, explicit fallbacks, a stated default for unanticipated collisions. Talk about evaluating both adherence and over-refusal.
Own the policy for which constraints in a product are user-overridable at all, and how that is decided and reviewed. Treat over-refusal as a product defect with the same weight as a violated constraint, and set the metrics for both.
## Conflicts are not all the same kind The useful first move is to classify the constraint the user is pushing against, because the ordering rule ("system outranks user") tells you who wins but not *how* the assistant should behave while winning. **Preference constraints** exist to make the product feel a certain way: tone, length, format, whether to use headings, whether to open with a summary. The system author would almost certainly say "fine" if asked. The user is the beneficiary of the constraint, so a user who wants it changed has legitimate standing. **Policy constraints** exist to protect someone other than the person typing: a required disclaimer, a limit on the advice a non-clinician assistant gives, verifiability of a claim, a rule about what data may be echoed back. The user is not the only stakeholder, so their preference does not dissolve the constraint. Citations sit in either bucket depending on the product, which is exactly what makes it a good interview question. "Always cite sources" in a general writing assistant is a preference. The same words in a clinical-summary tool are policy. ## The four outcomes **Obey the higher layer and say nothing.** Appropriate only when the user's instruction was incidental — a passing formatting aside where silently keeping the constraint costs them nothing. Rarely the right choice for an explicit request. **Partial compliance with disclosure.** The workhorse. Find the largest part of the user's intent you can satisfy inside the constraint, do that, and state in one sentence what you kept and why. For the citation case: drop the footnote apparatus and the inline bracket noise, keep a short source list, and say that the assistant is required to attribute claims. The user gets most of what they wanted and, crucially, an accurate picture of what they got. **Refuse the specific removal.** Reserved for policy constraints where compliance would mislead or expose someone. Refuse the *modification*, not the underlying task — decline to drop the disclaimer, then answer the question. A refusal that also abandons the user's actual request is a much worse answer than one that separates the two. **Ask.** Correct when the conflict is genuinely ambiguous — most often when the system prompt did not anticipate the situation and both readings are defensible, or when the user's instruction is broad enough that the cheap fix is to learn which part they cared about. Asking is expensive: it costs a turn and, in an autonomous agent, may cost the whole run. Use it when the ambiguity is real, not as a way to avoid deciding. ## Choosing well Three questions decide it. *Who does the constraint protect?* If only the user, lean toward compliance. *Is the conflict total or partial?* Most "conflicts" are partial — the user objects to the visible form of a constraint, not its substance, and there is usually a form that satisfies both. *Would the system author be surprised?* If you can predict the answer confidently, act on the prediction rather than asking. ## Design the resolution in advance The deeper lesson, and the one a strong candidate volunteers: most conflict-handling quality is decided when the system prompt is written, not at runtime. A system prompt that lists constraints without saying which are user-overridable forces the model to guess, and it will guess inconsistently across phrasings and across model versions. Write the resolution in. Mark each constraint as negotiable or fixed. State the fallback explicitly: "If the user asks for a format that would drop attribution, keep a minimal source list and tell them why." Distinguish the substance from the presentation so the model has somewhere to compromise. And state a default for unanticipated conflicts, because there will be some — usually "prefer the user's request unless it would remove a disclosure, exceed a stated limit, or misstate a fact." This has a testing payoff too: once resolutions are written down, they are assertable. You can eval that the citation constraint survives a user asking for prose, and separately that the tone constraint does *not* survive a user asking for something blunter — over-refusal is a real failure mode and it needs its own tests. ## Anti-patterns **Silent partial compliance.** The model half-honours the request without saying so. The user reads the answer as having complied and builds on a false belief about what the system does. This is the failure mode worth naming explicitly in an interview. **Blanket refusal on a preference conflict.** Refusing to shorten an answer because a system prompt mentioned thoroughness reads as broken, not as safe, and it trains users to fight the assistant. **Lecturing.** One clause explaining what was kept is enough. A paragraph on the assistant's guidelines makes the constraint feel like an obstacle. **Asking as a reflex.** Every clarification is a round trip. In an agent loop, a clarifying question at a point where no human is watching stalls the run entirely.
- When is asking for clarification the wrong response even though the conflict is genuine?Whenever nobody is there to answer, or when you can predict the answer. In an autonomous agent run a clarifying question can stall the whole task, so the better move is to act on the most likely reading and record the assumption in the output. Asking is also wrong when it is really an avoidance of a decision the system prompt should already have made — if the same clarification comes up repeatedly, the fix is to write the resolution into the prompt, not to keep asking.
- How would you write a system prompt so the model does not have to guess how to resolve conflicts?Mark each constraint as negotiable or fixed, separate substance from presentation so there is room to compromise, and state the fallback for the common collisions explicitly. Add a default rule for unanticipated conflicts — typically prefer the user's request unless it would drop a disclosure, exceed a stated limit, or misstate a fact. The payoff is testability: written-down resolutions become eval assertions, including tests that negotiable constraints actually do yield.
- Why is over-refusal treated as a real defect rather than the safe side of the tradeoff?Because it destroys the product's usefulness and pushes users to work around the assistant, which is worse for safety than a cooperative one. An assistant that refuses to shorten an answer, or that appends a disclaimer to a harmless question, teaches users that its constraints are noise to be defeated. Good evals therefore test both directions: that fixed constraints survive pressure, and that negotiable ones actually yield when a user asks.
saying these in an interview costs you the question
- Treating every system-user conflict as either full obedience or refusal
- Silently keeping a constraint without telling the user it was kept
- Refusing the whole task when only one modification was objectionable
- Asking for clarification on conflicts the system prompt already settles
- Assuming the user's latest instruction supersedes standing constraints