How does modelPreferences steer model choice in an MCP sampling/createMessage request?
answer
- you can ask, but you cannot choose
- fuzzy name matching, not exact ids
- three axes to trade off against each other
- values run from zero to one
- the result tells you what really ran
basics
~20 smodelPreferences carries a hints array of loose model-name substrings plus costPriority, speedPriority and intelligencePriority weights between 0 and 1. All of it is advisory: the client selects the model, and the result's model field reports what actually ran.
solid answer
~50 sA server cannot pick the model — it does not own the account — so `modelPreferences` is how it expresses what it wants. Two mechanisms sit inside it. `hints` is an ordered array of objects with a `name` field holding a substring of a model name; the client matches loosely and may map a hint onto an equivalent model from a different provider entirely. Alongside that are three normalised weights, `costPriority`, `speedPriority` and `intelligencePriority`, each between 0 and 1, saying which axis matters when a tradeoff has to be made. The client combines these with its own availability, policy and cost constraints and makes the final call, so a server must never assume it got what it asked for. Read `model` on the `CreateMessageResult` to learn what actually generated the message. Sampling as a whole is deprecated in MCP revision 2026-07-28.
go deeper
Recall that modelPreferences holds name hints plus cost, speed and intelligence priorities, and that it only suggests — the client picks the model.
Explain that hints are loose substrings that may map onto another provider's equivalent, that the three weights run from 0 to 1 and express tradeoffs, and that the result's model field is where you learn what really ran.
Demonstrate defensive design: validate the completion instead of trusting a requested tier, log the returned model, and recognise that any behaviour requiring a guaranteed model class should not be built on an advisory hint.
Treat this as a control-boundary question. The server can express intent but cannot enforce it, which is precisely why cost, capability and provider guarantees push toward owning the model integration rather than borrowing one through a deprecated capability.
## Why hints instead of a model name Sampling exists so a server can use a model it does not pay for. That immediately creates a mismatch: the server knows what the task needs, but the client knows what models exist, what they cost, and what the user's policy allows. A field that let the server name a model outright would fail constantly — the named model might not be configured, might be from a provider the host does not use, or might be one the user has deliberately disabled. `modelPreferences` resolves this by letting the server describe its needs in two complementary ways, both advisory. ## The hints array `hints` is an ordered array of objects each carrying a `name`. The name is treated as a **substring** matched loosely against model names, not as an exact identifier. A hint of `"sonnet"` says "something in that family"; the client is explicitly free to interpret it flexibly and may substitute a comparable model from a different provider. Ordering expresses preference — earlier hints are tried first. The practical consequence: hints are a way to say "a mid-sized general model" or "a small fast one" in the only vocabulary both sides share, which is fuzzy model naming. They are not a contract. ## The priority weights Three normalised values, each from 0 to 1: - `costPriority` — how much cheapness matters. - `speedPriority` — how much low latency matters. - `intelligencePriority` — how much raw capability matters. They are independent weights rather than a distribution that must sum to one, and they exist because the interesting question is not "which model" but "which axis do I sacrifice". A classification step inside a tool call might send `costPriority: 0.9, speedPriority: 0.8, intelligencePriority: 0.2`; a step that drafts prose a human will read inverts that. A server that fills in only hints, or only priorities, is fine — both parts are optional, and so is the whole `modelPreferences` object. ## Everything here is advisory This is the point interviewers are checking. The client makes the selection. It weighs the server's preferences against model availability, its own cost controls, the user's configuration and any policy the host enforces. It may honour the hints exactly, approximate them, or ignore them and use its single configured model. There is no error for "could not satisfy preferences" and no negotiation round trip — sampling has no handshake in which to negotiate one. So the server must be written defensively. Two habits follow: 1. **Read the result's `model` field.** Every `CreateMessageResult` reports the model that generated the message. This is the only reliable signal about what you actually got. 2. **Validate the output rather than trusting the tier.** If your logic genuinely breaks when a weak model answers, sampling is the wrong mechanism for it. ## How this interacts with statelessness MCP revision 2026-07-28 is stateless: every request is self-contained and servers must not rely on prior requests over the same connection. There is no "model chosen for this session" that persists — preferences travel on each `sampling/createMessage`, and a client is free to route two consecutive requests to two different models. A server that caches "the client uses model X" from an earlier completion and stops sending preferences is making exactly the assumption the stateless rule forbids. ## The deprecation angle Sampling is deprecated in revision 2026-07-28, with direct LLM provider integration as the migration path. That reframes `modelPreferences` neatly for an interview answer: the reason the field is only advisory is the reason the feature is going away. A server that truly needs guaranteed model characteristics — a specific capability tier, a known cost profile, a provider it has an agreement with — was never going to get them from a hint. Owning the provider call gives the server the control that `modelPreferences` could only ever request. ## Answering well Name both mechanisms (fuzzy hints, three priority weights), say clearly that they are advisory, and close with what the server does about it: check `model` on the result and degrade gracefully. That sequence shows you understand the control boundary rather than just the field list.
- What happens if the client cannot satisfy any of the hints?Nothing special — it simply uses whatever model it has. There is no error code for unsatisfiable preferences and no negotiation step; sampling has no handshake in which to negotiate. The server discovers what happened only by reading the `model` field on the `CreateMessageResult`, so any logic that depends on model class has to validate the output rather than trust the request.
- Do the three priority values have to sum to 1?No. `costPriority`, `speedPriority` and `intelligencePriority` are independent normalised weights, each between 0 and 1. Setting all three high is legal, it just expresses no tradeoff and gives the client no useful guidance. The field is most valuable when the values differ, because it tells the client which axis to sacrifice when it cannot have everything.
- Can a server learn once which model a client uses and stop sending preferences?No. MCP revision 2026-07-28 is stateless: every request is self-contained and servers must not rely on prior requests to establish context. A client may route consecutive completions to different models, and there is no session in which a model choice persists. Preferences belong on every `sampling/createMessage` that has an opinion about them.
saying these in an interview costs you the question
- Treats a hint as an exact model identifier that must be matched
- Believes the client is obliged to honour the priority weights
- Expects an error when preferences cannot be satisfied
- Thinks the three priorities must sum to one
- Assumes the chosen model stays fixed across successive calls