How would you set ttlMs and cacheScope for server/discover across a fleet of MCP servers?
answer
- two fields, two different decisions
- can two users see different tools?
- TTL is the blast radius of a change
- default private, justify public
- shorten during a rollout, lengthen after
basics
~20 sSet cacheScope by whether the discovery answer can differ per caller — private whenever capabilities vary by authorization, public only when every caller sees the same profile. Set ttlMs from how fast the profile actually changes, shortening it during rollouts.
solid answer
~50 sTreat the two fields as separate decisions. `cacheScope` is a correctness call: because a server's advertised capability set MAY vary by the authorization presented in revision 2026-07-28, any server whose profile differs between an admin and a read-only caller must declare `"private"`, and `"public"` is a promise that every caller sees the same answer. Get it wrong in the permissive direction and one privileged discovery seeds a host cache that other users read. `ttlMs` is an operations call, trading staleness against discovery load: a stable server pinned to one revision can name hours, while a server mid-migration — about to add a revision to `supportedVersions` or retire a capability — should name minutes, because clients that cached the old profile will keep acting on it for the whole window. Across a fleet I would default to private, shorten TTLs deliberately for the duration of a rollout, and lengthen them again once every node is serving the new profile.
go deeper
Know the two values you are choosing between and what they mean: a TTL in milliseconds, and a public-or-private sharing rule for the cached discovery answer.
Explain the deciding test for scope — whether two callers with different authorization could receive different answers — and that the TTL should follow how fast the server's profile actually changes.
Show the rollout discipline: shorten the TTL before a change to supportedVersions or capabilities so caches drain, and account for a heterogeneous fleet serving inconsistent profiles while a deploy is in flight.
Own the fleet policy — a private-by-default posture with public as a justified opt-in, a documented rollout TTL, and the clear statement that cache scope is hygiene while per-request authorization checks are the actual boundary.
## Why this is a judgment question `ttlMs` and `cacheScope` are required on any `DiscoverResult` in MCP revision 2026-07-28, so every server operator must answer them — there is no "leave it unset" path. The values are not derivable from the spec; they encode what you know about your own deployment. That makes them a small but real platform decision, and one that gets more interesting the more servers you run. ## cacheScope is a correctness decision, not a tuning knob The governing fact is that the tool, prompt and resource sets a server exposes MUST NOT vary per connection but MAY vary by the authorization presented. That is what makes per-caller discovery answers legitimate — and what makes a wrong scope a leak. Decision rule: if *anything* in the result — the capability object, the instructions, even the supported-version list under a per-tenant deployment — can differ between two callers, the scope is `"private"`. `"public"` is a claim that an anonymous caller and your most privileged caller receive byte-equivalent answers. Across a fleet I would default every server to `"private"` and treat `"public"` as an opt-in a service owner justifies. The cost of being wrong is asymmetric: `"private"` on a genuinely uniform server costs a few extra discovery calls, while `"public"` on a per-authorization server exposes the existence and shape of privileged capabilities to every user of a host that shares its cache. A useful review question is simply "can two of our users see different tools here?" — if the honest answer is "today, no, but the roadmap says yes", declare private now. ## ttlMs is a staleness-versus-load decision Start from the real question: how long can a client act on a stale profile before something goes wrong? Because the client is entitled to reuse the answer for the whole window, `ttlMs` is the blast radius of any change to the profile. - **Steady state, pinned revision, fixed capability set** — long is fine. Hours of reuse cost you nothing, and discovery traffic disappears. - **During a revision rollout**, when `supportedVersions` is about to gain or lose an entry, shorten well before the change lands so that in-flight caches drain, and keep it short until every node serves the new profile. This matters more than it sounds because a fleet behind a load balancer has no per-client affinity in 2026-07-28 — any node may serve any request — so a heterogeneous fleet mid-deploy hands out inconsistent profiles that clients then cache. - **When capabilities are derived from a fast-moving upstream**, the TTL should track that upstream's own volatility, not your deployment cadence. A practical shape is a fleet default in the tens of minutes, a documented "rollout mode" of a minute or two that a service owner flips for the duration of a migration, and a longer value for genuinely frozen servers. ## What you cannot fix with these fields Neither field is enforceable. `cacheScope` is an instruction to a cooperating client, and nothing stops a badly written host from sharing a private answer. So the scope is a hygiene control, never an authorization control: the actual gate on privileged capabilities is that the server checks the authorization presented on each request and refuses what that caller may not do. A candidate who describes `"private"` as a security boundary has the wrong model. Equally, a long TTL never makes a stale client dangerous in the protocol sense, because every request still declares its own version and capabilities and the server still evaluates it independently. The failure mode of staleness is a wasted round-trip and a confusing user experience — a tool the UI still lists and the server now rejects — not a bypass. ## The complementary lever Caching is not the only way to keep clients current. A host can open a `subscriptions/listen` stream and opt into the list-changed notifications, which lets it refresh a profile on a real change rather than on a timer. Where a host does that, a moderate TTL plus change notifications beats a very short TTL for everyone: you pay for discovery when something actually changed instead of on a fixed cadence.
- Is declaring cacheScope "private" a security control?No — it is hygiene, not enforcement. Nothing in the protocol stops a badly written host from sharing a private answer, so the real gate on privileged capabilities is the server checking the authorization presented on every request and refusing what that caller may not do. Private scope reduces accidental exposure by cooperating clients; it does not create a boundary.
- Why does a fleet behind a load balancer make TTL choice harder during a rollout?Because MCP 2026-07-28 statelessness means any node may serve any request, with no client affinity. Mid-deploy, two discovery calls from the same client can hit an old and a new node and return different profiles, either of which the client then caches for the full ttlMs. Shortening the TTL before the rollout keeps that inconsistency window small.
- What can reduce the need for very short TTLs?Change notifications. A host that opens a subscriptions/listen stream and opts into the list-changed filters learns when a server's tool, prompt or resource set actually changed, and can re-discover then. That lets you keep a moderate TTL and pay for discovery on real changes rather than on a fixed timer.
saying these in an interview costs you the question
- Treats cacheScope private as an authorization boundary
- Declares public because capabilities are the same today
- Sets one global TTL and never revisits it during rollouts
- Assumes a stale cached profile can bypass server-side checks
- Ignores that any fleet node may serve any request mid-deploy