How do namespaces and toolsets stop overlapping agent tools from being selected wrongly?
answer
- catalogs accrete, they are not designed
- the prefix says who owns the intent
- not loaded means not misselected
- routing can be wrong too
- duplicates should be merged, not renamed
basics
~20 sNamespaces prefix each tool with its owning domain, so two similar capabilities read as different things and the model can see who owns which intent. Toolsets go further and keep the colliding pair out of the same prompt by loading only the tools eligible for the current role or request class.
solid answer
~50 sWhen several teams contribute to one catalog, near-collisions are the norm: HR's `request_leave` and payroll's `request_advance` both talk about time and pay, and the model has to break a tie nobody ever resolved. Namespacing attacks the ambiguity in the naming layer — every tool carries its owning domain as a prefix, so `hr.request_leave` and `payroll.request_advance` are visibly different objects and the domain word itself becomes selection signal. Toolsets attack it by eligibility: you group tools and load only the set relevant to the current role, surface or request class, so a leave request never sees the payroll toolset at all. Namespacing helps the model reason; toolsets remove the decision entirely, which is stronger but needs a routing signal you can trust. Where two tools genuinely do the same thing, neither mechanism is the answer — merge them and retire one, because a catalog with true duplicates will keep producing coin flips.
code
json · 10 lines[
{
"name": "hr.request_leave",
"description": "Submit a paid-time-off request for the signed-in employee."
},
{
"name": "payroll.request_advance",
"description": "Request early payout of already-earned wages before payday."
}
]go deeper
Know that prefixing tools with their domain — hr. or payroll. — makes similar tools easier to tell apart, and that you can load only the tools relevant to the current user.
Explain the difference between the two levers: namespacing helps the model reason about ownership, while toolsets remove the choice by not loading the alternative at all. Name the case where merging beats both.
Own the tradeoff: scoping moves the failure upstream into routing, where a mistake reads as a missing capability. Verify changes on per-pair confusion rather than aggregate accuracy, and never treat a namespace as an authorization boundary.
Set catalog governance across teams: who owns a namespace, how boundaries between overlapping domains are arbitrated, when a duplicate is retired rather than renamed, and how eligibility routing is kept coarse enough that users never hit a phantom missing feature.
## Where the collisions come from A large tool catalog is rarely designed; it accretes. Each team ships tools that are unambiguous inside their own domain and ambiguous across it. HR's leave request, payroll's wage advance, facilities' desk booking and IT's out-of-office configuration all touch "I'm away next week". None of these teams did anything wrong. The catalog as a whole has a boundary problem, and boundary problems are structural — you cannot fix them one tool at a time. ## Namespacing: making tools visibly different The first lever is naming. Every tool carries its owning domain: `hr.request_leave`, `payroll.request_advance`, `facilities.book_desk`. Three things follow. - **Lexical separation.** Two names that were near-identical now differ in their most prominent token. That helps the model, and it helps keyword search over the catalog, which can now match on the domain word. - **A place to put the boundary.** The namespace answers "whose responsibility is this?" without the model having to infer it from prose. - **Ownership.** The prefix names the team that owns the tool, which matters enormously when a selection error needs fixing and you must find someone to fix it. Namespacing is cheap and non-invasive, but it is advisory: the tools are still all present, and the model can still pick across namespaces. It reduces error rates; it does not eliminate the decision. ## Toolsets: removing the decision The stronger lever is eligibility. Group tools into sets and load only the set that applies to the current context — the user's role, the product surface they are in, the classification of the request, or the phase of a workflow. A user asking about leave in an HR surface never has the payroll tools in context, so the collision cannot occur. This is strictly more powerful and strictly more dangerous. Powerful because a tool that is not loaded cannot be misselected. Dangerous because the routing decision has moved upstream, and now *it* can be wrong — and when it is wrong, the agent does not misselect, it reports that the capability does not exist. That failure is quieter and more confusing than a wrong tool call. So: keep the sets coarse, prefer a union when the request is ambiguous, and instrument how often the agent reports a missing capability that the full catalog actually contains. ## The third option: merge Sometimes two tools are not overlapping, they are duplicates — two teams built the same capability against the same backend. No amount of naming or scoping fixes that; the model is being asked to choose between indistinguishable options and will keep flipping coins. Merge them, retire one, and point the retired name at the survivor. Catalogs need pruning as much as they need governance, and pre-production systems especially should delete duplicates rather than keep both alive with disambiguating prose. ## Choosing between them - A handful of near-neighbours in one otherwise clean catalog: namespace, and check the selection suite. - Distinct user populations or surfaces with mostly disjoint needs: toolsets, because the routing signal already exists and is reliable. - Genuine duplicates: merge. - A catalog large enough that even the eligible set is big: combine toolsets with deferred definitions and search, so the eligible set defines the search space rather than the prompt. ## How you know it worked The measurable artefact is a per-pair confusion matrix over a labelled selection set. Before the change, errors concentrate on specific pairs; after, either the pair separates or it does not, and you can see which. Aggregate accuracy alone hides this — a two-point overall improvement can be a big win on one pair and a regression somewhere else. Track the pairs. ## What this does not cover Namespacing is about structure, not prose. How a single tool's description is worded, which enums it exposes and how its schema constrains arguments is a separate discipline; here the lever is the shape of the catalog. And namespaces are not a permission boundary — a tool the model can see is a tool the model may call, so scoping for safety needs real authorization, not a naming convention.
- What is the risk of aggressively scoping tools into per-role toolsets?You move the error upstream. A misrouted request no longer produces a wrong tool call — it produces the agent claiming the capability does not exist, which is harder for users to report and for you to detect. Keep sets coarse, prefer the union when the request is ambiguous, and instrument how often the agent reports a missing capability the full catalog actually provides.
- Do namespaces give you any security benefit?No. A namespace is a naming convention, not an authorization boundary — any tool present in context can be called, whatever its prefix. Restricting who may invoke what has to be enforced where the tool executes, through real authorization checks on the caller's identity. Treat namespacing as a clarity mechanism and never as a permission model.
- How would you verify that a namespacing change actually helped?Compare per-pair confusion on a labelled single-turn selection set before and after. Aggregate accuracy hides the story: a small overall gain can combine a large win on the pair you targeted with a regression elsewhere. Look at whether the specific colliding pair separated, and check that no previously clean pair started confusing.
- When should you merge two tools instead of disambiguating them?When they genuinely do the same work against the same system. If a labelled set cannot produce a rule that says which one a request should hit, the model cannot either, and every disambiguating sentence you add is a workaround for a duplicate. Retire one and redirect its name — catalogs need pruning as much as governance.
saying these in an interview costs you the question
- Treats a namespace prefix as an access-control mechanism
- Assumes scoping tools per role has no downside
- Keeps two duplicate tools alive and disambiguates them with prose
- Judges the change on aggregate accuracy without looking at pairs
- Believes renaming alone can separate tools that do the same thing