Cohere's v1 web-search connector is gone in v2 — how do you ground on the web now?
answer
- v1 had a connectors list
- managed web-search plus custom connectors
- v2 grounds through documents or tools
- you run the search, you return results
- citations then point at tool output
basics
~20 sYou run the search yourself. Cohere's v1 chat accepted a connectors list naming a managed web-search connector; v2 drops connectors, so web grounding becomes an ordinary tool-calling loop where your code performs the search and returns the results, which the model then cites.
solid answer
~50 sOn the v1 chat endpoint you could pass a `connectors` list — the managed `web-search` connector, or a custom connector you had registered as an HTTP service — and Cohere would issue the search, fold the hits into the grounded context, and cite them. The v2 chat API removed that layer. The replacement is tool use: you declare a search tool in the request, the model emits a tool call with a query, your code hits whatever search backend you choose, and you return the results in a tool message. The model then answers over those results and the citations point at the tool's output rather than at entries in the `documents` array. This is more code and one extra round trip, but it hands you the parts that mattered in production anyway — which search provider, what filters and freshness rules, what you cache, what you redact — and it removes a vendor-managed step from your data path.
go deeper
Recall that the managed connector belonged to the older v1 endpoint and that on v2 you supply grounding material yourself, either up front or through a tool.
Describe the loop: declare a search tool, receive a tool call with a query, run it, return the results in a tool message, and let the model answer and cite over them.
Argue the trade concretely — provider choice, filtering, caching, redaction and telemetry — and show you handle web text as attacker-controlled input inside your own perimeter.
Decide whether owning retrieval belongs in your system at all, and what your policy is for a vendor withdrawing a managed layer you had come to depend on.
## The v1 arrangement Cohere's first-generation chat endpoint bundled retrieval into the call. Alongside your message you passed a `connectors` list; naming the managed web-search connector caused Cohere to generate search queries from the conversation, run them, and use the hits as grounding material, returning citations against them. You could also register **custom connectors** — an HTTP service of your own exposing a search endpoint that Cohere would call during generation, so an internal wiki or ticket system could be grounded against without your application orchestrating anything. It was a genuinely appealing demo: one request, and a model that could answer about this morning's news over your intranet. It was also an unusual amount of control to give away, because the retrieval step — arguably the part that determines answer quality — ran inside the vendor. ## What v2 asks you to do instead The v2 chat API is organised around two grounding inputs: the `documents` parameter for material you already have, and tool calling for material that must be fetched during the turn. Connectors are not part of it. Web grounding therefore becomes an explicit loop you own: 1. Declare a search tool in the request, with a JSON Schema for its arguments — typically a query string, perhaps a recency or domain filter. 2. Send the user's question. If the model decides it needs to search, the response comes back with a tool call carrying the query it wants run. 3. Your code executes that query against whatever backend you have chosen — a commercial search API, your own index, a crawler cache. 4. Append the assistant's tool-call turn and a tool message carrying the results, keyed to the tool call, and call the endpoint again. 5. The model answers over those results, and citations attribute spans to the tool output rather than to documents you passed up front. The mechanics of that loop — schemas, tool-choice behaviour, multi-turn structure — are ordinary Cohere tool use. What matters for grounding is the last point: the citation machinery still works, but the source behind a span is a tool result. Code that assumes every citation source is a `documents` entry will break here, which is why source-handling code should branch on the source type. ## Why the change is defensible Once you have run a search-grounded assistant in production, the things you need are exactly the things a managed connector hides: - **Provider choice and cost.** Search APIs differ enormously in price, freshness and coverage, and you will want to switch or blend them. - **Filtering.** Domain allow-lists and deny-lists, recency windows, language and region constraints, paywalled-content rules. - **Caching.** Repeating identical searches across users is pure waste; a cache in front of your own search step is trivial and impossible if the vendor searches for you. - **Redaction and compliance.** If the conversation contains customer data, you decide what leaves your perimeter in a search query. - **Observability.** You get the queries, the hit lists and the latencies as your own telemetry. So the migration reads less like a regression and more like the vendor withdrawing from a layer it should not have been in. ## Prompt-injection is now visibly yours The most important consequence is a security one. Web results are attacker-controlled text: any page can contain instructions addressed to a model. Under a managed connector this risk existed but was easy to forget. Once your own code fetches and forwards the page text, the exposure is unmistakably in your system, and you must handle it — constrain what the model is permitted to do in the system message, never let a web-grounded answer trigger a privileged action without an independent check, strip active content, and cap how much of any single page you forward. The upside of owning the loop is that you *can* do all of that. ## Interview framing The answer that lands is not a recitation of the removed parameter. It is: v1 offered managed connectors, v2 asks you to implement retrieval as a tool, the citation model survives the change with tool outputs as sources, and the trade is more orchestration in exchange for control over provider, filtering, caching and the trust boundary. If asked whether you would want the connector back, the honest answer is that you would want it for prototypes and not for production.
- What did a custom connector look like in the v1 arrangement?It was an HTTP service you stood up and registered with Cohere, exposing a search endpoint that took a query and returned documents, with auth configured at registration. During a chat call Cohere would call it and fold the results into the grounded answer. The appeal was zero orchestration in your app; the cost was that your search backend had to be reachable from the vendor.
- How do citations differ when the grounding came from a tool rather than the documents array?The span offsets work identically, but the supporting source is a tool output instead of an entry in documents. So your mapping code has to branch on the source type and resolve tool-sourced citations through your own result identifiers. Normalising both shapes into one internal citation type early keeps the rendering layer from caring which grounding path produced the answer.
- What is the main risk you inherit by fetching web results yourself?Prompt injection. Web pages are attacker-controlled text and can contain instructions aimed at the model, which you are now forwarding into the context. Constrain the model's permitted actions in the system message, never let a web-grounded answer trigger a privileged operation without an independent check, strip active content, and cap how much of any one page you forward.
saying these in an interview costs you the question
- Thinks v2 still accepts a connectors parameter
- Assumes web grounding is a boolean flag on the request
- Believes citations stop working when grounding comes from a tool
- Treats fetched web text as trusted because your own code fetched it
- Sends unfiltered conversation content into a third-party search query