skip to content

When is Cohere's in-API grounded generation the wrong choice for a RAG service?

level: principalimportance: nice to knowfreq 24%

answer

  1. the model, or the citation contract, or both
  2. is the provider swappable here
  3. mode dial is the only policy knob
  4. citations do not evaluate retrieval
  5. normalise into your own citation type

basics

~20 s

When attribution has to outlive the vendor. Cohere's documents-and-citations contract is specific to its chat API, so a service that must run the same grounding across several model providers, or needs citation rules the API does not expose, is better served by attribution it owns.

solid answer

~50 s

The feature buys you real things: no prompt-assembly convention to invent, no marker parsing, machine-readable spans with your own document ids, and a quality of attribution that is hard to match with prompting alone. The cost is that the contract is Cohere-shaped. Route the same question to a second provider and there is no `documents` parameter and no citations array on the other side, so you either keep two grounding paths or rebuild attribution generically. You also cannot change how attribution behaves beyond the mode dial — no per-claim confidence, no custom granularity, no rule that a numeric claim must cite. Decide by asking whether the model is a swappable component in your architecture. If it is, put grounding in a layer you own and treat any vendor citations as an optimisation; if you are committed to Cohere, take the feature — reimplementing it worse is not a win.

go deeper

for a junior

Know that citations here come from Cohere's own API contract, so an answer from a different provider will not arrive in the same shape.

for a middle

Name the concrete work you would redo elsewhere — source formatting, a citation instruction, a marker parser — and why that is more than porting the model call.

for a senior

Show the normalising design: one internal citation type filled by every grounding path, with the vendor object kept out of the UI and the database.

for a principal

Tie the decision to whether the provider is pinned or swappable and whether attribution is a product promise, and price the migration honestly rather than invoking lock-in as a slogan.

## The real question behind this one This is not "is the feature good" — it is good. It is a build-versus-buy question about where the grounding boundary sits in your architecture, and the answer follows from one thing: **is the model provider a component you expect to swap?** ## What you get by using it Be honest about the value before arguing against it. - **A solved prompt-assembly convention.** Everyone who builds this themselves invents a format for presenting sources and a marker syntax for citing them, then debugs models that half-follow it. - **No marker parsing.** Inline `[1]` markers must be parsed out of prose, and models emit malformed ones. Offsets plus source ids arrive structured. - **Better attribution than prompting usually achieves.** A dedicated attribution step generally beats asking a model to remember to cite, particularly on long answers. - **Your ids come back.** The join to your store is direct, which makes the citation actionable in the UI and loggable as a metric. Rebuilding all of that in prompt engineering is real work, and doing it badly is worse than using the feature. ## What it costs **Shape lock-in.** The `documents` parameter and the citations array exist on Cohere's chat API. Nothing else you might call has that exact contract. A multi-provider service — one that routes by cost, or keeps a second vendor for failover, or lets enterprise customers choose a model — either maintains a Cohere-specific grounding path plus a generic one, or normalises. The migration cost is not the API call; it is that your product's citation semantics were defined by a vendor's behaviour, and the fallback path will not reproduce them exactly. **No control over attribution policy.** Beyond the mode dial you cannot express "cite at clause granularity", "every numeric claim must carry a citation", or "give me a confidence per span". If your domain needs a specific attribution rule, you will end up post-processing the API's output — at which point you own attribution logic anyway, just split across two places. **It does not evaluate retrieval.** The most persistent misconception is that in-API grounding reduces the need to measure your pipeline. It does the opposite of nothing: citations tell you which of the documents you supplied supported a span, and are silent on whether the right document was retrieved at all. A confidently cited answer built on the wrong three chunks looks perfect in the response object. **A hidden coupling.** Because it is easy, teams push grounding decisions into the request and out of any layer they test. Nobody has a component that answers "how does our system attribute claims", because the answer is "the vendor does it". ## A decision rule **Use it when:** Cohere is a committed choice for this workload; the surface benefits from attribution but does not have a domain-specific citation policy; the team is small and the alternative is inventing a marker convention; you want production attribution quickly. **Own it when:** you route across providers or expect to; attribution is a regulated or contractual product feature with its own rules; you need attribution over outputs that are not a single chat completion, such as a multi-step chain where the final text is assembled from several calls; or your citation UI must be identical regardless of which model answered. ## The architectural compromise The pattern that survives both worlds is a thin internal citation type — span, source id, provenance — that every path normalises into. Cohere's grounded path fills it directly from the citations array. Another provider fills it from your own marker parsing or a post-hoc attribution step. The UI, the logging, and the coverage metrics consume only the internal type and never learn which path produced them. That costs one adapter per provider and buys you the ability to change your mind, which on a two-year-old service is usually the thing you most wish you had. The corollary is a discipline: never let the vendor's citation object leak into your rendering layer or your database. The day it appears in a React component is the day switching providers becomes a rewrite. ## What a strong answer sounds like Not "avoid lock-in" as a reflex — that argument, applied uniformly, produces lowest-common-denominator systems that use nothing well. The strong version names what specifically is locked in (the citation contract, not the model), estimates what re-implementing it costs (an attribution step and a marker parser, both mediocre), and ties the decision to a concrete architectural fact about the service: whether the provider is pinned or swappable, and whether attribution is a product promise or a nice affordance.

  • What exactly would you have to rebuild if you moved this workload to another provider?
    The grounding prompt convention, a citation instruction the model actually follows, a parser for whatever markers it emits, and an offset mapper from markers back to spans — plus tests for all of it. The model call itself is trivial to port. That asymmetry is the argument for keeping an internal citation type that both paths fill, so only an adapter changes.
  • Does using in-API grounding reduce the need to evaluate your retrieval pipeline?
    No, and assuming it does is the expensive mistake. Citations report which of the documents you supplied supported a span; they say nothing about whether the right document was ever retrieved. An answer grounded confidently in the wrong three chunks produces a flawless-looking response object. You still need retrieval metrics — recall at k on a labelled set — measured independently of generation.
  • Is there a middle path between using the feature and owning attribution outright?
    Yes: normalise. Define an internal citation type of span plus source id plus provenance, fill it directly from Cohere's citations on that path and from your own attribution step elsewhere, and let the UI, logs and metrics consume only the internal type. You pay one adapter per provider and keep the option to switch without rewriting the rendering layer.

saying these in an interview costs you the question

  • Rejects the feature on reflexive lock-in grounds with no specifics
  • Assumes another provider will accept a documents parameter too
  • Thinks vendor citations remove the need for retrieval metrics
  • Lets the vendor citation object reach the UI and the database
  • Plans to reimplement attribution by prompting, and expects parity

context