skip to content

How do you call Cohere's v2/embed endpoint and read the vectors from the response?

level: juniorimportance: must knowfreq 58%

answer

  1. one POST, one batch of texts
  2. vectors come back grouped, not flat
  3. order is the only join key
  4. embeddings.float, not data[].embedding
  5. input_type is required, not optional

basics

~20 s

Send texts plus a model and an input_type to POST /v2/embed. The response groups vectors by type: a float request returns them under embeddings.float, one vector per input text, in the order you sent them.

solid answer

~40 s

A Cohere embed call is a single POST to `/v2/embed` carrying `texts`, a `model` (for example an Embed v4 model id), an `input_type` describing what the text is for, and `embedding_types` listing the representations you want back. In the Python SDK that is `cohere.ClientV2(api_key=...)` followed by `co.embed(...)`. The response is not a flat list: it keys vectors by type, so a `["float"]` request puts them under `embeddings.float`, with one vector per input in the same order as the input array — you zip them back to your documents by position, never by any id in the response. `meta.billed_units.input_tokens` tells you what the call cost. The two mistakes that bite newcomers are expecting an OpenAI-shaped `data[].embedding` and forgetting `input_type`, which the current Embed models require.

code

python · 12 lines
python
import cohere

co = cohere.ClientV2(api_key="CO_API_KEY")

res = co.embed(
    model="embed-v4.0",
    texts=["how do I reset my password?", "where is my invoice?"],
    input_type="search_query",
    embedding_types=["float"],
)

print(res.meta.billed_units.input_tokens)

go deeper

for a junior

Be able to name the four things in a request — model, texts, input_type, embedding_types — and say that the response groups vectors by type and preserves input order.

for a middle

Explain why the response is keyed by type rather than flat, and why position is the only safe way to join vectors back to documents.

for a senior

Show that you treat ordering as a correctness hazard in an ingest pipeline, and that you track cost from meta.billed_units.input_tokens rather than estimating.

for a principal

Own the abstraction boundary: an internal embedding client should hide vendor response shapes so swapping providers does not leak Cohere's embeddings-by-type structure through your ingest code.

## What the endpoint is Cohere's embedding endpoint is `POST /v2/embed` on the Cohere API. It turns text (and, on the multimodal models, images) into fixed-length numeric vectors that you store in a vector index and compare later. Unlike a chat endpoint it is stateless, synchronous, and batch-shaped: one request carries many inputs and returns many vectors. ## The request Four things matter in a normal call: - **`model`** — the embed model id, for example an Embed v4 model. The model decides dimensionality, language coverage, context length and which compressed output types are available. - **`texts`** — the array of strings to embed. It is a batch: send many per call rather than one call per string. - **`input_type`** — what the text is *for* (`search_document`, `search_query`, `classification`, `clustering`, and `image` for image inputs). The current Embed generations require it; it is not an optional hint. - **`embedding_types`** — a list of the representations you want, such as `["float"]`, `["int8"]` or `["binary"]`. You may ask for several in one call. Optionally `truncate` (`NONE`, `START`, `END`) decides what happens to inputs longer than the model's per-input limit. ## The response shape — the part people get wrong The v2 response keys embeddings **by type**, not as a bare list: ``` { "id": "...", "embeddings": { "float": [[0.02, -0.01, ...], [...]] }, "texts": ["...", "..."], "meta": { "billed_units": { "input_tokens": 42 } } } ``` So the vectors for a float request live under `embeddings.float`; if you also asked for `binary`, that list sits beside it under `embeddings.binary`. Engineers arriving from other vendors reach for `data[0].embedding` (the OpenAI shape) or `embedding.values` (the Gemini shape) and get an attribute error. Cohere is its own shape. **Ordering is the join key.** The response carries no per-item identifier you can rely on; the i-th vector corresponds to the i-th string you sent. If you build the request by iterating a list of chunks, keep that list and zip it against the returned vectors. Reordering or filtering the batch between building it and consuming the response is a classic source of silently mislabelled vectors — every document ends up with its neighbour's embedding, retrieval still "works", and results are quietly wrong. ## Billing and usage `meta.billed_units.input_tokens` reports the tokens the call was billed for. Embeddings are billed on input only — there is no output-token side — so cost tracking for an ingest job is just the sum of that field across calls. Asking for several `embedding_types` in one request does not multiply the token bill; you pay for reading the text once. ## SDK vs raw HTTP The Python SDK's v2 client (`cohere.ClientV2`) mirrors the HTTP body one-for-one: the keyword arguments are the JSON field names, and the returned object mirrors the JSON structure. That symmetry is useful in interviews — if you can describe the JSON, you can describe the SDK call, and vice versa. There are also TypeScript, Go and Java clients over the same surface, plus Cohere models served through cloud marketplaces where the auth changes but the body does not. ## Errors you will actually see - **401** — missing or wrong API key (`Authorization: Bearer <key>`). - **400** — a bad or missing `input_type`, too many texts in one batch, or an over-long input with `truncate` set to `NONE`. - **429** — rate limited; back off and retry rather than hammering. ## What a junior should walk away with One endpoint, one batch of texts, a required `input_type`, a response keyed by embedding type, and order-based joining. Everything else — compressed vector types, batching strategy at corpus scale, multimodal inputs — builds on that call.

  • How do you match returned vectors back to your source documents?
    By position. The i-th vector in the response corresponds to the i-th string in your `texts` array; the response carries no per-item id to join on. Keep the exact list you sent and zip it against the response, and never filter or reorder that list between building the request and consuming the result — doing so shifts every document onto its neighbour's vector without any error.
  • Where do you find what an embed call cost?
    In `meta.billed_units.input_tokens` on the response. Embeddings are billed on input tokens only, so summing that field across an ingest run gives you the whole cost. Requesting several `embedding_types` in a single call does not multiply that number — you pay once for reading the text and get each representation you asked for.
  • What is the difference between this and calling the chat endpoint?
    Embed is stateless and batch-shaped: no roles, no messages, no streaming, no output tokens. You send an array of strings and get an array of vectors of fixed length. Cohere's generation endpoint (`/v2/chat`) is the conversational surface with messages and streamed events; the two share auth and nothing else structurally.

saying these in an interview costs you the question

  • Expecting an OpenAI-style data[].embedding field
  • Calling embed once per document instead of batching
  • Assuming input_type is optional metadata
  • Joining vectors to documents by some id in the response
  • Thinking embeddings are billed on output tokens

context