How do you embed images with Cohere's Embed API, and what makes them comparable to text?
answer
- bytes go in the request, not a URL
- base64 inflates the body by a third
- one model, one shared vector space
- batch sizes far smaller than for text
- still billed in tokens, not per image
basics
~20 sImages are passed to Cohere's embed endpoint as base64-encoded data URIs rather than as remote URLs or file uploads. Because one multimodal model produces both text and image vectors in a single shared space, a text query can retrieve images directly.
solid answer
~50 sCohere's multimodal Embed models accept images as **base64-encoded data URIs in the request body** — there is no separate upload endpoint and no fetching of remote URLs, so your pipeline reads the bytes, encodes them and sends them inline. The `input_type` parameter still applies: image content is submitted with the image role, while the text you will search with stays `search_query`. The reason a text query can retrieve an image at all is that a single model projects both modalities into **one shared vector space**, so a caption-free screenshot and the sentence describing it land near each other without any intermediate captioning step. Operationally the constraints are the ones inline base64 implies: request bodies get large, so image batches are far smaller than text batches, and you should resize and re-encode before sending. Billing is still token-based — an image is converted to a token count, so large images cost more.
go deeper
Know that images are sent inline as base64 data URIs in the request body, and that the same model embeds both text and images.
Explain the shared vector space — one model places text and images near each other — which is why a text query can retrieve an image with no captioning step.
Show the operational side: payload inflation forces much smaller image batches, images need downscaling and their own timeouts, and cost is still measured in tokens.
Own the modality strategy — direct image embedding versus an OCR or captioning path, one mixed collection versus per-modality collections — and require evaluation on your own assets before committing.
## What multimodal embedding means here A multimodal embedding model maps more than one kind of input into the **same** vector space. That is the whole feature: the vector for a photograph of a red sneaker and the vector for the phrase "red running shoe" are close, because one model was trained to place them close. Practically, it means a product-image catalogue can be searched with plain text, or an image can be used as the query against a corpus of both images and descriptions, with no captioning model in between. This is different from the common workaround, where you run images through a vision-language model to produce captions and then embed the captions. Captioning loses whatever the captioner did not think to mention; direct image embedding keeps the visual signal. ## How images reach the API Images are supplied **inline as base64-encoded data URIs** in the request body — the `data:image/jpeg;base64,...` form. There is no separate file-upload step to reference later, and the API does not fetch remote HTTPS URLs on your behalf. Your ingest code therefore: 1. reads the image bytes, 2. optionally resizes and re-compresses them, 3. base64-encodes, 4. sends the data URI in the embed request with the image role on `input_type`. The exact field name for image content has changed across Embed generations — earlier multimodal support used a dedicated images array with a very small per-call limit, while the later generation accepts mixed text-and-image content in a unified inputs structure. Pin your model, read its reference for the current field, and do not assume a shape you saw in an older tutorial. ## The practical constraints **Base64 inflates payloads by about a third.** A 2 MB photo becomes roughly 2.7 MB of JSON. That single fact drives most of the operational advice: - **Batches are tiny compared with text.** Where a text batch might carry dozens of chunks, an image batch is a handful or even one. Build your batcher to be modality-aware rather than reusing the text batch size. - **Resize before sending.** Embedding models work at a fixed internal resolution, so shipping a 12-megapixel original wastes bandwidth, latency and often tokens without improving the vector. Downscale to a sensible bound and re-encode as JPEG or PNG in the ingest step. - **Latency per call is higher.** Upload time dominates, so a fixed timeout tuned for text calls will spuriously fail image calls. Give image requests their own timeout and retry budget. - **Memory in the ingest worker.** Holding several base64 strings of multi-megabyte images at once is an easy way to blow a container's memory limit. Stream and release. ## Billing Embed billing is token-based, and image inputs are converted into a token count for that purpose — so image embedding costs more per item than a short text chunk, and larger or more detailed images cost more than small ones. The response's `meta.billed_units.input_tokens` remains the authoritative number; pilot on a sample of real images and read it rather than estimating from file size. ## Index design consequences Because both modalities land in one space, you can put image vectors and text vectors in the **same** collection and search across both with one query. Whether you should is a design decision: a mixed index gives genuine cross-modal recall, while separate collections give you per-modality control over ranking and filtering. Either way, keep the invariants that apply to all embeddings — one model, one embedding type, consistent input roles — and store the modality and the source reference as metadata alongside each vector so you know what a hit actually points at. ## Where the limits are Text-image alignment is genuinely useful but coarse: it captures subject matter and gross visual attributes far better than fine detail, embedded text within an image, or precise counts. For document images dense with text, an OCR-then-embed-text path often beats pure image embedding. Knowing that boundary — and testing it on your own data instead of assuming — is what distinguishes a considered answer from an enthusiastic one. ## What a strong answer sounds like Say the images go inline as base64 data URIs, that one shared vector space is what makes cross-modal search work, that batches must be far smaller than text batches because of payload inflation, and that billing is still token-based. Then name a limitation, such as text-heavy document images.
- Why not caption images with a vision model and embed the captions instead?You can, and for text-heavy documents it often works better, but a caption is a lossy summary — whatever the captioner did not mention is gone from the vector. Direct image embedding keeps visual signal the captioner would have skipped, and removes a whole model from the ingest path along with its latency, cost and failure modes. The tradeoff is that captions are inspectable and debuggable while raw image vectors are not.
- Can image vectors and text vectors live in the same collection?Yes, provided the same model produced them, since they occupy one shared space and distances between them are meaningful. That gives cross-modal retrieval from a single query. The alternative — separate collections per modality — buys you independent ranking, filtering and scaling at the cost of merging results yourself. Either way, store the modality as metadata so you know what each hit refers to.
- What should your ingest do before base64-encoding an image?Resize to a sensible bound and re-compress. The model works at a fixed internal resolution, so a 12-megapixel original adds upload time, memory pressure and token cost without improving the vector, and base64 inflates whatever you send by roughly a third. Downscaling in the ingest step typically cuts payloads by an order of magnitude with no measurable retrieval loss.
saying these in an interview costs you the question
- Expecting the API to fetch images from public URLs
- Reusing the text batch size for image requests
- Assuming images are billed per image rather than in tokens
- Thinking the API captions the image and embeds the caption
- Sending full-resolution originals without downscaling