skip to content

What do Cohere's embedding_types int8 and binary return, and how much smaller are they?

level: middleimportance: should knowfreq 48%

answer

  1. it is a list, not a single value
  2. one byte per dimension versus one bit
  3. the array gets shorter, not just smaller
  4. roughly 4x and 32x
  5. your vector store must support the type

basics

~20 s

int8 returns one signed byte per dimension, about four times smaller than float32. binary packs one bit per dimension into bytes, so a 1024-dimension vector becomes 128 values — roughly thirty-two times smaller. You can request several types in one call.

solid answer

~50 s

`embedding_types` on Cohere's `/v2/embed` request is a **list**, and each entry is a different representation of the same vector. `float` is the baseline: 32 bits per dimension, so a 1024-dimension vector costs 4 KB. `int8` and `uint8` quantise each dimension into a single byte — same number of dimensions, about a quarter of the bytes. `binary` and `ubinary` go further: each dimension becomes one bit, packed eight to a byte, so 1024 dimensions come back as 128 packed values, roughly thirty-two times smaller than float. Because it is a list, one call can return `["float", "binary"]` together, keyed separately in the response under `embeddings.float` and `embeddings.binary`, billed as a single read of the input text. The catch is downstream: binary vectors must be compared with a bit-wise distance in your vector store, so the store has to support that type, and compression costs some retrieval accuracy that a rescore step usually recovers.

go deeper

for a junior

Know that float is the default full-precision form and that int8 and binary are smaller versions of the same vector, requested through the embedding_types list.

for a middle

Explain the mechanics: one byte per dimension for int8, one packed bit per dimension for binary, giving roughly 4x and 32x reductions, and that the binary array is dimensions divided by eight.

for a senior

Show you check vector-store support for the type before committing an ingest run, and that you quantify accuracy loss on your own eval set rather than trusting a published figure.

for a principal

Frame it as an index-generation decision: capture multiple representations in the one billed pass so a future storage or accuracy change does not force a full corpus re-embed.

## The parameter `embedding_types` is a list on the Cohere embed request naming the representations you want back. The available values on current Embed models are `float`, `int8`, `uint8`, `binary` and `ubinary`. The response keys each requested representation separately — ask for two, get two lists side by side. ## What each type actually is **float** — the model's native output, 32-bit floats, one per dimension. For a 1024-dimension model that is 1024 × 4 = 4096 bytes per vector. This is the highest-fidelity form and the baseline everything else is measured against. **int8 / uint8** — scalar quantisation. Each dimension is squeezed from a 32-bit float into a single 8-bit integer (signed for `int8`, unsigned for `uint8`). The vector keeps its full dimension count; only the precision per dimension drops. Size: 1024 bytes for a 1024-dimension vector, a 4× reduction. **binary / ubinary** — one bit per dimension, essentially the sign of each component, then packed eight bits to a byte. A 1024-dimension vector arrives as 128 byte-sized values (signed for `binary`, unsigned for `ubinary`). Size: 128 bytes, a 32× reduction against float. Note the count you get back is dimensions ÷ 8, which surprises people the first time — the array is *shorter*, not just smaller-typed, and its length is not the model's dimensionality. ## The size arithmetic that matters For 100 million vectors at 1024 dimensions: | type | bytes/vector | total | |------|--------------|-------| | float | 4096 | ~410 GB | | int8 | 1024 | ~102 GB | | binary | 128 | ~13 GB | That is the whole point. The float index needs a fleet of memory-heavy machines; the binary index fits on one. Compression also speeds up search, because distance computation is memory-bandwidth-bound and bit-wise comparison over packed bytes is far cheaper per candidate than float arithmetic. ## Requesting several types in one call Because `embedding_types` is a list, `["float", "binary"]` in a single request returns both representations of every input, under separate keys in the response. Two things follow: 1. **You are billed once.** Embed pricing counts input tokens, and you read the text once. Asking for a second representation is essentially free compared with making a second call — and far cheaper than re-embedding the corpus months later when you discover you need it. 2. **It future-proofs the ingest.** A common pattern is to store binary vectors in the hot index and keep the float (or int8) vectors in cheaper storage for rescoring and for future re-indexing, all captured in one pass over the corpus. ## The downstream constraints Compression is not free on the consumption side: - **Your vector store must support the type.** Binary vectors are compared with a bit-wise distance rather than the float distance used for the uncompressed form, and not every store or library implements that. Check before you commit an ingest run to a type. - **Accuracy drops.** Quantisation discards information, so the nearest-neighbour ordering shifts slightly. int8 typically loses very little; binary loses more. How much depends on your corpus and queries, which is why the honest answer is always "measure it on your own eval set". - **Model support varies.** Compressed types are a property of the embed model generation, not of the API in the abstract. Confirm the model you are pinning actually offers the type you plan to store. - **Mixed types do not interoperate.** A binary query vector is compared with binary document vectors; you cannot search a float index with a binary query. Query-time and index-time representations must match, exactly as the input type must. ## The standard mitigation The accepted pattern is two-stage: search the compressed index broadly to get a generous candidate set, then rescore those few hundred candidates with the higher-fidelity vectors you kept. You pay compressed-index cost for the expensive part (scanning millions of vectors) and float accuracy for the cheap part (scoring a few hundred). This is why storing both representations from the same call is such a natural default. ## What a strong answer sounds like Give the two mechanisms — one byte per dimension versus one bit per dimension packed — give the 4× and 32× numbers, mention that one call can return several types for one token bill, and finish on the constraint that the vector store must support the type and that accuracy loss is measured, not assumed.

  • Why does a binary embedding come back with fewer values than the model has dimensions?
    Because the bits are packed. Each dimension becomes a single bit, and eight bits are stored per byte-sized value, so a 1024-dimension vector arrives as 128 values. The array length is dimensions ÷ 8, not the dimensionality — code that validates vector length against the model's dimension count will reject perfectly good binary output unless it accounts for the packing.
  • Does asking for two embedding_types in one call double the cost?
    No. Cohere bills embed calls on input tokens, reported in `meta.billed_units.input_tokens`, and you read the text once regardless of how many representations you ask for. That makes `["float", "binary"]` a cheap hedge: you capture both forms in a single pass over the corpus instead of re-embedding everything later when your index design changes.
  • Can you convert float vectors to binary yourself instead of asking the API?
    You can threshold the components yourself, and for a rough index it may be adequate, but you lose the guarantee that your derivation matches what the API produces and would return for a query. Since requesting the type costs nothing extra on the same call, take the API's version so index-time and query-time representations are produced identically.

saying these in an interview costs you the question

  • Thinking embedding_types takes only one value per call
  • Expecting a binary vector to have as many values as dimensions
  • Assuming a second embedding type doubles the token bill
  • Believing compression is lossless because search still returns results
  • Searching a float index with a binary query vector

context