skip to content

How do you choose between Cohere's float and binary embeddings at 100M scale?

level: principalimportance: should knowfreq 30%

answer

  1. storage arithmetic first, opinion second
  2. the conjunction should be and, not or
  3. one billed pass, two representations
  4. search wide cheap, rescore narrow precise
  5. re-embedding is the real lock-in

basics

~20 s

Do not choose one. Request both types in a single embed call — the token bill is the same — then serve a compressed index for first-pass recall and rescore the top candidates with the higher-fidelity vectors. Validate the recall loss on your own eval set.

solid answer

~50 s

At a hundred million vectors the decision is an infrastructure decision, not a model one: 1024-dimension float vectors are around 410 GB while their binary form is around 13 GB, which is the difference between a memory-heavy fleet and a single machine. But framing it as float *or* binary is the trap. Because `embedding_types` is a list and Cohere bills embed calls on input tokens, one pass over the corpus can return `["float", "binary"]` for a single bill — so capture both. Then serve binary in the hot index for a wide first-pass retrieval and rescore the top few hundred candidates with float, which recovers most of the accuracy the compression cost. The number that decides the design is recall on **your** eval set at each configuration, not a published benchmark. And the real lock-in is not storage but re-embedding: changing the model later means re-reading the whole corpus, so keep the source text and record the model, input type and embedding types as index metadata.

go deeper

for a junior

Know that binary vectors are far smaller than float ones and that smaller vectors mean a cheaper index but some loss of accuracy.

for a middle

Do the storage arithmetic for a given corpus size and explain the two-stage pattern: retrieve wide from the compressed index, then rescore the top candidates with higher-fidelity vectors.

for a senior

Insist on measuring recall on your own eval set at several configurations, and show you know how candidate depth trades first-pass recall against rescoring latency.

for a principal

Own the reversibility argument: capture every representation in the one billed pass, record index provenance, retain source text, and treat re-embedding as the migration cost that constrains all future choices.

## Frame the question correctly Junior framing: "which type is better?" Senior framing: "what does each cost and what does each lose?" Principal framing: "what does this commit us to, and what will it cost to change our mind?" All three matter, but the last one is what the level is being probed for. ## The infrastructure arithmetic For 100M vectors at 1024 dimensions: - **float** — 4 KB each, roughly **410 GB**. - **int8** — 1 KB each, roughly **102 GB**. - **binary** — 128 bytes each, roughly **13 GB**. ANN indexes generally want their vectors resident in memory. 410 GB means sharding across a fleet of large-memory nodes, with the operational tax that implies: shard routing, rebalancing, replica cost multiplied per shard, longer index builds, slower recovery. 13 GB fits comfortably on a single machine with room for the graph structure on top, and replicas become cheap enough to over-provision for availability. That is not a 30× storage saving; it is a different system architecture, with different failure modes and a different on-call burden. Search throughput moves the same way. Distance computation over millions of candidates is memory-bandwidth-bound, so packed binary vectors scan far faster per candidate than float arithmetic. Compression buys latency headroom as well as capacity. ## Why "or" is the wrong conjunction `embedding_types` accepts a list, and embed pricing counts input tokens — you pay to read the text, not per representation returned. So `["float", "binary"]` in the same call yields both forms of every vector for the same token bill as either one alone. That converts the decision from an irreversible commitment into a storage question. Keep binary in the hot ANN index; keep float (or int8) in cheap storage — object storage, a columnar store, a disk-backed keyed lookup. Later, if evaluation says binary-only recall is inadequate, or if a new store gains a better quantisation mode, you already hold the high-fidelity vectors and never re-read the corpus. The alternative — embedding float-only now and adding binary in six months — means re-billing 100M documents' worth of input tokens and re-running an ingest that takes days. That asymmetry is the entire argument for capturing both up front. ## The two-stage serving pattern The standard design once you hold both: 1. **Search the binary index wide.** Retrieve substantially more candidates than you need — hundreds where you want ten — because compression perturbs ranking near the boundary. 2. **Rescore the candidates with float.** Fetch the float vectors for those few hundred ids and score them precisely. You pay compressed cost for the expensive step (scanning 100M vectors) and full fidelity for the cheap one (scoring a few hundred). Recall recovers close to the float-only baseline while the index stays small. The knob is the candidate count: too small and rescoring cannot recover what the first pass missed; too large and you lose the latency benefit. Tune it empirically. This is also where a reranking model would sit in a full search stack — but note that rescoring with float vectors and reranking with a cross-encoder are different stages solving different problems, and neither substitutes for the other. ## What actually decides it: your evaluation No published recall figure transfers to your corpus. Build a labelled eval set from real queries against real documents — a few hundred query/relevant-document pairs is enough to see differences — and measure recall@k and end-task metrics for: float-only, int8-only, binary-only, and binary-plus-rescore at a couple of candidate depths. Then read the numbers against your latency and cost budgets. Corpora with many near-duplicate documents suffer more from compression, because the discarded precision was exactly what separated the near-duplicates. Corpora of diverse topics suffer much less. Budget for re-running that evaluation whenever the model, the chunker or the corpus mix changes materially. ## The real lock-in Storage type is cheap to change if you kept both forms. What is genuinely expensive: - **Changing embedding model.** Every vector must be regenerated. That is the migration to plan for, and the reason to keep raw source text alongside the vectors forever. - **Changing input type.** Same cost, same reason. - **Losing provenance.** An index whose model id, input type and embedding types were never recorded cannot be safely extended or migrated — you cannot tell whether new vectors are compatible with old ones. So the principal-level deliverable is not a chosen type. It is: capture both representations in the one billed pass, record the generation metadata, keep the source text, define an eval set, and design the index so swapping the served representation is a config change rather than a project. ## What a strong answer sounds like Do the storage arithmetic out loud, refuse the false binary choice by pointing at the multi-type request and token-based billing, describe the two-stage rescore, insist on measuring recall on your own data, and close on re-embedding as the cost that actually constrains future choices.

  • How many candidates should the binary first pass return before rescoring?
    Enough that the relevant document is almost always somewhere in the set — typically an order of magnitude more than you finally serve, so hundreds for a top-ten result. Tune it by measuring recall of the first pass alone at several depths against your eval set: pick the smallest depth where first-pass recall plateaus, because everything beyond that only adds rescoring latency without adding correct answers.
  • When is binary compression clearly the wrong choice?
    When the corpus is full of near-duplicates that differ in fine detail — legal clauses, product variants, versioned policy documents. The precision binary discards is exactly what separated those neighbours, so ranking collapses among them. It is also wrong when the index is small enough that float fits comfortably in memory: you would be trading accuracy for a saving you did not need.
  • What metadata must you record alongside a production vector index?
    The embedding model id, the input type used for the corpus, the embedding types stored, the dimensionality, the chunking configuration, and the ingest date. None of it is recoverable from the vectors themselves, and all of it is needed to decide whether newly embedded documents can be added to the existing index or whether a change forces a full rebuild.
  • Why is re-embedding, rather than storage, the constraint that shapes this decision?
    Because storage is a config and capacity choice you can revisit, whereas regenerating 100M vectors costs a full re-read of the corpus in tokens plus days of ingest and a parallel index build. That asymmetry is why you capture every representation you might plausibly want during the single pass you are already paying for, and why the raw source text must be retained indefinitely.

saying these in an interview costs you the question

  • Treating it as a binary either/or choice
  • Quoting published recall numbers instead of measuring your own
  • Assuming a second embedding type doubles the bill
  • Ignoring that changing model means re-embedding everything
  • Discarding source text once vectors are stored

context