skip to content

When would you pick text-embedding-3-large over text-embedding-3-small?

level: middleimportance: must knowfreq 68%

answer

  1. Small is the default; large is the upgrade
  2. 1536 versus 3072 dimensions
  3. Roughly a 6x price gap
  4. Storage cost outlives the embedding cost
  5. Measure recall@k on your own corpus

basics

~20 s

Pick text-embedding-3-large when retrieval quality is the bottleneck and the corpus is small enough that its higher per-token price and 3072-dimensional vectors are affordable. text-embedding-3-small is the sane default: far cheaper, 1536 dimensions, and close on benchmark quality.

solid answer

~40 s

Both are current-generation OpenAI embedding models. `text-embedding-3-small` returns 1536 dimensions by default and is roughly six times cheaper per token than `text-embedding-3-large`, which returns 3072. On OpenAI's published MTEB averages the large model leads by a couple of points — meaningful for hard retrieval, invisible for easy retrieval. My default is small, because the cost difference compounds twice: once on the one-off corpus embedding, and permanently on vector storage and index memory, where 3072 floats per chunk is double 1536. I move to large when measured recall on a real evaluation set is the limiting factor — long or technical documents, multilingual corpora, fine-grained distinctions between near-duplicate chunks. The legacy `text-embedding-ada-002` is not a live option for new work: it scores below both, costs more than small, and does not support the `dimensions` parameter.

go deeper

for a junior

Know the two current models by name, that small returns 1536 dimensions and large 3072, and that small is the cheaper default most projects start with.

for a middle

Explain the price and dimension gap, that the quality difference is a couple of MTEB points on average, and that shortening the large model is a third option between them.

for a senior

Demonstrate that you decide by measuring recall@k on a domain evaluation set, and that you count vector-store memory and query latency alongside the API bill.

for a principal

Own the framing that embedding width is a long-lived architectural commitment: it fixes storage, index memory and migration cost for the life of the corpus, so choose it with an eviction and re-embedding plan already in mind.

## The three models on the table As of mid-2026 OpenAI's embedding line-up is two current models plus one legacy: | Model | Default dimensions | `dimensions` support | Position | |---|---|---|---| | `text-embedding-3-small` | 1536 | yes | cheap default | | `text-embedding-3-large` | 3072 | yes | highest quality | | `text-embedding-ada-002` | 1536 | no | legacy, superseded | All three share an 8192-token maximum input and return L2-normalised vectors. ## The quality gap, honestly stated OpenAI's published benchmark numbers put `text-embedding-3-large` around 64.6% average on MTEB and `text-embedding-3-small` around 62.3%, with the legacy `ada-002` around 61.0%. A two-point average on a benchmark suite is a real but modest gap, and — this is the part interviewers listen for — it is an *average across tasks*, not a promise about yours. What that means practically: if your retrieval task is easy (short, distinct documents; queries that share vocabulary with the answer), small and large will return nearly the same top-k and you are paying six times more for noise. If your task is hard (long technical prose, many near-duplicate chunks, cross-lingual queries, questions phrased nothing like the source), the gap widens and large earns its price. The only defensible way to choose is to build a small labelled evaluation set — a few hundred query/expected-chunk pairs from your own domain — and measure recall@k for both. That measurement costs an afternoon and settles the argument permanently. ## The cost difference has two halves People usually reason only about the API bill. There are two costs: **Embedding cost.** Roughly a 6x per-token difference between small and large. This is a one-time cost at index build, plus a trickle for new documents and for every query you embed at runtime. For a corpus of a few million tokens it is negligible in absolute terms for either model; at hundreds of millions of tokens the multiplier starts to matter. **Storage and serving cost.** This one is permanent and often larger. A 3072-dimension float32 vector is 12 KB; a 1536-dimension one is 6 KB. Across ten million chunks that is 120 GB versus 60 GB before index overhead, and vector indexes are typically memory-resident. Doubling the dimension also roughly doubles the work per distance computation, so query latency and CPU cost rise. Many teams that "chose large for quality" discover the real bill arrived from their vector database, not from OpenAI. ## The middle path The choice is not binary, because both `text-embedding-3-*` models accept a `dimensions` parameter that returns a shorter vector. That turns a two-way model choice into a two-axis decision: which model, and at what output width. A common and defensible configuration is `text-embedding-3-large` shortened to 1024 or 1536 dimensions — you buy some of the large model's quality while keeping storage at or below small's footprint. OpenAI reports that `text-embedding-3-large` truncated to 256 dimensions still outperforms `text-embedding-ada-002` at its full 1536, which is the clearest evidence that dimension width and model quality are separable dials. ## Why ada-002 is not a choice for new work It is strictly dominated: lower benchmark quality than small, higher price than small, and no `dimensions` support so you cannot trade width for cost. Its only remaining role is compatibility — if you already have a large production index built with it, the question becomes whether a migration is worth it, not whether to start there. ## Constraints that override the quality argument - **A fixed vector-store dimension.** Some managed indexes are provisioned at a fixed width. If you are pinned to 1536, small at default width or large shortened to 1536 are your candidates, and the choice is decided by measurement. - **Latency on the query path.** Every user query needs one embedding call before retrieval can start. Both models are fast, but the wider vector also costs more in the nearest-neighbour search itself. - **Existing index.** You cannot mix models in one index. Switching means re-embedding everything. ## The answer that lands "Start on small, hold out an evaluation set, and only upgrade — model or width — when recall@k on that set says the retriever is what is limiting answer quality." That shows you know the models and that you know a benchmark average is not a substitute for measuring your own corpus.

  • How would you actually decide, rather than argue from benchmarks?
    Build a small labelled evaluation set from your own domain — a few hundred queries with the chunk that should be retrieved — and measure recall@k and MRR for each candidate configuration. Embedding a few hundred queries twice costs almost nothing, and the result is specific to your corpus rather than to MTEB's mixture of tasks.
  • Is text-embedding-3-large at 1536 dimensions better than text-embedding-3-small at 1536?
    Generally yes on quality, since you keep the stronger model's representation and only shorten it, and the storage footprint is identical. You pay the large model's per-token rate at embed time. It is a common sweet spot precisely because it decouples the storage decision from the model decision — but it still needs verifying on your own evaluation set.
  • Where does text-embedding-ada-002 still fit?
    Only in already-built indexes. It is beaten by text-embedding-3-small on published quality while costing more per token, and it does not support the dimensions parameter. For new work it has no role; for existing systems the live question is whether re-embedding the corpus is worth the quality gain.

saying these in an interview costs you the question

  • Assumes more dimensions always means better retrieval
  • Ignores vector storage and index memory when comparing cost
  • Treats an MTEB average as a promise about their own corpus
  • Recommends text-embedding-ada-002 for new projects
  • Thinks you can mix small and large vectors in one index

context