skip to content

What does the dimensions parameter do in OpenAI's Embeddings API?

level: middleimportance: must knowfreq 56%

answer

  1. A width dial, not a price dial
  2. Only on the text-embedding-3 models
  3. Front-loaded information, nested prefixes
  4. Storage and latency shrink, tokens do not
  5. Mixed widths cannot share an index

basics

~20 s

Passing dimensions asks a text-embedding-3-small or text-embedding-3-large call to return a shorter vector — for example 256 instead of 3072. The models are trained so the leading coordinates carry the most information, so quality degrades gradually rather than collapsing.

solid answer

~50 s

`dimensions` is an integer on the `text-embedding-3-small` and `text-embedding-3-large` requests that truncates the returned vector to that width. The legacy `text-embedding-ada-002` does not support it. This works because those models are trained with a Matryoshka-style objective: information is packed front-loaded, so the first N coordinates are themselves a usable embedding rather than an arbitrary slice. OpenAI reports that `text-embedding-3-large` cut to 256 dimensions still beats `ada-002` at its full 1536. In practice it is a storage and latency dial. Halving the width halves vector-store bytes and roughly halves distance-computation cost, at some measurable recall cost you should quantify on your own evaluation set. Two hard constraints: it does not reduce the token bill, since you still pay for the input text; and every vector in one index must share the same model and the same width, so this is a decision you make before building the index, not after.

code

python · 11 lines
python
from openai import OpenAI

client = OpenAI()
for width in (3072, 1024, 256):
    resp = client.embeddings.create(
        model="text-embedding-3-large",
        input="quarterly revenue recognition policy",
        dimensions=width,
    )
    vec = resp.data[0].embedding
    print(width, len(vec), resp.usage.prompt_tokens)

go deeper

for a junior

Recall that dimensions makes the returned vector shorter, that it works on text-embedding-3-small and text-embedding-3-large, and that shorter vectors take less storage.

for a middle

Explain why a prefix is still a valid embedding — the models front-load information into the earliest coordinates — and that the parameter cuts storage and search cost but not token cost.

for a senior

Show you pick the width by sweeping a labelled evaluation set against storage, and that you record model and width with every vector because the width is baked into the index.

for a principal

Treat width as a one-way architectural commitment tying together storage budget, query latency and future migration cost, and set the policy that no index accepts vectors whose model and width are not recorded.

## What the parameter is `dimensions` is an optional integer on the embeddings request: ``` {"model": "text-embedding-3-large", "input": "...", "dimensions": 1024} ``` The response comes back with 1024 floats instead of the model's native 3072. It is supported on `text-embedding-3-small` (native 1536) and `text-embedding-3-large` (native 3072) only. The legacy `text-embedding-ada-002` predates the feature and rejects it. ## Why a shortened vector still works Naively, truncating a vector should destroy it — you are throwing away coordinates that carry meaning. That is true for embeddings trained conventionally, where information is spread more or less evenly across the dimensions. The `text-embedding-3-*` models are trained differently, with an objective in the Matryoshka Representation Learning family: the loss is computed not only on the full vector but simultaneously on nested prefixes of it. The effect is that the model is forced to put the most discriminative information into the earliest coordinates, with later coordinates adding progressively finer detail — the same way the first few principal components of a dataset carry most of its variance. A 256-dimension prefix is therefore not a mutilated 3072-dimension vector; it is a genuine, coarser embedding of the same text. The headline evidence OpenAI published for this: `text-embedding-3-large` truncated to 256 dimensions still scores higher on MTEB than `text-embedding-ada-002` at its full 1536. That is a twelvefold storage reduction against the old model with a quality gain, which is why the parameter got attention. ## What it buys you **Storage.** Vectors are the dominant cost of a large index. At float32, one 3072-dimension vector is 12,288 bytes; at 1024 dimensions it is 4,096. Over ten million chunks that is roughly 120 GB versus 40 GB, before index structures. **Query latency and CPU.** Approximate nearest-neighbour search cost scales with vector width, both in the distance computations and in memory bandwidth. Narrower vectors mean more of the index fits in cache and RAM. **Vector-store compatibility.** Some indexes are provisioned at a fixed width, or you may want to swap in the large model without re-provisioning an index built at 1536. Setting `dimensions: 1536` on `text-embedding-3-large` gives you a drop-in upgrade at identical storage. ## What it does not buy you **It does not reduce your API bill.** Billing counts input tokens. The model still reads and processes the whole text; you simply receive fewer numbers back. A team that sets `dimensions: 256` expecting a cheaper invoice has misread the parameter. **It does not make vectors interchangeable.** A 256-dimension vector and a 1536-dimension vector from the same model cannot be compared — cosine similarity is undefined between vectors of different length, and there is no padding rule that would make it meaningful. Every vector in an index, including the query embedding computed at request time, must share both the model and the width. Store both as metadata. ## Truncating client-side versus using the parameter You can also slice the array yourself after the fact. If you do, you must re-normalise: the model's output is unit length across all its coordinates, and a prefix of a unit vector is not itself unit length. Divide the prefix by its own L2 norm before storing it, or every downstream cosine computation is subtly wrong — and worse, wrong by an amount that varies per vector, which distorts rankings rather than merely rescaling them. Using the `dimensions` parameter is the simpler path precisely because it removes that footgun from your code. ## Choosing a width There is no universal right answer, and the parameter is not free quality. The honest method is a sweep: embed a labelled evaluation set at several widths — say 3072, 1536, 1024, 512, 256 — and plot recall@k against storage. Corpora differ sharply in where the curve knees. A small corpus of distinct documents often loses nothing at 512; a large corpus of near-duplicate technical chunks may degrade visibly below 1536, because fine distinctions are exactly what the trailing coordinates encode. ## The decision is a one-way door Because the width is baked into every stored vector, changing it later means re-embedding the entire corpus and rebuilding the index. Treat the width choice with the same seriousness as the model choice: measure first, and record the chosen model and width alongside the data so a future migration is a known quantity rather than an archaeology project.

  • Does setting dimensions reduce what you are billed?
    No. Billing counts input tokens, and the model reads the entire input regardless of how many numbers it returns. What shrinks is response payload, vector-store bytes and distance-computation cost. If cost is the goal, the levers are the model choice, chunking strategy, and avoiding re-embedding unchanged content — not the dimensions parameter.
  • If you slice the vector yourself instead of using the parameter, what must you remember?
    Re-normalise. The model returns a unit-length vector, but a prefix of a unit vector generally is not unit length, and the shortfall differs per vector. Skipping the L2 re-normalisation makes cosine similarity inconsistent across your corpus and distorts rankings rather than merely rescaling scores. Using the API parameter avoids the issue entirely.
  • How would you pick a width for a new RAG index?
    Sweep it. Embed a labelled query/chunk evaluation set at 3072, 1536, 1024, 512 and 256, then plot recall@k against storage per million chunks and pick the knee. Corpora differ: distinct short documents often hold up at 512, while large corpora of near-duplicate technical text degrade well before that.

Think of it like a Russian nesting doll: each smaller doll is a complete doll, not a broken piece of the big one. The model is trained so the first 256 numbers already stand on their own as a coarser portrait of the text.

saying these in an interview costs you the question

  • Thinks dimensions lowers the token bill
  • Believes truncation is arbitrary and destroys the vector
  • Assumes ada-002 also accepts the dimensions parameter
  • Compares vectors of different widths in the same index
  • Slices vectors client-side without re-normalising

context