skip to content

Embeddings API

Turning text into vectors for semantic search and RAG retrieval. Expect questions on choosing small vs large models, the dimensions parameter as a cost/accuracy dial, why cosine similarity is the default metric, and batching inputs to keep the bill down.

on this pageshow

questions

6

What does OpenAI's /v1/embeddings endpoint take as input and return?

level: juniorimportance: must knowfreq 72%

answer

  1. One call in, vectors out
  2. No roles, no streaming, no finish reason
  3. data[] entries carry index
  4. usage counts input tokens only
  5. input may be a string or an array

basics

~20 s

A POST to /v1/embeddings sends a model plus input (a string or an array of strings) and returns a data array of vectors, each carrying its embedding and the index of the input it came from, plus a usage object counting input tokens.

solid answer

~40 s

The embeddings endpoint is a single-shot transform, not a conversation. You send `model` (for example `text-embedding-3-small`) and `input`, which may be one string, an array of strings, or pre-tokenised integer arrays. Optional fields are `dimensions` (on the `text-embedding-3-*` models), `encoding_format` (`float`, the default, or `base64`), and `user`. The response is an object with `object: "list"`, a `data` array, the `model` actually used, and `usage`. Each `data` element has `object: "embedding"`, the `embedding` array of floats, and `index` — the position of the corresponding input. Always key results off `index` rather than assuming array order if you post-process asynchronously. `usage` reports `prompt_tokens` and `total_tokens` only; there are no output tokens, because nothing is generated. There is no `messages` array, no roles, no streaming, and no `finish_reason` here.

code

python · 11 lines
python
from openai import OpenAI

client = OpenAI()
resp = client.embeddings.create(
    model="text-embedding-3-small",
    input=["the cat sat on the mat", "a feline rested on a rug"],
)

for item in resp.data:
    print(item.index, len(item.embedding))
print(resp.model, resp.usage.prompt_tokens, resp.usage.total_tokens)

go deeper

for a junior

Be able to write the call from memory: model plus input, then read resp.data[0].embedding. Say plainly that nothing is generated — text goes in, a fixed-length array of numbers comes out.

for a middle

Explain the polymorphic input field, the index on each data element, and why usage only counts input tokens. Know that oversized inputs error rather than truncate.

for a senior

Show that you persist the model name and dimension alongside every vector, validate and filter chunks before sending, and join batch results on index so parallel workers cannot silently corrupt an index.

for a principal

Frame the endpoint as a build-time cost with a long-lived artefact: the vectors outlive the call, so the decisions that matter are what you store beside them and how you keep that store self-describing enough to migrate later.

## What this endpoint is for An embedding is a fixed-length array of floating-point numbers that represents a piece of text as a point in a high-dimensional space, positioned so that texts with similar meaning land near each other. OpenAI's Embeddings API is the HTTP surface that produces those arrays. You feed it text, it hands back numbers. Nothing is generated, nothing is sampled, and there is no notion of a conversation — this is a pure encoder call, which is why the request and response shapes look nothing like chat completions. The typical downstream use is semantic search or retrieval-augmented generation: you embed every chunk of your corpus once, store the vectors in a vector database, then at query time embed the user's question and retrieve the nearest stored vectors. ## Request shape A request to `POST /v1/embeddings` requires exactly two fields: - `model` — the embedding model identifier, e.g. `text-embedding-3-small`, `text-embedding-3-large`, or the legacy `text-embedding-ada-002`. - `input` — the text to embed. This is polymorphic: a single string, an array of strings, an array of token integers, or an array of token-integer arrays. Passing an array is how you batch. Optional fields: - `dimensions` — an integer requesting a shorter output vector. Supported only on the `text-embedding-3-*` models. - `encoding_format` — `"float"` (default) returns plain JSON numbers; `"base64"` returns the vector as a base64-encoded binary blob, which is materially smaller over the wire and faster to parse for large batches. - `user` — an opaque end-user identifier for abuse monitoring. Notably absent: there is no `temperature`, no `max_tokens`, no `stream`, no `tools`, no `response_format`. Those belong to the generative endpoints. An embedding call is deterministic in intent — the same input to the same model is meant to give you the same vector. ## Response shape ``` { "object": "list", "data": [ {"object": "embedding", "index": 0, "embedding": [0.0023, -0.009, ...]}, {"object": "embedding", "index": 1, "embedding": [ ... ]} ], "model": "text-embedding-3-small", "usage": {"prompt_tokens": 12, "total_tokens": 12} } ``` Three things matter here. First, `data` is a list parallel to your `input` list, and each element carries an explicit `index`. The API returns them in order, but the `index` field is the contract: if you fan requests out across threads or workers and reassemble, join on `index`, not on arrival order. Second, `model` echoes the model that served the request. Record it alongside the vectors you persist — a stored vector is meaningless without knowing which model produced it, because different models produce different, mutually incomparable spaces. Third, `usage` has only `prompt_tokens` and `total_tokens`, and for this endpoint they are equal. Embedding billing counts input tokens only; there is no completion side. That also makes cost estimation trivial: total corpus tokens times the model's per-token rate, one time, at index-build. ## The vectors themselves The returned floats are L2-normalised — the vector has unit length. That is why similarity is usually computed with cosine similarity or, equivalently for unit vectors, a plain dot product. The length of the array depends on the model and on any `dimensions` value you asked for: `text-embedding-3-small` returns 1536 floats by default, `text-embedding-3-large` returns 3072. ## Error cases a junior should recognise - An empty string input is rejected; filter empty chunks before sending. - An input longer than the model's maximum context (8192 tokens for the `text-embedding-3-*` models) returns a 400 error rather than being silently truncated. You must chunk long documents yourself. - Sending an array with more than 2048 elements is rejected; split the batch. ## SDK ergonomics In the Python SDK the call is `client.embeddings.create(model=..., input=...)`, returning a typed object where `resp.data[i].embedding` is a `list[float]` and `resp.usage.prompt_tokens` is an int. The JavaScript SDK mirrors this as `client.embeddings.create({ model, input })`. Both are thin wrappers over the same JSON; knowing the raw shape is what lets you debug when the SDK types and the wire disagree.

  • Why would you ever set encoding_format to base64?
    Because a 3072-float vector rendered as JSON decimal text is several times larger than its binary form. On large batches, base64 cuts response payload size and JSON parse time noticeably. You decode it back to a float array client-side. The vectors are identical either way; it is purely a transport optimisation.
  • The response echoes a model field. Why does that matter for what you persist?
    Because a vector is only interpretable relative to the model that produced it. Storing the model identifier (and the dimension count) alongside each vector is what lets you detect a mixed index, validate that a query embedding matches its corpus, and plan a re-embedding migration. Without it you have anonymous numbers.
  • What happens if one string in a batched input array is empty?
    The request fails rather than returning a placeholder vector for that element, so the whole batch is lost. In production you filter or replace empty and whitespace-only chunks before the call, and validate chunk length, because one bad element wastes the tokens of every good element in the same request.

saying these in an interview costs you the question

  • Thinks the endpoint takes a messages array with roles
  • Expects streaming or a finish_reason on embedding responses
  • Assumes temperature affects the returned vector
  • Believes usage reports completion tokens for embeddings
  • Reassembles batch results by arrival order instead of index

context

open as a page

What does the dimensions parameter do in OpenAI's Embeddings API?

level: middleimportance: must knowfreq 56%

basics

~20 s

Passing dimensions asks a text-embedding-3-small or text-embedding-3-large call to return a shorter vector — for example 256 instead of 3072. The models are trained so the leading coordinates carry the most information, so quality degrades gradually rather than collapsing.

open as a page

When would you pick text-embedding-3-large over text-embedding-3-small?

level: middleimportance: must knowfreq 68%

basics

~20 s

Pick text-embedding-3-large when retrieval quality is the bottleneck and the corpus is small enough that its higher per-token price and 3072-dimensional vectors are affordable. text-embedding-3-small is the sane default: far cheaper, 1536 dimensions, and close on benchmark quality.

open as a page

Why does OpenAI recommend cosine similarity for comparing its embeddings?

level: middleimportance: should knowfreq 50%

basics

~20 s

OpenAI embeddings are returned normalised to unit length, so cosine similarity reduces to a plain dot product and gives exactly the same ranking as Euclidean distance. Cosine is recommended because it is the cheapest and least error-prone of the equivalent choices.

open as a page

How would you embed a large document corpus through OpenAI's Embeddings API?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Chunk documents under the model's 8192-token input limit, pack many chunks per request as an array (up to 2048 entries), reassemble results by the index field, and for very large corpora submit the work through the asynchronous Batch API instead of live calls.

open as a page

How do you migrate a production vector index to a new OpenAI embedding model?

level: principalimportance: should knowfreq 36%

basics

~20 s

Vectors from different embedding models are not comparable, so the whole corpus must be re-embedded. Build the new index alongside the old one, keep queries on the old index until the new one is complete and evaluated, then cut over and only afterwards delete the old vectors.

open as a page