How does output_dimensionality work for gemini-embedding-001 vectors?
answer
- Shorter vector, same model
- Prefix, not a distilled variant
- Nested dolls of decreasing size
- Only full length arrives normalised
- Saves storage, not API spend
basics
~20 soutput_dimensionality truncates the model's full-length embedding to a shorter prefix instead of running a smaller model. The model is trained so early dimensions carry most of the signal, and truncated vectors should be re-normalised to unit length before storage.
solid answer
~50 s`gemini-embedding-001` returns 3072 floats by default. Setting `output_dimensionality` in the embed request asks for a shorter vector — 1536 or 768 are the usual choices — and the API returns a prefix of the full vector rather than switching to a different model. That works because the model is trained with a Matryoshka-style objective: the leading dimensions are made to carry most of the semantic signal on their own, so a truncated vector degrades gracefully instead of losing a random quarter of the meaning. One catch: only the full-length output is unit-normalised, so after truncating you should re-normalise the vector yourself before writing it to the index, otherwise cosine similarity and dot-product scores drift. The payoff is storage and search cost, which scale with dimension count — not API price, which is billed on input tokens.
code
python · 18 linesimport numpy as np
from google import genai
from google.genai import types
client = genai.Client(api_key="YOUR_KEY")
resp = client.models.embed_content(
model="gemini-embedding-001",
contents="Quarterly credential rotation runbook.",
config=types.EmbedContentConfig(
task_type="RETRIEVAL_DOCUMENT",
output_dimensionality=768,
),
)
v = np.array(resp.embeddings[0].values, dtype=np.float32)
print(v.shape, float(np.linalg.norm(v))) # 768, and not 1.0
v = v / np.linalg.norm(v) # re-normalise before storinggo deeper
Know that Gemini embeddings can be returned shorter than the default by setting an output dimensionality, and that the shorter vector is a prefix of the longer one rather than a different model's output.
Explain the Matryoshka training that makes prefixes usable, name the default 3072 for gemini-embedding-001, and state that truncated vectors must be re-normalised before storage.
Show that you would measure recall@k at each candidate dimension, treat dimension count as part of index identity, and know that the savings land in storage, RAM and latency rather than in the API bill.
Own the capacity argument: quantify index RAM and storage at corpus scale, set the quality bar you are willing to trade away, and make dimension a governed schema decision with metadata assertions so a deploy cannot silently split the index.
## The parameter Gemini's embedding request accepts `output_dimensionality` (`outputDimensionality` in REST JSON, a field on `EmbedContentConfig` in the google-genai SDK). As of mid-2026 the current model, `gemini-embedding-001`, emits 3072 floats per input by default; the parameter lets you ask for fewer. The important mental model is that this is **truncation, not a different model and not a projection computed over your data**. You are not getting a distilled 768-dimension variant, and there is no PCA over your corpus. You are getting the first N components of the same vector the model would have produced anyway. ## Why truncation works: Matryoshka representations With an ordinary embedding model, chopping off dimensions is vandalism — every coordinate carries roughly equal weight, so dropping three quarters of them destroys most of the geometry. Matryoshka Representation Learning trains the model so that nested prefixes of the vector are each independently useful: the first 768 numbers are optimised to be a good 768-dimension embedding, the first 1536 a better one, and the full 3072 the best. Like nested dolls, each prefix is a complete, smaller version of the same thing. That is what makes `output_dimensionality` safe. The quality curve is not flat — 768 really is somewhat worse than 3072 on hard retrieval — but it degrades gently and predictably rather than falling off a cliff. ## The normalisation catch This is the detail interviews probe, and the one teams miss. The model's **full-length** output is unit-normalised (length 1). A prefix of a unit vector is not itself a unit vector — its norm is less than 1, and by a different amount for every input. Consequences if you skip re-normalisation: - **Dot product stops equalling cosine similarity.** Many vector stores use inner product as the metric precisely because normalised inputs make the two identical. With ragged norms, dot-product scores now mix "how similar" with "how long the vector happened to be", and ranking degrades. - **Distance thresholds become meaningless.** Any tuned cutoff ("treat scores above 0.82 as a match") assumes a fixed scale. - **Mixed batches misbehave.** If some rows were normalised and others were not, the unnormalised ones are systematically mis-ranked. The fix is one line: divide the returned vector by its L2 norm before storing it, and do exactly the same to query vectors. Cosine similarity implementations that normalise internally are unaffected, but you rarely control that end to end, so normalise at the boundary. ## What you actually buy Dimension count drives everything downstream of the API call: - **Storage.** 3072 float32 values is ~12 KB per vector; 768 is ~3 KB. Across ten million chunks that is 120 GB versus 30 GB, before index overhead. - **Memory-resident index size,** which is often the real constraint — an HNSW graph over 3072-dim vectors needs four times the RAM per node's vector payload. - **Query latency,** since distance computation is linear in dimensions. - **Not** the API bill. Embedding calls are priced on input tokens; asking for 768 outputs instead of 3072 does not make the call cheaper. So the decision is a retrieval-quality-versus-infrastructure-cost trade, and it should be made by measuring recall@k on your own evaluation set at each candidate dimension, not by picking the default because it is the default. ## Index identity, again Dimension count is fixed for the lifetime of an index, exactly like the model name and the task type. Most vector stores declare the dimension in the collection schema and will reject a mismatched insert outright — which is the good failure. The bad case is a store that accepts anything: then a deploy that quietly changes `output_dimensionality` splits your corpus into two incomparable halves. Record the dimension alongside the model name in index metadata and assert on it at write time. A subtle trap: truncating an *existing stored* 3072-dim vector to 768 by hand and comparing it against fresh 768-dim API output is *almost* right — the prefixes match — but only if you re-normalise both consistently. Doing this to avoid a backfill is a false economy; run the backfill. ## Rule of thumb Start at the default; drop to 1536 or 768 when index size, RAM or latency is a measured problem; validate with recall@k before and after; always re-normalise anything you truncated.
- Why does skipping re-normalisation break a store that uses inner product as its metric?Inner product equals cosine similarity only when both vectors have length 1. A truncated prefix has a norm below 1 that varies per input, so the dot product now blends semantic similarity with incidental vector length. Rankings shift and any tuned score threshold stops meaning what it used to. Normalising both stored and query vectors at the boundary restores the equivalence.
- Does asking for 768 dimensions instead of 3072 reduce your Gemini API bill?No. Embedding requests are priced on input tokens, so the same text costs the same regardless of how many floats come back. The savings are entirely downstream: disk, index RAM, and distance-computation time during search. Choose the dimension against those constraints and against measured recall, not against the invoice.
- Can you truncate already-stored 3072-dimension vectors instead of re-embedding at 768?Mathematically the prefix is the same numbers, so it can work — but only if you re-normalise every truncated vector and the query side identically, and only if nothing else in the pipeline assumed the old norms. It is cheaper than a backfill but easy to get subtly wrong; for a production index, re-embed into a fresh collection and cut over.
- How would you decide between 3072 and 768 for a new corpus?Build an evaluation set of real queries with known-relevant chunks, embed the corpus at both dimensions, and compare recall@k and MRR. Then weigh the quality delta against measured index RAM, storage and p99 search latency at your corpus size. If quality loss is within noise, take the smaller vector; if retrieval is the product, keep the full length.
Matryoshka embeddings are nested dolls: the small doll inside is a complete, lower-detail version of the big one, so opening it up and keeping only the inner doll still leaves you with something recognisable.
saying these in an interview costs you the question
- Believing a smaller dimension runs a lighter, cheaper model
- Assuming truncated vectors stay unit length
- Thinking output_dimensionality lowers the per-call API price
- Expecting PCA-style reduction computed over your corpus
- Changing dimension count without re-embedding the index