skip to content

Embeddings

embedContent, distinguished by a task_type that tells the model whether the text is a stored document, a search query, or a classification input — the same string embeds differently for different jobs. Questions cover matching task types between index and query time, and trimming output dimensionality.

on this pageshow

questions

5

What does task_type change in a Gemini embedContent request?

level: middleimportance: must knowfreq 70%

answer

  1. Same text, different job, different vector
  2. Declared per request, not per model
  3. Asymmetric pair for search
  4. Document side and query side must match
  5. RETRIEVAL_DOCUMENT indexes, RETRIEVAL_QUERY searches

basics

~20 s

task_type tells Gemini's embedding model what the text is for — a stored passage, a search query, a classification input — and the model places the same text differently in vector space for each. Index and query sides must use the matching pair.

solid answer

~50 s

Gemini's `embedContent` endpoint takes an optional `task_type` (`taskType` over REST) next to the text. The response shape never changes — you still get one float vector — but the model applies a different training objective, so the same sentence embeds to different coordinates depending on the declared job. The values include `RETRIEVAL_DOCUMENT`, `RETRIEVAL_QUERY`, `SEMANTIC_SIMILARITY`, `CLASSIFICATION`, `CLUSTERING`, `QUESTION_ANSWERING`, `FACT_VERIFICATION` and `CODE_RETRIEVAL_QUERY`. The retrieval pair is asymmetric on purpose: a short question and the long passage that answers it do not look alike, so the document side and the query side are projected to meet. The operational rule is that whatever you used when writing vectors into the index, you must use its counterpart at search time — embed chunks with `RETRIEVAL_DOCUMENT`, embed the user's query with `RETRIEVAL_QUERY`. Mixing them produces no error, just quietly worse recall.

code

python · 18 lines
python
from google import genai
from google.genai import types

client = genai.Client(api_key="YOUR_KEY")

doc = client.models.embed_content(
    model="gemini-embedding-001",
    contents="Rotate the database password every quarter using the runbook.",
    config=types.EmbedContentConfig(task_type="RETRIEVAL_DOCUMENT"),
)

query = client.models.embed_content(
    model="gemini-embedding-001",
    contents="how often should I change the db password?",
    config=types.EmbedContentConfig(task_type="RETRIEVAL_QUERY"),
)

print(len(doc.embeddings[0].values), len(query.embeddings[0].values))

go deeper

for a junior

Know that Gemini embeddings take a task_type, and that stored passages use RETRIEVAL_DOCUMENT while search strings use RETRIEVAL_QUERY. Say plainly that the same sentence embeds differently depending on the value.

for a middle

Explain that task_type selects a training objective at inference, that the retrieval pair is deliberately asymmetric so short queries land near long answers, and name several of the real values.

for a senior

Show the operational judgment: task_type is part of the index's identity, a mismatch is silent rather than an error, and changing it means re-embedding the corpus. Mention centralising the query-side value in one helper.

for a principal

Own the tradeoff framing — the correct pair is a free recall win because pricing follows input tokens, and index identity (model, dimensions, task type) should be stored as metadata so drift after a deploy is detectable instead of mysterious.

## What the parameter is The Gemini API embeds text through `models/<model>:embedContent` (and `:batchEmbedContents` for many texts in one HTTP call). Beyond the model name and the content, the request carries an optional `task_type` — `taskType` in REST JSON, `task_type` on the `EmbedContentConfig` object in the google-genai SDK. It does not alter the response shape. You get back an embedding whose length is the model's dimensionality (3072 by default for `gemini-embedding-001`, the current model as of mid-2026, which superseded `text-embedding-004`). What changes is *where in that space the text lands*. The model was trained against several objectives at once, and `task_type` selects which one is applied at inference time. Feed the identical string twice with two different task types and you get two different vectors. ## The values and what each is for - `RETRIEVAL_DOCUMENT` — text you are storing and will later search over: a chunk, a passage, a knowledge-base article. - `RETRIEVAL_QUERY` — the user's search string, matched against vectors stored as documents. - `SEMANTIC_SIMILARITY` — symmetric comparison of two texts of the same kind, e.g. near-duplicate detection or sentence-pair scoring. - `CLASSIFICATION` — vectors that will be fed to a downstream classifier as features. - `CLUSTERING` — vectors intended for k-means or similar grouping. - `QUESTION_ANSWERING` and `FACT_VERIFICATION` — retrieval-shaped variants tuned for finding an answer to a question or evidence for a claim. - `CODE_RETRIEVAL_QUERY` — a natural-language query searching over code. ## Why asymmetry exists This is the concept behind the whole parameter. A search query ("how do I rotate the database password") and the passage that answers it (three paragraphs of runbook prose that may never contain the word "rotate") are not similar texts. A purely symmetric embedding model is optimised so that things which *look alike* land near each other — which is the wrong objective for search, where a short interrogative fragment must land near a long declarative answer. The `RETRIEVAL_DOCUMENT` / `RETRIEVAL_QUERY` pair is trained as exactly that: two projections that are designed to meet. The document projection is what you use when writing to the index; the query projection is what you use when reading. Neither is useful alone; the pairing is the point. `SEMANTIC_SIMILARITY`, by contrast, is genuinely symmetric — you use it on both sides when comparing like with like. ## The rule that matters in production **task_type is part of your index's identity, alongside the model name and the dimension count.** Whatever you embedded a corpus with, you are locked into its counterpart at query time until you re-embed. Consequences: 1. Embed every stored chunk with `RETRIEVAL_DOCUMENT` and every incoming query with `RETRIEVAL_QUERY`. Using `RETRIEVAL_DOCUMENT` on both sides is the single most common bug, and it is silent — retrieval still returns ten results, they are just measurably worse. 2. If you decide later that your workload is symmetric and switch to `SEMANTIC_SIMILARITY`, you must re-embed the whole corpus. You cannot mix task types in one index. 3. A/B a task-type change on offline recall metrics, not by eyeballing a handful of queries; the effect is a shift in ranking quality, not a visible failure. ## The title field With `RETRIEVAL_DOCUMENT` you may also supply an optional `title` for the chunk. It is folded into the document representation, so a chunk from "Payments runbook — refunds" is embedded with that framing rather than as an anonymous paragraph. The field is only meaningful for the document task type; there is no equivalent on the query side. ## If you omit it Omitting `task_type` is legal — the API applies a general-purpose default rather than rejecting the call. That is precisely why teams ship RAG pipelines that never set it and only discover the parameter when someone asks why recall is mediocre. Setting the correct pair is one of the cheapest retrieval-quality wins available on this API, because it costs nothing extra: the request price is driven by input tokens, not by which objective you selected. ## Checklist - Same task type at write time for every document in one index. - The paired query task type at read time, applied in exactly one place in the code so the two cannot drift. - Store the model name, dimension count and task type as index metadata, so a mismatch after a deploy is detectable rather than mysterious.

  • What happens if you embed both the corpus and the queries with RETRIEVAL_DOCUMENT?
    Nothing visibly breaks — the API returns normal vectors of the right length and search still returns results. You lose the asymmetric projection that was trained to put short questions near long answers, so ranking quality degrades measurably while every request looks healthy. It is a silent recall bug, which is why the query-side task type belongs in exactly one shared helper function rather than being set at each call site.
  • You decide to change the task_type your index was built with. What does that cost?
    A full re-embed of the corpus. Vectors produced under different task types are not comparable, so you cannot mix old and new rows in one index. Treat task type like the model name: part of the index's identity, stored as metadata, changed only via a backfill into a fresh index with a cutover.
  • When would you use SEMANTIC_SIMILARITY rather than the retrieval pair?
    When both sides of the comparison are the same kind of text — deduplicating support tickets, clustering headlines, scoring paraphrase pairs. There is no query/document asymmetry to exploit there, and the symmetric objective is a better fit. For a search index where short queries hit long passages, the retrieval pair wins.
  • What does the optional title field do on an embedding request?
    It is only meaningful with RETRIEVAL_DOCUMENT, where it supplies the chunk's document title so the passage is embedded with that framing instead of as an anonymous paragraph. It typically helps when your chunks are small fragments of larger titled documents. There is no counterpart on the query side.

It is like filing a document versus writing the search slip you take to the archivist: the same words get organised differently depending on whether they are being shelved or used to find something on a shelf.

saying these in an interview costs you the question

  • Thinking task_type only tags the request as metadata
  • Using RETRIEVAL_DOCUMENT for queries as well as chunks
  • Believing task_type changes the vector's length
  • Assuming a task_type mismatch raises an API error
  • Mixing task types within a single vector index

context

open as a page

How do you embed text with the Google GenAI SDK's embed_content method?

level: juniorimportance: should knowfreq 58%

basics

~10 s

Create a genai.Client, call client.models.embed_content with a model name and contents, and optionally pass an EmbedContentConfig for task type and dimensionality. Read the floats from response.embeddings[i].values, one entry per input text.

open as a page

How does output_dimensionality work for gemini-embedding-001 vectors?

level: middleimportance: should knowfreq 50%

basics

~20 s

output_dimensionality truncates the model's full-length embedding to a shorter prefix instead of running a smaller model. The model is trained so early dimensions carry most of the signal, and truncated vectors should be re-normalised to unit length before storage.

open as a page

How do you embed a large document corpus with Gemini's embedding API?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Chunk documents to fit the model's input token limit, send many chunks per request via the batch embedding endpoint, run bounded concurrency with backoff on rate-limit errors, and make the job resumable by chunk ID so an interruption does not restart it.

open as a page

Your Gemini-embedded RAG index needs a new embedding model. What breaks?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Vectors from two different embedding models are not comparable, so the entire corpus must be re-embedded — there is no conversion. Plan a backfill into a second index and a cutover, not an in-place, gradual replacement.

open as a page