skip to content

How do you embed text with the Google GenAI SDK's embed_content method?

level: juniorimportance: should knowfreq 58%

answer

  1. Same client object as generation
  2. models service, embed_content method
  3. Config object carries task type
  4. One entry per input, in order
  5. No streaming, no candidates

basics

~10 s

Create a genai.Client, call client.models.embed_content with a model name and contents, and optionally pass an EmbedContentConfig for task type and dimensionality. Read the floats from response.embeddings[i].values, one entry per input text.

solid answer

~40 s

In the google-genai Python SDK you build a client once — `genai.Client(api_key=...)`, or with no key argument when the API key is in the environment — and then call `client.models.embed_content(model=..., contents=..., config=...)`. `contents` takes a single string or a list of strings; passing a list embeds them in one round trip, which maps onto the REST `:batchEmbedContents` endpoint rather than `:embedContent`. Options that shape the vector go on `types.EmbedContentConfig` — notably `task_type` and `output_dimensionality`. The result is an `EmbedContentResponse` whose `embeddings` list has one `ContentEmbedding` per input, in input order, each exposing the floats as `.values`. Note this is a different shape from generation calls: there are no candidates, no finish reason, and no streaming — embedding is a single synchronous request in, fixed-length vectors out.

code

python · 16 lines
python
from google import genai
from google.genai import types

client = genai.Client(api_key="YOUR_KEY")

resp = client.models.embed_content(
    model="gemini-embedding-001",
    contents=[
        "Refunds are issued to the original payment method.",
        "Password rotation happens every quarter.",
    ],
    config=types.EmbedContentConfig(task_type="RETRIEVAL_DOCUMENT"),
)

for emb in resp.embeddings:
    print(len(emb.values))

go deeper

for a junior

Be able to write the call from memory: build a client, call models.embed_content with a model and contents, and read the floats from response.embeddings[0].values.

for a middle

Explain that config is a separate EmbedContentConfig carrying task type and dimensionality, that a list of contents becomes one batched round trip, and that results are positional.

for a senior

Demonstrate pipeline discipline: one shared client, model and task type from configuration, immediate zipping of inputs to vectors so records never desynchronise, and chunking before the model's input limit.

for a principal

Frame it as contract management — model, task type and dimension define index identity, so those settings belong in governed configuration with metadata assertions rather than being an implementation detail of an ingestion script.

## The call The google-genai SDK exposes embeddings on the `models` service, alongside generation: - `client.models.generate_content(...)` produces text. - `client.models.embed_content(...)` produces vectors. The naming parallel is deliberate, but the response shapes have nothing in common. An embedding call has no candidates, no finish reason, no safety ratings and no streaming mode — you send text, you get numbers, once. ## Client construction `client = genai.Client(api_key="...")` creates the client. The key can also come from the environment so the literal never appears in source, which is the form you want in any real service. One client is enough for the process; it holds the HTTP configuration, so constructing one per request wastes connections. The same client object reaches every model service, so there is no separate "embeddings client" to build. ## Arguments **`model`** — the embedding model's name, e.g. `"gemini-embedding-001"`, the current model as of mid-2026 (it superseded `text-embedding-004`). A generation model name here is an error; embedding models and generation models are disjoint sets. **`contents`** — the text. Pass a single string for one vector, or a list of strings to embed several in one HTTP round trip. The list form is what you want for a corpus: the per-request overhead is amortised across the batch, and results come back in input order so you can zip them against your source records. **`config`** — a `types.EmbedContentConfig`. This is where `task_type` (what job the text is for) and `output_dimensionality` (how many floats to return) live. Both are optional; omitting the config gives a general-purpose, full-length vector. ## The response `embed_content` returns an `EmbedContentResponse`. The field you want is `embeddings`: a list of `ContentEmbedding` objects, one per input, each with a `values` attribute holding the floats. So for a single string you read `response.embeddings[0].values`; for a batch you iterate. Two shapes people confuse it with: - `response.data[0].embedding` — that is the OpenAI SDK's layout, not this one. - `response.candidates[0].content` — that is Gemini's *generation* response. Getting this wrong is the most common first-attempt error, and it fails loudly with an attribute error, which is the merciful case. ## Order and pairing Because results are positional, the safest pattern is to zip the input list against `response.embeddings` immediately and carry the record ID with the vector from that moment on. Any code path that re-sorts, filters or retries a partial batch between call and zip is a chance to associate the wrong vector with the wrong document — a bug that produces plausible-looking nonsense at search time and is miserable to trace. ## What is not here - **No streaming.** There is no incremental delivery of a vector; the request either returns the whole embedding or fails. - **No system instruction, no temperature, no tools.** Those are generation concepts; the embedding call has no sampling behaviour at all, and the same text with the same config yields the same vector. - **No conversation.** Each call is stateless and independent. ## Errors and limits The usual API failure modes apply: an invalid key gives an authentication error, an unknown model name gives a not-found error, exceeding your quota gives a rate-limit response you should retry with exponential backoff, and text longer than the model's input token limit must be chunked by you before the call. None of these are embedding-specific, but the last one bites hardest, because a document that is merely long is a perfectly ordinary input to a generation model and an over-limit input here. ## Minimal end-to-end shape Build the client once at start-up; write one helper that takes a list of strings plus a task type and returns a list of vectors; keep `model`, `task_type` and `output_dimensionality` in configuration rather than scattered at call sites, so the whole pipeline cannot drift out of agreement with the index it writes to.

  • What is the difference between passing a single string and a list of strings to contents?
    A single string yields one embedding; a list embeds every element in one HTTP round trip, mapping onto the batch embedding endpoint. Results come back in input order, one per element. The list form is what you use for corpus ingestion, since it amortises per-request overhead — but per-request batch limits apply, so a large corpus still has to be chunked into several calls.
  • Why is there no streaming mode for embed_content?
    An embedding is a single fixed-length vector computed from the whole input, not a token sequence generated left to right. There is nothing partial to deliver. That is also why there is no temperature or sampling here: the call is deterministic in a way generation is not, and the same text with the same config returns the same numbers.
  • What is the most likely cause of an AttributeError when reading the result?
    Reaching for the wrong response shape. Gemini generation responses expose candidates, and the OpenAI SDK exposes data[0].embedding — neither applies. The google-genai embedding response puts the vectors on .embeddings, a list of objects each carrying .values.
  • Where should model name and task type live in a real ingestion service?
    In configuration, read in one place, not repeated at call sites. The model, the task type and the output dimensionality together define what an index's vectors mean; if two code paths disagree after a refactor you get an index containing incomparable vectors with no error anywhere. One helper function, one config source, and the values recorded as index metadata.

saying these in an interview costs you the question

  • Reading the vector from response.data[0].embedding
  • Expecting candidates or a finish reason on the response
  • Building a new client for every request
  • Passing a generation model name to the embedding call
  • Assuming embeddings can be streamed incrementally

context