skip to content

In the Gemini API, what does CachedContent store and how do you call it?

level: juniorimportance: should knowfreq 45%

answer

  1. cache the prefix, not the answer
  2. one create, many generate calls
  3. handle looks like cachedContents/...
  4. cached_content on the generate config
  5. bound to one model, expires on TTL

basics

~10 s

CachedContent stores a large unchanging request prefix — system instruction, documents, uploaded files, tool declarations — server-side under a name like cachedContents/abc123. Later generateContent calls reference that name instead of resending those tokens.

solid answer

~40 s

Gemini's explicit context cache is created with `client.caches.create(model=..., config=CreateCachedContentConfig(contents=[...], system_instruction=..., ttl="3600s"))`. What you put in it is the **stable prefix** of your prompt: the long document or uploaded file, the system instruction, tool declarations — never the user's changing question. The call returns a `CachedContent` whose `name` looks like `cachedContents/abc123`. You then generate normally but pass that handle: `client.models.generate_content(model=..., contents="the question", config=GenerateContentConfig(cached_content=cache.name))`. The model behaves as if the cached content were prepended to your request. Caches are bound to the exact model they were created for, are immutable apart from their expiry, and disappear when the TTL runs out (default one hour) or when you call `client.caches.delete`. Check `response.usage_metadata.cached_content_token_count` to confirm the prefix really came from the cache.

code

python · 22 lines
python
from google import genai
from google.genai import types

client = genai.Client()
handbook = open("handbook.txt").read()  # tens of thousands of tokens

cache = client.caches.create(
    model="gemini-2.5-flash",
    config=types.CreateCachedContentConfig(
        display_name="handbook-v7",
        system_instruction="Answer only from the handbook.",
        contents=[handbook],
        ttl="3600s",
    ),
)

resp = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="What is the refund window?",
    config=types.GenerateContentConfig(cached_content=cache.name),
)
print(resp.usage_metadata.cached_content_token_count)

go deeper

for a junior

Know the shape: create a cache with contents plus a TTL, then pass its name in cached_content on generate calls. Say plainly that only the stable prefix goes in the cache.

for a middle

Explain the lifecycle — immutable contents, updatable TTL, default one-hour expiry, model binding — and show how you confirm a hit via cached_content_token_count in usage metadata.

for a senior

Demonstrate that you handle expiry and model upgrades in code: recreate-on-miss, expiry tracked alongside the handle, and cache-hit rate logged so a silent regression to full-price prompts is visible.

for a principal

Own the question of what the cache boundary is for the product — which prefix is genuinely shared, who pays for storage, and how document versioning maps onto immutable caches across tenants.

## The problem it solves A Gemini prompt over a long document is mostly the same bytes every turn. If you attach a 200-page manual and ask ten questions about it, a naive client uploads and pays for that manual ten times. Explicit context caching lets you register the unchanging part once and refer to it by name, so the repeated turns carry only the new question. This is a request-side mechanism, not an answer cache. It does not remember what the model said; it remembers the tokens you sent and (behind the scenes) the model's processed state for them. ## The three calls **Create.** `client.caches.create()` takes the target `model` and a `CreateCachedContentConfig`. Inside that config you can set `contents` (the parts you want cached — text, or references to files you uploaded through the Files API), `system_instruction`, `tools`, `display_name`, and `ttl` (a duration string such as `"3600s"`) or an absolute `expire_time`. The response is a `CachedContent` object; its `name` field, of the form `cachedContents/<id>`, is the handle you keep. **Use.** Every normal `client.models.generate_content()` call can carry `cached_content=<that name>` inside `GenerateContentConfig`. The `contents` you pass on that call are appended after the cached prefix. So the pattern is: big stable material in the cache, small variable turn in `contents`. **Manage.** `client.caches.get()`, `client.caches.list()`, `client.caches.update()` and `client.caches.delete()` round out the surface. `update` only accepts a new `ttl` or `expire_time` — the contents themselves are immutable. To change the document you create a new cache and delete the old one. ## What belongs inside, and what does not Inside: the corpus, the long system instruction that never varies, tool/function declarations, uploaded PDFs, audio or video file references. Outside: the user's question, retrieved snippets that differ per query, anything that changes between calls. A cache whose contents shift per request will never be reused and just wastes a round trip plus storage. There is also a **minimum size**. A prefix below the model's threshold cannot be cached at all — as of mid-2026 that is roughly 1,024 tokens for Gemini 2.5 Flash and 2,048 for 2.5 Pro. Trying to cache a two-sentence system prompt is a category error; that is what implicit caching is for. ## Model binding A cache belongs to one model. You create it for, say, `gemini-2.5-flash`, and every request that uses it must target that same model. Point the handle at a different model — or migrate to a new model generation — and the request is rejected. This matters operationally: a model upgrade invalidates every cache you hold, so your code needs a create-on-miss path rather than a hard-coded name. ## Lifetime Caches expire. If you set no `ttl`, the default is one hour. When the TTL elapses the cache is gone; a later request carrying the stale `cachedContents/...` name fails rather than silently falling back to full-price input. Production code therefore wraps generation so that a failed lookup triggers a fresh `caches.create` and a retry, and it stores the handle with its own expiry timestamp so it can pre-emptively recreate. ## Verifying it worked The response's `usage_metadata` carries `cached_content_token_count` alongside `prompt_token_count` and `candidates_token_count`. If the cached count is zero on a request you believed was a hit, either the handle was ignored, the cache expired, or you never passed `cached_content` at all. Logging that field per request is the cheapest possible cache-hit monitor. ## Common mistakes People assume the cache is free once created — it is not, storage is billed for as long as the cache lives. They assume it caches responses, so identical questions get identical answers without inference — it does not; the model still runs. They put the question in the cache and the document in `contents`, which is exactly backwards. And they treat the handle as permanent, which breaks the first time a TTL lapses or a model version rolls forward.

  • What happens to a request that passes a cache handle whose TTL has already elapsed?
    It fails rather than degrading gracefully — the handle no longer resolves, so you get an error instead of a full-price uncached run. Production code should catch that, call `caches.create` again with the same contents, store the new name, and retry the generation. Tracking each cache's expiry locally lets you recreate before the failure instead of after it.
  • Can you edit the contents of an existing CachedContent?
    No. Contents, system instruction and tools are immutable once the cache exists; `client.caches.update()` accepts only a new `ttl` or `expire_time`. To change the document you create a second cache and delete the first. That immutability is why cache keys in real systems include a document version — a new version simply means a new cache.
  • How do you know whether a prefix is even large enough to cache explicitly?
    Count it before creating anything: `client.models.count_tokens(model=..., contents=...)` gives the token total for the exact material you plan to cache. Compare it against the model's minimum cacheable size — around 1,024 tokens for Gemini 2.5 Flash and 2,048 for 2.5 Pro as of mid-2026. Below that, rely on implicit caching instead.

It is like uploading an attachment to a shared drive once and then linking to it in every email, instead of re-attaching the same 200-page PDF to each message.

saying these in an interview costs you the question

  • Thinks the cache stores model responses rather than input tokens
  • Expects a cache created for one model to work on another
  • Puts the changing user question inside the cached content
  • Assumes a cache costs nothing once created
  • Treats the cache handle as permanent, with no recreate path

context