How is Gemini context caching billed, and what does the cache TTL cost?
answer
- two meters, not one
- discounted input plus hourly storage
- TTL is a bill, not a setting
- break-even is reuses per hour
- delete the cache when the session ends
basics
~20 sGemini charges cached input tokens at a reduced per-token rate on each request, plus a separate storage fee for holding the cache — priced per million cached tokens per hour. A longer TTL means a bigger storage bill, whether or not anyone reads the cache.
solid answer
~40 sThere are two meters. First, on every `generateContent` that carries `cached_content`, the tokens served from the cache are billed as input but at a **discounted** rate versus fresh input tokens; the new tokens in that turn are billed normally. Second, the cache itself carries a **storage charge proportional to cached tokens × time alive**, quoted per million tokens per hour. That second meter is what makes TTL an economic decision rather than a convenience setting: a one-hour TTL over a 300k-token document costs the same whether you ask one question or two hundred. Caching therefore pays off exactly when reuse count during the TTL is high enough that the per-request savings exceed storage. Verify with `usage_metadata.cached_content_token_count`, and extend or shorten a live cache with `client.caches.update(name=..., config=UpdateCachedContentConfig(ttl=...))` rather than guessing at creation time.
code
python · 19 linesfrom google.genai import types
# Extend a live cache only while the session is active
cache = client.caches.update(
name=cache.name,
config=types.UpdateCachedContentConfig(ttl="600s"),
)
resp = client.models.generate_content(
model="gemini-2.5-flash",
contents="Summarise clause 12.",
config=types.GenerateContentConfig(cached_content=cache.name),
)
u = resp.usage_metadata
fresh = u.prompt_token_count - u.cached_content_token_count
print("cached:", u.cached_content_token_count, "fresh:", fresh)
client.caches.delete(name=cache.name) # stop the storage metergo deeper
Remember that caching gives a discount, not free tokens, and that a cache costs money for as long as its TTL keeps it alive.
Explain both meters and derive the break-even: caching wins when the prefix is re-read enough times per hour of storage to outweigh the storage fee.
Show the operational discipline — TTL matched to session length, sliding extension on activity, explicit delete on completion, and cached-versus-fresh token counts logged per request.
Own the spend model: decide which prefixes are worth caching at all across the product, who owns the storage budget, and how the long-prompt price tier changes the build-versus-retrieve calculus.
## Two meters, not one Engineers usually reach for context caching expecting a single discount. Gemini's explicit cache actually bills on two independent axes, and confusing them is the classic interview slip. **Meter 1 — cached input tokens, per request.** When a `generateContent` call carries a `cached_content` handle, the tokens that came from the cache are still counted as input tokens, but they are priced below the normal input rate. The rest of the turn — your question, retrieved extras, and the generated output — is billed at ordinary rates. So a turn that reads a 200k-token cached document and adds a 40-token question pays the discounted rate on 200k plus the full rate on 40, plus output. **Meter 2 — cache storage, per unit time.** Holding the cache alive costs money on a schedule, quoted per million cached tokens per hour. This charge accrues whether the cache is read a thousand times or never. It is the direct price of the TTL you chose. ## Why that makes TTL an economic knob The break-even is a small piece of arithmetic. Let *N* be the number of cached tokens, *r* the saving per cached token versus full input price, *s* the storage price per token-hour, and *h* the hours the cache is alive. Caching wins when ``` reuses × N × r > N × s × h ``` which reduces to `reuses > s × h / r`. The cache size cancels out — what actually matters is **how many times you re-read the prefix per hour of storage**. That is why a chat session over one contract, asking twenty questions in ten minutes, is a textbook win, while a cache created for a single question and left with a one-hour TTL is a straight loss: you paid storage for an hour to save on one call. Two practical corollaries. First, set the TTL to the expected session length, not to a round number that feels safe — a ten-minute session does not need a one-hour cache. Second, delete caches when the session ends. `client.caches.delete(name=...)` stops the meter immediately and is the single most effective cost control in a caching layer. ## Adjusting a live cache Contents are immutable, but expiry is not. `client.caches.update(name=..., config=types.UpdateCachedContentConfig(ttl="600s"))` resets the remaining lifetime from now, and an absolute `expire_time` works too. That supports a sliding-window policy: create with a short TTL, and extend on each hit while the user is still active. You pay storage only for the time the session genuinely spans, instead of provisioning for the worst case up front. ## What the usage object tells you `response.usage_metadata` reports `prompt_token_count` (the whole input, cached part included), `cached_content_token_count` (how much of it was served from cache), and `candidates_token_count` for the output. To reconstruct a bill you need the split: fresh input is `prompt_token_count - cached_content_token_count`. Logging both numbers per call is how you catch a cache that has silently stopped hitting — the total stays flat while the cached count drops to zero and the invoice quietly doubles. ## The long-context price tier Separately from caching, Gemini prices very long prompts higher. As of mid-2026, Gemini 2.5 Pro bills prompts above roughly 200k tokens at an elevated per-token input rate, and output for those requests is also priced higher. That interacts with caching in a way people miss: caching reduces the *rate* you pay on repeated tokens, but a 700k-token cached prefix still sits in the long-prompt tier on every call, and the prompt is still 700k tokens of prefill latency. Caching is a discount, not a free context window. ## Minimums and wasted attempts Explicit caches have a floor — around 1,024 tokens for Gemini 2.5 Flash and 2,048 for 2.5 Pro as of mid-2026. Below that the create call is rejected, so a caching layer that blindly caches every system prompt will throw on the small ones. Count tokens with `client.models.count_tokens` first and fall through to implicit caching for anything under the threshold. ## How to talk about it A strong answer names both meters, states that TTL drives the second one, and gives the reuse-per-hour break-even rather than a vague "caching saves money". A weak answer says cached tokens are free.
- Your cached prefix is 400k tokens and each session asks two questions. Is caching worth it?Almost certainly not with a long TTL. Two reads spread the storage cost over very few requests, and storage accrues on wall-clock time regardless of traffic. Either shorten the TTL to the real session length and delete on completion, or skip explicit caching and let implicit caching pick up whatever prefix reuse exists for free. Measure the two meters before committing.
- Does caching reduce the latency of a long-context request as well as the cost?Yes, materially — the model does not have to re-prefill the cached prefix, so time to first token drops on cache hits. But it does not eliminate the cost of length: the cached tokens still count toward the prompt size, still sit in the long-prompt price tier if the prompt is large, and still consume context budget. Latency improves; the prompt does not get smaller.
- Why should a cache-hit metric be part of production monitoring?Because a cache regression is invisible in behaviour and loud in billing. If a model upgrade, an expiry, or a refactor drops cached_content_token_count to zero, responses stay correct while every request reverts to full-price input. Emitting cached versus fresh token counts per call turns a silent doubling of spend into an alert.
saying these in an interview costs you the question
- Claims cached tokens are free rather than discounted
- Ignores the hourly storage charge tied to TTL
- Sets a long TTL by default because it feels safer
- Assumes caching removes the long-prompt price tier
- Never deletes caches after a session ends