skip to content

For document QA on Gemini, how do you scope cache keys and TTLs per tenant?

level: principalimportance: should knowfreq 33%

answer

  1. cache identity is an authorisation boundary
  2. hash the content into the key
  3. TTL from measured session length
  4. create-on-miss, never hard-fail
  5. cap cached tokens per tenant

basics

~20 s

Key each cache by tenant, document version and target model, since caches are immutable and model-bound. Set TTL to the observed session length, extend on activity, delete on close, and only cache prefixes whose reuse per hour beats the storage fee.

solid answer

~50 s

Treat Gemini caches as short-lived, per-session resources rather than a persistent index. The cache key should combine **tenant/document identity, content version and exact model name** — content is immutable so a document edit means a new cache, and a cache is bound to the model it was created for so a model roll-forward invalidates every handle. Set the TTL from measured session duration rather than a default hour, extend it with `caches.update` while the user is active, and call `caches.delete` on session end to stop the storage meter. Guard creation behind a reuse test: below the model's minimum cacheable size, or with only one or two expected reads, caching loses money and implicit caching is the better default. Operationally you need three things: a create-on-miss path so an expired handle degrades to a rebuild rather than an outage, cached-versus-fresh token counts in metrics, and a per-tenant cap so one large corpus cannot dominate storage spend.

code

python · 13 lines
python
import hashlib

def cache_key(tenant, doc_bytes, model, system_instruction):
    h = hashlib.sha256(doc_bytes + system_instruction.encode()).hexdigest()[:16]
    return f"{tenant}:{model}:{h}"

def get_or_create(key, build_config, model):
    entry = store.get(key)
    if entry and entry.expires_at > now():
        return entry.cache_name
    cache = client.caches.create(model=model, config=build_config)
    store.put(key, cache.name, expires_at=now() + build_config.ttl)
    return cache.name

go deeper

for a junior

Understand that caches expire and are tied to one model, so any code using them needs a path that rebuilds when the handle stops working.

for a middle

Explain why immutability forces content-versioned keys and why TTL should be derived from session length rather than left at the one-hour default.

for a senior

Design the request path end to end: create-on-miss with retry, sliding TTL on activity, delete on session close, orphan sweeps, and hit-rate plus rebuild-rate metrics.

for a principal

Own the economics and the boundaries — per-tenant storage caps, which traffic classes get explicit caches at all, shared handle storage across instances, and the policy for model rollouts that invalidate every cache at once.

## Why this needs a design, not a helper function A single-user script can create a cache, use it, and forget it. A multi-tenant product cannot: caches cost storage per hour, are immutable, are bound to a model version, and expire out from under callers. Those four properties drive the whole design. ## Cache identity The key must capture everything that would make a cached prefix wrong or unusable: - **Tenant / document identity.** Never share a cache across tenants unless the content is genuinely global. A cached prefix is content that will be shown to whoever uses the handle, so the cache boundary is an authorisation boundary. - **Content version.** Contents are immutable, so a document edit cannot update a cache — it must produce a new one. Hashing the exact cached material into the key makes staleness structurally impossible and gives you free deduplication when two sessions open the same version. - **Model name.** A cache belongs to one model. Include it so that rolling forward to a new model generation simply misses on every key and rebuilds, instead of throwing errors from stale handles. - **Prefix composition.** If the cached block includes a system instruction or tool declarations that vary by product surface, they belong in the key too. A local map from that key to `{cache_name, expires_at}` is what your request path consults. Anything not found, or found expired, goes through create-on-miss. ## TTL policy Defaulting to an hour is the most common waste. Better: 1. **Measure real session length.** If the p90 document session is eleven minutes, an hour of storage per session is roughly a five-fold overpay. 2. **Create short, extend on activity.** Start at something like the p50 session, and call `client.caches.update(name=..., config=UpdateCachedContentConfig(ttl=...))` on each hit to slide the window forward. Idle sessions then expire on their own. 3. **Delete explicitly on close.** A logout, tab close or job completion should trigger `client.caches.delete`. This is the single highest-leverage cost control, because storage bills on wall-clock time regardless of traffic. 4. **Sweep orphans.** `client.caches.list()` on a schedule catches handles your process lost track of after a crash or deploy. ## The admission test Not every prefix deserves a cache. Gate creation on: - **Size floor.** Below the model minimum — roughly 1,024 tokens on Gemini 2.5 Flash and 2,048 on 2.5 Pro as of mid-2026 — creation is rejected. Count with `client.models.count_tokens` first. - **Expected reuse per hour.** Caching wins when the per-request saving times the number of reads outweighs storage times hours alive. The prefix size cancels out of that comparison, so the real question is how many times you will re-read the prefix before it expires. One or two reads is a loss. - **Access pattern.** A document opened by many concurrent users of the same tenant amortises one cache across all of them, which is far better than per-user caches of the same bytes — another reason to key on content rather than on session. ## Failure handling Expired or invalidated handles must never surface as user-visible errors. The request path is: look up key → if missing or the call fails, create → retry once → serve. Because a rebuild costs a full-price prefill, log it: a spike in rebuild rate usually means a TTL set below actual session length, or a model rollout in progress. Deploys deserve explicit thought. If instance A holds handles in memory and you deploy, instance B has none and rebuilds everything — a cost spike and a latency spike at once. A shared store for the key→handle map (Redis or similar), keyed by content hash, keeps caches alive across restarts and lets several application instances share one cache. ## Metrics that matter Three series answer nearly every question you will be asked about this system: cache hit rate (fraction of requests with non-zero `cached_content_token_count`), rebuild rate per session, and cached versus fresh tokens per tenant. The third is the one finance will ask for, because storage spend is per-tenant even when the code is not. ## Guardrails Cap cached tokens per tenant so a single customer uploading a huge corpus cannot consume the storage budget; degrade over the cap to uncached or retrieval-backed calls rather than failing. Consider a tier policy — explicit caching for interactive sessions where latency is felt, implicit caching alone for batch and background work where nobody is waiting. ## What separates a strong answer Naming the immutability and model-binding constraints as the reason keys look the way they do; treating TTL as a spend decision with a measured input; and being explicit that an expired cache is an expected event with a defined recovery path, not an exception.

  • A deploy rolls your service and cache hit rate collapses. What is the likely cause?
    Handles were held only in process memory, so the new instances start with an empty map and rebuild every cache at full-price prefill. Move the key-to-handle map into a shared store keyed by content hash and model, so surviving caches are rediscovered rather than duplicated. It also lets multiple instances share one cache instead of each creating its own copy of the same bytes.
  • How would you stop one tenant's storage spend from dominating the bill?
    Cap cached tokens per tenant and enforce it at admission: track the sum of cached token counts for that tenant's live handles, and above the cap either refuse new caches and fall back to uncached calls, or evict the least recently used ones by deleting them early. Report cached tokens per tenant so the cap can be tuned per plan rather than guessed globally.
  • Why key the cache on a content hash rather than a document ID?
    Because caches are immutable — an edited document can never update its cache, so an ID-keyed entry silently serves stale content until it expires. Hashing the exact cached bytes makes a new version a new key by construction, eliminating a whole class of staleness bugs, and makes two sessions over identical content share one cache automatically.

saying these in an interview costs you the question

  • Treats caches as a long-lived index rather than session-scoped
  • Keys on document ID, so edits serve stale cached content
  • Leaves TTL at the default without measuring session length
  • Holds cache handles only in process memory across deploys
  • Caches every prefix without checking expected reuse

context