skip to content

How do you turn on Helicone's response cache and control how long entries live?

level: juniorimportance: must knowfreq 60%

answer

  1. one header switches it on
  2. TTL comes from a standard HTTP header
  3. match is exact, not semantic
  4. response header tells you HIT or MISS
  5. cannot work in async logging mode

basics

~10 s

Send Helicone-Cache-Enabled: true on the request, and set the lifetime with a standard Cache-Control: max-age=<seconds> header. Caching only works through the Helicone proxy, because something has to answer in place of the provider.

solid answer

~50 s

Helicone's cache is header-driven, not SDK-driven. Point your `base_url` at the Helicone gateway, authenticate with `Helicone-Auth`, and add `Helicone-Cache-Enabled: true` to the request. The entry's lifetime comes from a normal `Cache-Control: max-age=<seconds>` header on the same request; if you omit it Helicone applies its own default TTL, which is far longer than most people expect, so set it explicitly. The cache key is derived from the request itself — model, messages, and sampling parameters — so anything that varies per call (a timestamp, a user's name, a request id inside the prompt) means you will never see a hit. On the way back Helicone sets a `Helicone-Cache` response header of `HIT` or `MISS`, which is the only honest way to confirm caching is working. Because a hit is served without calling the provider, it costs nothing and returns in gateway latency, and the dashboard reports the saving.

code

python · 17 lines
python
from openai import OpenAI

client = OpenAI(
    base_url="https://oai.helicone.ai/v1",
    api_key="sk-your-openai-key",
    default_headers={
        "Helicone-Auth": "Bearer sk-your-helicone-key",
        "Helicone-Cache-Enabled": "true",
        "Cache-Control": "max-age=3600",
    },
)

resp = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "What is the capital of France?"}],
)
print(resp.choices[0].message.content)

go deeper

for a junior

Know the two headers by name: Helicone-Cache-Enabled turns it on, Cache-Control max-age sets how long the entry lives. Say plainly that a hit is served without calling the provider, so it is free.

for a middle

Explain what forms the cache key — model, messages and parameters, matched exactly — and why any per-request variation inside the prompt drives the hit rate to zero. Mention the HIT/MISS response header as the way to verify.

for a senior

Show judgment about which calls belong in the cache. Talk about TTL choice against how fast the underlying data changes, and about the fact that enabling the cache means completion bodies are now stored at a third party for the TTL's duration.

for a principal

Frame it as a control-plane decision: caching only exists because Helicone sits in the request path, which is the same reason it is a new dependency and a new data store. Be ready to argue when that trade is worth a percentage off the model bill.

## Why the cache is a header Helicone's whole design premise is that you should not have to import a library. You change the `base_url` of your existing OpenAI-compatible client to Helicone's gateway, add an auth header, and every request flows through Helicone on its way to the provider. Once Helicone is physically in the request path, it can do more than record what happened — it can answer. Response caching is the first of those behaviour-changing features, and like everything else in the product it is switched on with an HTTP header rather than a config file or a client object. ## The two headers `Helicone-Cache-Enabled: true` opts this request into the cache. It is per request, not per account, so you can cache the deterministic classification call and leave the creative-writing call uncached simply by sending different headers on each. `Cache-Control: max-age=<seconds>` sets how long an entry stays valid. This is deliberately the standard HTTP cache header rather than a Helicone-specific one, so `max-age=3600` means one hour exactly as it would anywhere else. If you send `Helicone-Cache-Enabled` without a `Cache-Control`, Helicone falls back to a default TTL measured in days, and the practical consequence is a stale answer served long after you changed the prompt template and forgot the cache existed. Treat the TTL as mandatory even though the API treats it as optional. ## What forms the cache key The key is computed from the outgoing request: the model, the full message array, and the sampling parameters. Two calls hit the same entry only if those are byte-identical. This is exact matching, not similarity matching — there is no embedding involved, and nothing in Helicone's cache decides that two differently-worded prompts mean the same thing. That detail is the single biggest source of "caching doesn't work" tickets. A system prompt that interpolates `datetime.now()`, a prompt that embeds the user's display name, a request id threaded through for tracing — each of them makes every request unique and the hit rate zero. The fix is to move the varying material out of the prompt body and into headers or metadata, or to accept that this particular call is not cacheable. ## Confirming it works The response carries a `Helicone-Cache` header whose value is `HIT` or `MISS`. Read it in a smoke test rather than inferring caching from latency, because a fast provider response and a cache hit look similar from the outside. The dashboard also attributes cost savings to cached requests, which is the number a finance-minded reviewer will ask for. ## The proxy requirement Helicone offers two integration modes: the proxy, where requests travel through the gateway, and async logging, where your application calls the provider directly and ships a copy of the request and response to Helicone afterwards. Caching is impossible in the second mode by construction — by the time Helicone learns about the call, the provider has already been billed. The same is true of the other behaviour-changing gateway features. If your team chose async logging for latency or blast-radius reasons, the cache is simply not available to you, and it is worth knowing that before you promise a cost reduction. ## Operational consequences A cache hit is free and fast, which is exactly why it is dangerous in the wrong place. Cached answers do not reflect changes to the underlying model, changes to retrieved context that was not part of the prompt, or the passage of time. Any request whose correct answer depends on the current state of the world should either be uncached or carry a short TTL. Cached responses are still logged as requests in Helicone, so your trace volume does not fall when your provider bill does — the trace shows the cached status rather than a fresh generation. Finally, remember the trust boundary. Enabling the cache means Helicone stores the completion body, not just metadata about it, for the duration of the TTL. On a workload carrying regulated data that is a storage decision, not just a performance one, and it belongs in the same conversation as whether prompt content should be logged at all.

  • How would you prove in a test that a second identical call was actually served from Helicone's cache?
    Read the `Helicone-Cache` response header and assert it is `MISS` on the first call and `HIT` on the second. With the OpenAI Python SDK you get at the raw headers through `client.chat.completions.with_raw_response.create(...)`. Do not infer a hit from latency alone — a warm provider can return as fast as the gateway, and a passing latency assertion will flap in CI.
  • Which requests in a RAG service are actually safe to cache this way?
    Ones whose correct answer is a pure function of the prompt: classification, extraction, translation, query rewriting, and embedding-adjacent utility calls. The final answer generation usually is not, because its prompt embeds retrieved chunks that change as the index changes, and a hit would serve an answer built from documents that have since been edited or deleted.
  • What happens to caching if your team uses Helicone's async logging integration instead of the proxy?
    It is unavailable. In async logging your application calls the provider itself and posts a copy of the exchange to Helicone afterwards, so Helicone never has the chance to answer in the provider's place. Caching, gateway rate limits and gateway retries all require the request to physically traverse the proxy.

saying these in an interview costs you the question

  • Thinks Helicone matches semantically similar prompts, not exact requests
  • Assumes the cache works with async logging as well as the proxy
  • Omits Cache-Control and is surprised by a days-long default TTL
  • Leaves a timestamp in the prompt and reports the cache as broken
  • Claims cached requests stop appearing in the Helicone dashboard

context