skip to content

Helicone

Observability by proxy: point your base URL at Helicone and every request is logged with cost, latency, caching, and usage breakdowns without touching SDK code. Interviewers ask about the trade-off — near-zero integration effort against putting a third party in the request path.

on this pageshow

questions

12

How do you turn on Helicone's response cache and control how long entries live?

level: juniorimportance: must knowfreq 60%

answer

  1. one header switches it on
  2. TTL comes from a standard HTTP header
  3. match is exact, not semantic
  4. response header tells you HIT or MISS
  5. cannot work in async logging mode

basics

~10 s

Send Helicone-Cache-Enabled: true on the request, and set the lifetime with a standard Cache-Control: max-age=<seconds> header. Caching only works through the Helicone proxy, because something has to answer in place of the provider.

solid answer

~50 s

Helicone's cache is header-driven, not SDK-driven. Point your `base_url` at the Helicone gateway, authenticate with `Helicone-Auth`, and add `Helicone-Cache-Enabled: true` to the request. The entry's lifetime comes from a normal `Cache-Control: max-age=<seconds>` header on the same request; if you omit it Helicone applies its own default TTL, which is far longer than most people expect, so set it explicitly. The cache key is derived from the request itself — model, messages, and sampling parameters — so anything that varies per call (a timestamp, a user's name, a request id inside the prompt) means you will never see a hit. On the way back Helicone sets a `Helicone-Cache` response header of `HIT` or `MISS`, which is the only honest way to confirm caching is working. Because a hit is served without calling the provider, it costs nothing and returns in gateway latency, and the dashboard reports the saving.

code

python · 17 lines
python
from openai import OpenAI

client = OpenAI(
    base_url="https://oai.helicone.ai/v1",
    api_key="sk-your-openai-key",
    default_headers={
        "Helicone-Auth": "Bearer sk-your-helicone-key",
        "Helicone-Cache-Enabled": "true",
        "Cache-Control": "max-age=3600",
    },
)

resp = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "What is the capital of France?"}],
)
print(resp.choices[0].message.content)

go deeper

for a junior

Know the two headers by name: Helicone-Cache-Enabled turns it on, Cache-Control max-age sets how long the entry lives. Say plainly that a hit is served without calling the provider, so it is free.

for a middle

Explain what forms the cache key — model, messages and parameters, matched exactly — and why any per-request variation inside the prompt drives the hit rate to zero. Mention the HIT/MISS response header as the way to verify.

for a senior

Show judgment about which calls belong in the cache. Talk about TTL choice against how fast the underlying data changes, and about the fact that enabling the cache means completion bodies are now stored at a third party for the TTL's duration.

for a principal

Frame it as a control-plane decision: caching only exists because Helicone sits in the request path, which is the same reason it is a new dependency and a new data store. Be ready to argue when that trade is worth a percentage off the model bill.

## Why the cache is a header Helicone's whole design premise is that you should not have to import a library. You change the `base_url` of your existing OpenAI-compatible client to Helicone's gateway, add an auth header, and every request flows through Helicone on its way to the provider. Once Helicone is physically in the request path, it can do more than record what happened — it can answer. Response caching is the first of those behaviour-changing features, and like everything else in the product it is switched on with an HTTP header rather than a config file or a client object. ## The two headers `Helicone-Cache-Enabled: true` opts this request into the cache. It is per request, not per account, so you can cache the deterministic classification call and leave the creative-writing call uncached simply by sending different headers on each. `Cache-Control: max-age=<seconds>` sets how long an entry stays valid. This is deliberately the standard HTTP cache header rather than a Helicone-specific one, so `max-age=3600` means one hour exactly as it would anywhere else. If you send `Helicone-Cache-Enabled` without a `Cache-Control`, Helicone falls back to a default TTL measured in days, and the practical consequence is a stale answer served long after you changed the prompt template and forgot the cache existed. Treat the TTL as mandatory even though the API treats it as optional. ## What forms the cache key The key is computed from the outgoing request: the model, the full message array, and the sampling parameters. Two calls hit the same entry only if those are byte-identical. This is exact matching, not similarity matching — there is no embedding involved, and nothing in Helicone's cache decides that two differently-worded prompts mean the same thing. That detail is the single biggest source of "caching doesn't work" tickets. A system prompt that interpolates `datetime.now()`, a prompt that embeds the user's display name, a request id threaded through for tracing — each of them makes every request unique and the hit rate zero. The fix is to move the varying material out of the prompt body and into headers or metadata, or to accept that this particular call is not cacheable. ## Confirming it works The response carries a `Helicone-Cache` header whose value is `HIT` or `MISS`. Read it in a smoke test rather than inferring caching from latency, because a fast provider response and a cache hit look similar from the outside. The dashboard also attributes cost savings to cached requests, which is the number a finance-minded reviewer will ask for. ## The proxy requirement Helicone offers two integration modes: the proxy, where requests travel through the gateway, and async logging, where your application calls the provider directly and ships a copy of the request and response to Helicone afterwards. Caching is impossible in the second mode by construction — by the time Helicone learns about the call, the provider has already been billed. The same is true of the other behaviour-changing gateway features. If your team chose async logging for latency or blast-radius reasons, the cache is simply not available to you, and it is worth knowing that before you promise a cost reduction. ## Operational consequences A cache hit is free and fast, which is exactly why it is dangerous in the wrong place. Cached answers do not reflect changes to the underlying model, changes to retrieved context that was not part of the prompt, or the passage of time. Any request whose correct answer depends on the current state of the world should either be uncached or carry a short TTL. Cached responses are still logged as requests in Helicone, so your trace volume does not fall when your provider bill does — the trace shows the cached status rather than a fresh generation. Finally, remember the trust boundary. Enabling the cache means Helicone stores the completion body, not just metadata about it, for the duration of the TTL. On a workload carrying regulated data that is a storage decision, not just a performance one, and it belongs in the same conversation as whether prompt content should be logged at all.

  • How would you prove in a test that a second identical call was actually served from Helicone's cache?
    Read the `Helicone-Cache` response header and assert it is `MISS` on the first call and `HIT` on the second. With the OpenAI Python SDK you get at the raw headers through `client.chat.completions.with_raw_response.create(...)`. Do not infer a hit from latency alone — a warm provider can return as fast as the gateway, and a passing latency assertion will flap in CI.
  • Which requests in a RAG service are actually safe to cache this way?
    Ones whose correct answer is a pure function of the prompt: classification, extraction, translation, query rewriting, and embedding-adjacent utility calls. The final answer generation usually is not, because its prompt embeds retrieved chunks that change as the index changes, and a hit would serve an answer built from documents that have since been edited or deleted.
  • What happens to caching if your team uses Helicone's async logging integration instead of the proxy?
    It is unavailable. In async logging your application calls the provider itself and posts a copy of the exchange to Helicone afterwards, so Helicone never has the chance to answer in the provider's place. Caching, gateway rate limits and gateway retries all require the request to physically traverse the proxy.

saying these in an interview costs you the question

  • Thinks Helicone matches semantically similar prompts, not exact requests
  • Assumes the cache works with async logging as well as the proxy
  • Omits Cache-Control and is surprised by a days-long default TTL
  • Leaves a timestamp in the prompt and reports the cache as broken
  • Claims cached requests stop appearing in the Helicone dashboard

context

open as a page

How do you route OpenAI traffic through Helicone's proxy, and which header authenticates it?

level: juniorimportance: must knowfreq 72%

basics

~10 s

Point the OpenAI client's base URL at Helicone's proxy host, https://oai.helicone.ai/v1, and send a Helicone-Auth header holding "Bearer" plus your Helicone API key. Your provider key still travels in the usual Authorization header.

open as a page

When would you choose Helicone's async logging over its proxy integration?

level: middleimportance: must knowfreq 62%

basics

~20 s

Async logging keeps the provider call direct and ships the log separately, so Helicone adds no latency and cannot take your feature down. Choose it when the request path must stay untouched; choose the proxy when you want gateway behaviour, not just logs.

open as a page

What breaks when Helicone's proxy is slow or unreachable, and how do you limit the blast radius?

level: seniorimportance: must knowfreq 58%

basics

~20 s

With the base URL pointed at the proxy, every model call goes through it, so a degraded proxy degrades the feature itself — calls slow down, hang or fail. Limit the damage by making the base URL a flippable runtime setting, bounding timeouts, and moving latency-critical paths to async logging.

open as a page

In Helicone, what does the Helicone-Cache-Seed header change about cache hits?

level: middleimportance: should knowfreq 42%

basics

~20 s

Helicone-Cache-Seed partitions the cache into namespaces. Two identical requests sent with different seed values never share an entry, so you can isolate caches per user or tenant, and changing the seed instantly invalidates everything cached under the old one.

open as a page

How do you express a per-user quota in Helicone's Helicone-RateLimit-Policy header?

level: middleimportance: should knowfreq 55%

basics

~10 s

Send Helicone-RateLimit-Policy with the form quota;w=window;u=unit;s=segment — for example 1000;w=3600;u=cents;s=user caps each end user at $10 of spend per hour. Segmenting by user also requires a Helicone-User-Id header on the request.

open as a page

What do Helicone-Property-* headers do to a logged request, and why set them?

level: middleimportance: should knowfreq 55%

basics

~20 s

A header named Helicone-Property-<Name> attaches an arbitrary key/value tag to that request's log row. Those tags become filter and group-by dimensions in Helicone's dashboards, which is how you get cost and latency broken down per feature, environment or customer.

open as a page

How do Helicone's session headers group an agent's many calls into one trace?

level: middleimportance: should knowfreq 42%

basics

~20 s

Sending the same Helicone-Session-Id on every call of one run ties those requests together, Helicone-Session-Name labels the run, and Helicone-Session-Path places each call in a hierarchy so a multi-step agent reads as a tree instead of scattered rows.

open as a page

What does Helicone-Retry-Enabled retry, and what does it cost the caller?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Helicone-Retry-Enabled: true makes the proxy re-send a failed call to the provider on rate-limit 429s and 5xx errors, with exponential backoff. The retries happen inside one client request, so the caller sees a single much slower call and must widen its timeout.

open as a page

Should per-customer LLM spend caps live in Helicone's gateway or your own service?

level: principalimportance: should knowfreq 33%

basics

~20 s

Use both, for different jobs. A Helicone-RateLimit-Policy is a cheap backstop against runaway spend on traffic that passes through the proxy. The entitlement your billing and product depend on must live in a service you own, because the gateway sees only proxied calls.

open as a page

How would you standardise Helicone properties and user IDs across teams for cost attribution?

level: principalimportance: should knowfreq 35%

basics

~20 s

Agree a small fixed vocabulary — environment, service, feature, version — and enforce it in a shared client wrapper rather than by convention, with Helicone-User-Id carrying an opaque internal identifier. Decide up front what granularity of attribution the business will actually act on.

open as a page

Why can two identical Helicone cache hits return different completions?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Because of bucketing. Helicone-Cache-Bucket-Max-Size tells the gateway to store several distinct responses under one cache key and return one of them at random on a hit, trading determinism for variety so repeat callers do not all see identical text.

open as a page