How do you route OpenAI traffic through Helicone's proxy, and which header authenticates it?
answer
- two URLs, two keys
- the base URL does the work
- one header names your Helicone account
- oai.helicone.ai in front of OpenAI
- Helicone-Auth carries Bearer plus Helicone key
basics
~10 sPoint the OpenAI client's base URL at Helicone's proxy host, https://oai.helicone.ai/v1, and send a Helicone-Auth header holding "Bearer" plus your Helicone API key. Your provider key still travels in the usual Authorization header.
solid answer
~40 sHelicone's proxy integration is deliberately a two-line change: swap the base URL of your existing client to `https://oai.helicone.ai/v1`, and add the header `Helicone-Auth: Bearer <HELICONE_API_KEY>`. Nothing else about the call changes — the request body, the model name and your provider credential in `Authorization` are forwarded verbatim to OpenAI, and the provider's response comes back to you unchanged. Because Helicone terminates the request, it records the request and response bodies, latency, model, token counts and a computed cost, and groups them in its dashboard. Two distinct credentials are in play, which is the usual first mistake: the Helicone key identifies your Helicone account and the provider key identifies you to OpenAI. Being in the request path is also what lets Helicone's gateway features act on a call at all.
code
python · 16 linesimport os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["OPENAI_API_KEY"],
base_url=os.environ.get("OPENAI_BASE_URL", "https://oai.helicone.ai/v1"),
default_headers={
"Helicone-Auth": f"Bearer {os.environ['HELICONE_API_KEY']}",
},
)
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Say hello"}],
)
print(response.choices[0].message.content)go deeper
Be able to state the two-part setup out loud: change the client's base URL to Helicone's proxy host and add the Helicone-Auth header with your Helicone key. Remember the provider key stays where it always was.
Explain what the proxy can capture because it terminates the request — bodies, latency, tokens, computed cost — and why no code instrumentation is needed for that.
Show that you treat the base URL as configuration, and that you know cost figures are derived from observed usage, so streamed traffic without a usage block can drift from the provider invoice.
Own the decision of whether a hosted third party may sit in front of every model call at all, including where the request bodies are stored and who is allowed to read them.
## The idea Most observability tools ask you to install an SDK and wrap your call sites. Helicone's headline integration asks you to change a URL. It runs a hosted reverse proxy that speaks the provider's own API shape, so an existing client keeps working when you point it at Helicone instead of the provider. For OpenAI-shaped traffic the proxy host is `https://oai.helicone.ai/v1`. Helicone publishes a host per provider (for example `https://anthropic.helicone.ai` for Anthropic's API); the pattern is the same in each case — same paths, same request bodies, same responses. ## The two credentials A proxied call carries two authorizations, and confusing them is the classic failure: - `Authorization: Bearer <PROVIDER_KEY>` — your OpenAI key. Helicone passes it through so the provider can bill and authenticate you. - `Helicone-Auth: Bearer <HELICONE_API_KEY>` — your Helicone key. It tells Helicone which account and project the log belongs to. If you omit `Helicone-Auth`, the request is not attributable to your account. If you put the Helicone key in `Authorization`, OpenAI rejects it as an invalid key. Note that the Helicone key value itself includes the `Bearer ` prefix — it is a normal bearer-style header, not a bare token field. ## What Helicone records Because the proxy sits between you and the provider, it sees the whole exchange and needs no cooperation from your code to log: - the request body, including the rendered prompt and any tool definitions; - the response body, including the completion; - wall-clock latency as measured at the proxy; - the model, prompt/completion token counts, and a cost derived from the model's published price; - the status code, so failures and provider errors show up beside successes. That is where the cost dashboards come from: Helicone multiplies observed token usage by a per-model rate table, so "spend by model", "spend by day" and (once you add custom property headers) "spend by feature or customer" all fall out of the same log rows without you emitting a single metric. One accuracy caveat worth knowing: streamed responses only carry a usage block if you ask the provider to include it. When usage is absent, the cost number for that row is derived rather than reported, so a heavily streamed workload can drift from the provider's own invoice. ## Setting it up in a typical client Every mainstream client library has a base-URL option and a default-headers option, which is all the integration needs. Prefer setting the base URL from configuration rather than hardcoding it — you want the ability to point straight at the provider again without a code change (that becomes important the day the proxy is unhealthy). Setting the Helicone header as a client default means every call made through that client is logged. You can still add or override headers per request when you want to attach call-specific metadata. ## What being in the path buys and costs The price of the one-line setup is that a third party is now inside your critical request path: one more network hop of latency and one more thing that can be down. The benefit is symmetrical — because Helicone handles the request before the provider does, it can *change* what happens to it (serving a stored response, applying a policy, retrying a failure). Logging alone does not require the proxy; behaviour-changing gateway features do. That tradeoff, not the setup itself, is what interviewers actually push on: the setup is a URL and a header, and the interesting question is when you are willing to accept the dependency. ## Verifying it works After the swap, a successful call looks identical to your application — same response object, same latency order of magnitude. The check is on Helicone's side: the request should appear in the dashboard within seconds, with a model, a token count and a cost attached. If calls succeed but nothing appears, the `Helicone-Auth` header is the first thing to inspect; if calls start failing with an auth error from the provider, the two keys have been crossed.
- What happens if you send the Helicone key in the Authorization header instead?The provider receives it as your OpenAI credential and rejects the call with an authentication error, so the request never reaches a model. Helicone may still log the failed attempt if `Helicone-Auth` was set correctly. The two headers are independent: `Authorization` is forwarded to the provider untouched, `Helicone-Auth` is consumed by the proxy and identifies your Helicone account.
- Does the proxy require you to change your request payloads or model names?No. The proxy speaks the provider's own API shape, so paths, request bodies, model identifiers and response schemas are unchanged — that is the whole point of the integration. Client libraries built for the provider keep working, including streaming. The only thing you change is where the client points and one extra header.
- How would you make the proxy easy to turn off later?Read the base URL from configuration or an environment variable rather than hardcoding it, so flipping back to the provider's own host is a config change and a restart, not a deploy. Keep the Helicone header harmless when unset. That switch is the mitigation you reach for when the proxy is slow or unavailable, so it should exist before you need it.
saying these in an interview costs you the question
- Thinking one API key covers both Helicone and the provider
- Believing the proxy requires rewriting request payloads
- Assuming logging needs an SDK install as well as the URL swap
- Hardcoding the proxy base URL with no way to switch back
- Expecting logs to appear without the Helicone-Auth header