How do you express a per-user quota in Helicone's Helicone-RateLimit-Policy header?
answer
- one structured header, semicolon separated
- w is the window in seconds
- the unit can meter money, not calls
- segment defaults to the whole organization
- user segmentation needs an id header too
basics
~10 sSend Helicone-RateLimit-Policy with the form quota;w=window;u=unit;s=segment — for example 1000;w=3600;u=cents;s=user caps each end user at $10 of spend per hour. Segmenting by user also requires a Helicone-User-Id header on the request.
solid answer
~50 sThe policy is a single structured header: `Helicone-RateLimit-Policy: [quota];w=[seconds];u=[unit];s=[segment]`. The quota and the window `w` are required; `u` defaults to counting requests and can instead be `cents`, which meters spend rather than call volume; `s` defaults to global and can be `user` or the name of a custom property you already send. So `1000;w=3600;u=cents;s=user` means ten dollars per hour per end user, while a bare `500;w=60` means five hundred requests a minute across the whole organization. When you segment by user, Helicone needs to know who the user is, which is the `Helicone-User-Id` header — omit it and the request cannot be attributed. Over quota, the gateway returns 429 without forwarding to the provider, and the response carries `Helicone-RateLimit-*` headers describing the limit and what remains, which is what your client should read rather than guessing a backoff.
code
bash · 7 linescurl https://oai.helicone.ai/v1/chat/completions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Helicone-Auth: Bearer $HELICONE_API_KEY" \
-H "Helicone-User-Id: customer-42" \
-H "Helicone-RateLimit-Policy: 1000;w=3600;u=cents;s=user" \
-H "Content-Type: application/json" \
-d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"hi"}]}'go deeper
Recall that the quota is one header with a number and a window in seconds, and that Helicone answers with a 429 instead of calling the provider once the limit is passed.
Be able to read a policy string field by field: quota, w for the window, u for requests versus cents, s for the segment. Name the defaults and the fact that s=user needs Helicone-User-Id alongside it.
Show you have rolled one of these out. Talk about picking the number from measured usage, about the 429 needing a product-level response, and about how a cents quota is counted after the response rather than before the call.
Own the limits of a gateway quota: it only covers traffic that traverses the proxy, so it is a spend backstop rather than an authoritative budget. Argue where the real per-customer entitlement should live and what the gateway is for.
## One header, four fields Helicone expresses a quota as a structured value on `Helicone-RateLimit-Policy`, borrowing the shape of HTTP's rate-limit header conventions: `[quota];w=[time window in seconds];u=[unit];s=[segment]` The leading number and `w` are mandatory. `u` and `s` are optional and have defaults, which is where most of the confusion comes from — a policy that looks like it is per user often is not, because `s` was never set. ## Units: requests or cents The default unit counts requests. `u=cents` instead counts spend, using Helicone's own cost calculation for the model and token usage of each call. That distinction matters more here than in a conventional API gateway, because LLM requests are not fungible: a thousand short classification calls and a thousand long-context summarization calls differ in cost by orders of magnitude. If what you are protecting is a budget, count cents; if what you are protecting is a downstream capacity or a fairness property, count requests. Note that a cent-denominated quota can only be evaluated once a response exists and its usage is known, so it bounds spend after the fact rather than pre-authorizing it — a single very expensive call can carry you past the line before it is enforced. ## Segments: global, user, or a custom property With no `s`, the policy applies to the whole organization: one bucket for all traffic. `s=user` gives every distinct `Helicone-User-Id` its own bucket, which is the multi-tenant case. You can also segment by a custom property you are already attaching for analytics, which is how you build a quota per workspace, per feature, or per environment without introducing a new identity concept — the property name goes in `s`. The dependency is easy to miss: `s=user` is meaningless unless `Helicone-User-Id` is present on the request. If your service already sends that header for cost attribution, segmented quotas are free to add; if it does not, adding the policy alone silently fails to do what you meant. ## What happens at the limit When a request would exceed its bucket, the Helicone proxy short-circuits it: the provider is never called, and the caller receives a 429. The response includes `Helicone-RateLimit-*` headers reporting the limit, what remains in the window, and the policy in force. Your client should read those rather than inventing a backoff, and — more importantly — your application has to decide what a 429 means in product terms. For an interactive feature that is usually a specific message and possibly an upgrade path, not a stack trace. For a background job it is a signal to defer. ## Where this sits relative to everything else This is a gateway-side quota, which has three implications worth stating in an interview. First, it requires the proxy integration — in async logging mode Helicone learns about the call only after your application has already made and paid for it, so there is nothing to block. Second, it only sees traffic that goes through Helicone; a service that calls the provider directly, a batch job with its own credentials, or a developer's laptop is invisible to the policy, so the quota is a partial view of spend rather than a hard ceiling. Third, it composes awkwardly with retries: a policy that trips will return 429s, and gateway retries treat 429 as retryable, so an unthinking configuration can spend its retry budget hammering its own quota. ## Rollout advice Quotas are the one feature here that can take a product down, because the failure mode is refusing your own customers' traffic. Introduce a policy in a low environment first, watch how many requests would have been rejected, and pick the number from measured p99 usage rather than from a round figure that felt safe. Send the policy from a config value rather than hard-coding it, so raising a limit for one large customer is a deploy-free change. And keep the gateway quota as a coarse backstop against runaway spend, not as the mechanism your billing model depends on — the latter needs to live somewhere you own end to end.
- What does the policy 500;w=60 actually limit, with no u or s given?Five hundred requests per sixty seconds, counted across your entire organization in one shared bucket. The unit defaults to requests and the segment defaults to global. That is a useful blast-radius guard, but it is not per-customer fairness: one runaway tenant can consume the whole allowance and every other customer sees 429s.
- Why might a cents-denominated quota still let a customer overspend?Cost is only known once the response comes back with its token usage, so the quota is evaluated against completed calls. A single request with a very large context or a long generation can cross the line in one shot, and concurrent in-flight calls can all be admitted before any of them has been counted. It bounds sustained spend, not a single spike.
- How do you apply a quota per workspace rather than per end user?Segment on a custom property. If you already attach a Helicone-Property-* header naming the workspace for analytics, put that property's name in the policy's `s` field and each distinct value gets its own bucket. That reuses an identifier you are already sending rather than overloading the user id with a tenant concept.
saying these in an interview costs you the question
- Thinks the w value is milliseconds rather than seconds
- Assumes an unsegmented policy is per user by default
- Sets s=user without sending Helicone-User-Id
- Believes a cents quota pre-authorizes spend before the call
- Treats the gateway quota as a complete ceiling on organizational spend