How would you attribute and cap OpenRouter spend per team on one shared account?
answer
- One pool means one blast radius
- Keys are the budget boundary
- Provision keys with their own limits
- The gateway cannot see your tenants
- Alarm on runway, not dollars
basics
~10 sIssue a separate OpenRouter API key per team through the key-provisioning API, each with its own credit limit, and log the per-request cost from usage accounting against your own tenant and feature dimensions.
solid answer
~50 sOne prepaid balance is convenient until one workload drains it for everyone, so the design has two halves. **Containment:** use OpenRouter's key-provisioning API to mint a runtime key per team, service or environment, each carrying its own credit limit; when a key exhausts its limit only that key's calls fail, and the shared pool survives. `GET /api/v1/auth/key` reports the calling key's label, usage and limit, so each service can watch its own runway. **Attribution:** the gateway knows the key, not your business dimensions, so per-feature and per-tenant attribution can only live in your telemetry. Enable usage accounting, and log `usage.cost` with the generation id, model slug, team, feature and tenant on every request; that record is what answers "which feature doubled our bill". Reconcile it periodically against the account's credit usage. The judgement calls are granularity, whether limits are hard stops or alerts, and which high-volume routes move onto provider contracts instead.
code
python · 26 linesimport json, logging, os, requests
log = logging.getLogger("llm.cost")
def call(messages, *, team, feature, tenant):
r = requests.post(
"https://openrouter.ai/api/v1/chat/completions",
headers={"Authorization": f"Bearer {os.environ['OPENROUTER_API_KEY']}"},
json={
"model": os.environ["OPENROUTER_MODEL"],
"messages": messages,
"usage": {"include": True},
},
timeout=60,
)
body = r.json()
usage = body.get("usage", {})
log.info(json.dumps({
"generation_id": body.get("id"),
"model": body.get("model"),
"cost": usage.get("cost"),
"prompt_tokens": usage.get("prompt_tokens"),
"completion_tokens": usage.get("completion_tokens"),
"team": team, "feature": feature, "tenant": tenant,
}))
return bodygo deeper
Know that one OpenRouter account has a single credit balance, so separate API keys with their own limits are how different teams avoid draining each other's budget.
Describe the mechanics: provision a key per team with a credit limit, read that key's usage and limit from the auth key endpoint, and log the per-request cost from usage accounting.
Design the pipeline end to end — structured cost records with tenant and feature dimensions, runway-based alarms per key, cost logged on failures too, and periodic reconciliation against account usage to find uninstrumented traffic.
Own the policy: attribution granularity versus telemetry cost, which keys are allowed to fail closed, when a workload moves onto a provider contract for better isolation and economics, and whether showback or full chargeback is the right incentive.
## The failure mode you are designing against A single prepaid balance behind one API key means every workload is a noisy neighbour to every other. A runaway retry loop in a batch job, an accidental switch to an expensive model, or an abuse spike drains the pool, and then **every** product surface fails with insufficient credits simultaneously. The blast radius is the whole company. Everything below is about shrinking that radius and knowing whose fault it was. ## Containment: keys as budget boundaries OpenRouter supports provisioning runtime API keys programmatically, and each key can carry its own credit limit. That makes the key the natural budget unit. A workable policy: - One key per **team plus environment** — production and non-production budgets should never be the same pool, because experiments are exactly where runaway spend originates. - Limits sized to a real forecast plus headroom, not to a round number, and reviewed on the same cadence as capacity planning. - Programmatic creation and rotation, so onboarding a service is a config change rather than a screenshot of a dashboard pasted into a chat. The payoff is that exhaustion is scoped: the offending key starts failing and everyone else keeps serving. That converts a company-wide outage into one team's incident, which is the single most valuable property of the whole scheme. ## Observability per key `GET /api/v1/auth/key` returns the calling key's label, usage and limit — note that it describes the key making the call, so each service can self-report its own runway without being handed account-wide credentials. Export that as a metric per service and alarm on **runway in hours at current burn**, not on dollars remaining; a fixed dollar threshold means nothing when one team's burn rate is fifty times another's. Account-level solvency still comes from the credits endpoint, and it stays a separate, page-worthy alarm. ## Attribution: only your telemetry can do it The gateway sees a key and a model. It does not know that a call belonged to the onboarding summariser for tenant 4471. That mapping exists only in your process, at the moment of the call, and if you do not write it down it is unrecoverable. So: 1. Enable usage accounting on every request. 2. On response, emit one structured record: generation id, model slug, `usage.cost`, native token counts, team, feature, tenant, environment, and outcome. 3. Aggregate that into your normal metrics stack — cost becomes just another dimension beside latency and error rate. Two refinements matter. First, **also record cost on failures and cancellations**, since partially generated tokens can still be billed and an abandoned-stream bug is invisible if you only log successes. Second, keep the generation id so you can reconcile individual records against the platform's own generation stats during a dispute or an audit. ## Reconciliation Run a periodic check: does the sum of logged per-request cost for a window match the account's credit usage over the same window? Drift means something is calling outside your instrumented path — a script with a stray key, a vendor integration, a forgotten cron. Finding that is worth more than the accounting precision itself, because unattributed spend is usually unowned spend. ## What is not attribution The attribution headers OpenRouter accepts for app identification exist for public app ranking and identification, not for billing. Do not build a chargeback system on them. Likewise, the model slug is not a proxy for a team once two teams share a model, which they will within a quarter. ## The judgement calls a lead owns - **Granularity.** Per team is cheap and coarse; per feature is where decisions actually get made; per tenant is required if you resell. Each level is more telemetry cardinality — pick deliberately rather than logging everything at maximum resolution. - **Hard stop or soft alert.** A hard per-key limit protects the pool but will, one day, take a production feature down over money. Decide in advance which keys are allowed to fail closed and which get generous limits plus paging, and make that a documented tier rather than an accident of who set the number. - **Where spend should live.** Once a workload's volume is large and stable, moving it to a bring-your-own-key route puts it on a provider contract with its own quota, discount and invoice — better economics and better isolation, at the price of a second ledger to reconcile. - **Chargeback versus showback.** Publishing per-team cost usually changes behaviour on its own; formal chargeback adds finance machinery and often buys little more. Start with showback and escalate only if the incentives fail.
- Your logged per-request costs sum to noticeably less than the account's credit usage for the same week. What does that tell you?That traffic is reaching the account outside your instrumented path — a stray key in a script, a forgotten scheduled job, a vendor integration, or a code path that skips the logging wrapper. The gap is the finding, not a rounding error: unattributed spend is almost always unowned spend. Chase it by key first, since key-level usage is visible even when your own record is missing.
- Would you make per-key credit limits a hard stop or an alerting threshold?Tier it. Batch, experimental and internal keys fail closed, because containment matters more than availability there. Customer-facing keys get generous limits plus paging, so money never silently takes a product down. Write the tiers down as policy; when limits are set ad hoc, the key that fails at 3am is invariably the one nobody meant to cap.
- Why not attribute spend using the app-identification headers OpenRouter accepts?They exist for identifying and ranking the calling app publicly, not for billing, and nothing binds them to a tenant or feature in a way you can audit. Chargeback built on them is unverifiable and easy to spoof from within your own codebase. Attribution belongs in your own structured logs, joined to the per-request cost from usage accounting and the generation id.
It is the difference between one company credit card everyone shares and issuing each team its own card with its own limit: the second still bills to the same company, but a bad week stops at one team's card.
saying these in an interview costs you the question
- One shared key for every team and environment
- Attributing cost by model slug once teams share models
- Building chargeback on app-identification headers
- Alarming on dollars remaining rather than runway
- Logging cost only on successful requests