How do you forecast and cap OpenAI spend before shipping a new feature?
answer
- Measure first, price second
- Averages hide the expensive tail
- One key per feature, or no attribution
- Dollars lag, tokens lead
- Decide the fallback before launch
basics
~20 sBuild a unit-cost model from a pilot: measured p50 and p95 input and output tokens per request, multiplied by the model's per-million rates and forecast volume. Then enforce it with per-project budgets, separate keys per feature for attribution, and alerts on tokens per request.
solid answer
~50 sForecasting starts with measurement, not list price. Run a realistic pilot and record the `usage` object per request — input, output, cached input, and reasoning tokens on reasoning models — then take p50 and p95 rather than averages, because prompt sizes are long-tailed. Cost per request is tokens times the model's per-million rates; multiply by forecast volume per user and adoption to get a monthly range, and present a range, not a point estimate. Control comes from three places: structural levers (smaller model for cheap traffic, prompt caching on stable prefixes, the Batch API's discount for offline work, honest completion caps), hard guardrails (per-project usage limits and organisation budgets in the dashboard, plus a distinct API key or project per feature so spend is attributable), and monitoring — the Usage and Costs endpoints with an admin key for chargeback, and an alert on tokens-per-request drift, which moves before the invoice does.
code
python · 21 linesfrom openai import OpenAI
client = OpenAI()
stream = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Explain idempotency in one paragraph."}],
max_completion_tokens=300,
stream=True,
stream_options={"include_usage": True},
)
usage = None
for chunk in stream:
if chunk.usage is not None:
usage = chunk.usage # final chunk carries the totals
elif chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
print()
print(usage.prompt_tokens, usage.completion_tokens)go deeper
Know that spend is driven by tokens per request times request volume, and that the usage object on responses is where the real numbers come from.
Build the unit-cost model: measure input, output, cached and reasoning tokens from a pilot, apply the model's per-million rates, and name the levers that reduce each term.
Show operational control: streaming usage capture, per-project keys for attribution, budgets and usage limits as backstops, and alerting on tokens-per-request drift rather than on the invoice.
Own the economics end to end — unit cost per served user, which traffic classes justify the flagship model, what degrades if cost doubles, and how attribution and budgets are structured so the answer is a decision rather than an incident.
## Start from measured unit cost A credible forecast has one required input that a pricing page cannot give you: how many tokens **your** feature actually consumes per request. Get it by piloting the real prompt against real data and logging the `usage` object every time: - `prompt_tokens` and `completion_tokens` — the two priced quantities. - `prompt_tokens_details.cached_tokens` — the share of input billed at the discounted cached rate. - `completion_tokens_details.reasoning_tokens` — on reasoning models, hidden output that is billed at the output rate and often exceeds the visible answer. If the pilot streams, remember that a streamed Chat Completions call returns no usage object unless you set `stream_options` with `include_usage` true. Without it your highest-traffic path is precisely the one you cannot measure. Use **percentiles, not averages**. Prompt sizes in retrieval-augmented features are long-tailed: a mean hides that the top 5% of requests may cost ten times the median. Model cost as p50 for the typical bill and p95 for the exposure, and forecast a range. ## Assemble the model Cost per request = (input tokens × input rate) + (cached input × cached rate) + (output tokens × output rate), with rates quoted per million tokens per model. Then: *monthly cost ≈ cost per request × requests per active user per month × active users × adoption rate* The volume terms are where forecasts really go wrong — token counts are measurable, adoption is a guess. Say so explicitly, give a low/expected/high band, and state which assumption dominates the spread. That honesty is the difference between a forecast a business can plan against and a number that gets quoted back at you in a postmortem. ## Structural levers, applied before guardrails 1. **Model tiering.** Route classification, routing and extraction to a small model and reserve the flagship for work that demonstrably needs it. This is usually an order of magnitude, not a percentage. 2. **Prompt caching.** A stable prefix — system prompt, tool schemas, a fixed policy document — repeated across requests is billed at a reduced input rate; order the prompt so the volatile part comes last. 3. **Batch API.** Offline and delay-tolerant work runs at a discount and on a separate capacity allowance. 4. **Completion caps.** A realistic cap bounds the expensive side of the bill and stops a runaway generation from turning one request into a hundred requests' worth of output. 5. **Prompt hygiene.** Fewer retrieved chunks, trimmed history, no dead instructions. Prompt size is a recurring cost paid on every call, forever. ## Guardrails that actually stop spend - **Separate project (and key) per feature.** Attribution is impossible after the fact if everything shares one key. Per-project scoping also lets you set independent rate and usage limits, so a runaway job in one feature cannot consume another's budget. - **Budgets and usage limits.** Organisation-level monthly budgets with notification thresholds, and per-project usage limits, are the backstop for the case your model did not predict — a bug, an abusive user, a retry storm. - **Application-level quotas.** Per-user or per-tenant token budgets enforced in your own code are the only control that acts at the right granularity; the vendor's limits protect the company, not a single customer's experience. ## Monitor the leading indicator Dollars are a lagging indicator; **tokens per request** is the leading one. When someone adds three tools to the catalogue, raises the retrieval count from five chunks to fifteen, or lets conversation history grow unbounded, tokens per request steps up the day it ships and the invoice reflects it weeks later. Track it per feature, alert on step changes, and review it in the same place you review latency. For accounting, OpenAI exposes organisation-level usage and cost endpoints queried with an **admin key**, which is what you use for chargeback and dashboards rather than scraping the console. Reconcile your own logged token totals against them periodically; a persistent gap means something is calling the API outside your instrumented path. ## Decide the levers before you need them The strategic part is deciding, in advance, what you would sacrifice if cost doubled: which traffic classes drop to a smaller model, what quality degradation is acceptable, whether some capability becomes a paid tier of your own product. A feature whose unit economics are only understood after launch has no options left except turning it off. On some models OpenAI also offers alternative service tiers trading latency for price, which is worth checking for the specific model you have chosen when the workload tolerates slower responses.
- Why present a range rather than a single monthly cost number?Because the two inputs have very different confidence. Tokens per request you can measure to within a few percent from a pilot; volume and adoption are estimates that can be off by multiples. Give low/expected/high with the assumption that dominates the spread named explicitly, and say which lever you would pull if actuals track the high line. A single number invites false precision and gets quoted back at you.
- How do you make spend attributable across several features sharing one organisation?Give each feature its own project and API key. Project scoping carries independent rate and usage limits and makes the organisation usage and cost endpoints — queried with an admin key — able to report per-project spend without you reverse-engineering it from application logs. Add your own per-tenant token accounting on top, because vendor-side granularity stops at the project boundary.
- Costs jump 40% with flat request volume. Where do you look first?Tokens per request, split by feature. Flat volume with rising cost means each call got more expensive: a bigger system prompt, more retrieved chunks, an enlarged tool catalogue, unbounded conversation history, or a model change. Also check the cached-token share — if a prefix stopped being stable, previously discounted input is now billed at the full rate, which can move the bill sharply with no visible code change.
saying these in an interview costs you the question
- Forecasting from list price without measuring real token usage
- Using mean tokens per request and ignoring the tail
- One shared API key, so no feature-level attribution
- Watching only the invoice instead of tokens per request
- Assuming streamed calls report usage by default