skip to content

Rate Limits & Cost

What breaks once the feature is live: RPM/TPM/RPD limits per tier, 429s, retries with exponential backoff, counting tokens with tiktoken before you send, and predicting spend per model. This is the part that separates people who shipped an LLM feature from people who demoed one.

on this pageshow

questions

6

How is an OpenAI API call billed, and why do long chats cost more?

level: juniorimportance: must knowfreq 72%

answer

  1. Metered per token, not per request
  2. Two rates, and they differ
  3. Output side is the expensive one
  4. The API remembers nothing between calls
  5. Whole transcript re-sent, re-billed each turn

basics

~20 s

OpenAI bills per token, quoted per million tokens, and output tokens cost several times more than input tokens. The API is stateless, so every turn resends the whole conversation as input — the input bill grows with each turn.

solid answer

~40 s

Billing is per token in two buckets: input (the prompt you send) and output (what the model generates), each with its own per-million rate, and output is typically a few times more expensive than input. Every response carries a `usage` object with `prompt_tokens`, `completion_tokens` and `total_tokens`, which is the number you should log rather than guessing from character counts. The Chat Completions API keeps no server-side memory of the conversation, so to give the model context you resend the entire message history on every call. That makes turn 10 pay for turns 1 through 9 all over again, and cost grows roughly quadratically with conversation length. The usual controls are trimming or summarising old turns, capping the completion length, and picking a smaller model for cheap traffic.

go deeper

for a junior

Be able to say plainly that billing is per token, that input and output have different prices with output higher, and that resending the conversation each turn is why long chats cost more.

for a middle

Explain the usage object field by field and where the token count actually comes from, including framing tokens and tool schemas, and describe history trimming or summarisation as the standard control.

for a senior

Show that you instrument cost per request in production from the usage object, including the streaming case, and that you treat prompt size as a recurring cost driver rather than a one-off design choice.

for a principal

Own the economics: which traffic classes justify the flagship model, what the cost per served user is, and how prompt design, caching and history policy shift unit economics as volume grows.

## What is actually metered OpenAI does not bill per request, per character, or per second of compute. It bills per **token** — a sub-word chunk produced by the model's tokenizer. English text averages roughly 3–4 characters per token, so "approximately" is one token and a rare proper noun may be three or four. Code, JSON and non-Latin scripts tokenize less efficiently, which is why the same 1,000 characters can cost noticeably more in one language than another. Tokens are metered in two separate buckets with two separate prices: - **Input (prompt) tokens** — everything you send: system message, every prior user and assistant turn, tool/function JSON schemas, and the token cost of any images or files attached to the request. - **Output (completion) tokens** — everything the model generates, including tool-call arguments it emits, and, on reasoning models, the hidden reasoning tokens. Prices are published per **1 million tokens** per model, and output is consistently the more expensive side — commonly around three to five times the input rate for the same model. That asymmetry is the single most useful thing to remember when estimating cost: a long prompt with a short answer is usually cheap; a short prompt that produces pages of output is not. ## Where to read the real numbers Every non-streaming response includes a `usage` object: - `usage.prompt_tokens` — what you were charged for on the input side. - `usage.completion_tokens` — what you were charged for on the output side. - `usage.total_tokens` — their sum, a convenience field, not a separately-priced quantity. On newer models this object also carries breakdowns: `usage.prompt_tokens_details.cached_tokens` shows how much of the prompt was served from OpenAI's automatic prompt cache and therefore billed at a discounted input rate, and `usage.completion_tokens_details.reasoning_tokens` shows the hidden reasoning a reasoning model produced — billed as output even though you never see the text. Log these fields per request. Estimating cost from string lengths is a classic source of forecast error, because it ignores framing tokens, tool schemas, images and reasoning. ## Why the bill climbs across a conversation Chat Completions is **stateless**. The server does not remember the previous call; the only reason the model appears to remember turn 1 at turn 10 is that your client sent turn 1 again inside the `messages` array. So the input bill for turn *n* covers the whole transcript up to *n*: - turn 1 pays for the system prompt plus one user message; - turn 5 pays for the system prompt plus nine messages; - turn 20 pays for the system prompt plus thirty-nine messages. Summed over the session, input cost grows roughly with the square of the number of turns. A long-running assistant thread with a big system prompt and tool schemas can spend far more on re-sent history than on the answers it produces. ## The standard controls 1. **Trim the history.** Keep the system message plus the last N turns, or replace older turns with a short model-written summary. This is a product decision as much as a cost one — you are choosing what the assistant is allowed to forget. 2. **Cap the output.** Set a completion cap appropriate to the task; an unbounded cap invites the model to produce more expensive output than the feature needs. 3. **Right-size the model.** Classification, extraction and routing rarely need the flagship model; a smaller model in the same family is often an order of magnitude cheaper per token. 4. **Exploit caching.** A stable prefix (system prompt, tool schemas, a fixed document) that repeats across requests can be served from the automatic prompt cache at a reduced input rate — which is a reason to put the volatile part of the prompt last, not first. 5. **Batch the offline work.** Bulk jobs that tolerate a delay can go through the Batch API, which is billed at a discount relative to synchronous calls. ## The mental model Think of each request as: *cost = (input tokens × input rate) + (output tokens × output rate)*, where input tokens include everything you decided to re-send. Once you internalise that the transcript is re-billed on every turn, the design consequences follow naturally: prompts are a recurring cost, not a fixed one, and history management is a cost lever.

  • What does the cached_tokens field inside prompt_tokens_details tell you about the bill?
    It reports how many of that request's input tokens were served from OpenAI's automatic prompt cache. Cached input is billed at a lower rate than fresh input, so a high cached_tokens share on a request with a large stable prefix means the repeated part of your prompt is costing less than list price. It is also a diagnostic: if it stays at zero on requests you expected to hit, something before the stable prefix is changing.
  • A reasoning model returns a two-sentence answer but a huge bill. What happened?
    Reasoning models generate hidden reasoning tokens before the visible answer, and those are billed as output tokens at the output rate. Check usage.completion_tokens_details.reasoning_tokens: it is often many times the visible answer. It also counts against the completion cap you set, so a cap that is too low can consume the whole budget on reasoning and return an empty or truncated answer.
  • How do you get usage numbers when you are streaming the response?
    A streamed Chat Completions call does not include a usage object by default. Set stream_options with include_usage true, and the API sends a final chunk carrying the usage totals after the content chunks. Without it, teams commonly end up with a cost blind spot on exactly their highest-traffic path, because streaming is what interactive features use.

saying these in an interview costs you the question

  • Thinking the API remembers history server-side, so old turns are free
  • Assuming input and output tokens cost the same rate
  • Estimating cost from character or word counts instead of usage
  • Believing pricing is per request rather than per token
  • Ignoring that tool schemas and images consume input tokens

context

open as a page

How do you handle an OpenAI API 429 rate_limit_exceeded error?

level: middleimportance: must knowfreq 80%

basics

~20 s

Retry it with exponential backoff plus random jitter and a capped attempt count, using the x-ratelimit-reset headers to time the wait. First check the error code: rate_limit_exceeded is transient and worth retrying, while insufficient_quota is a billing problem no retry will clear.

open as a page

Why can an OpenAI call return 429 on TPM while RPM is barely used?

level: middleimportance: must knowfreq 63%

basics

~20 s

OpenAI enforces several limits at once — requests per minute, tokens per minute, and daily caps — per model and per project. Exceeding any single one returns 429, so a handful of very large prompts can exhaust the token budget while the request count stays trivial.

open as a page

How do you count an OpenAI request's tokens with tiktoken before sending?

level: middleimportance: should knowfreq 52%

basics

~20 s

Get the encoding for your model with tiktoken.encoding_for_model, then take len(encoding.encode(text)) for each message. Add the per-message framing overhead, because raw text encoding alone undercounts a chat request and ignores tool schemas and images entirely.

open as a page

Your live feature keeps hitting OpenAI 429s at peak — how do you fix it?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Stop relying on retries and add admission control: a shared token-bucket limiter sized to the tightest limit, so work waits in your queue instead of bouncing off the API. Then reduce demand — realistic completion caps, smaller models for cheap traffic, bulk work moved to the Batch API.

open as a page

How do you forecast and cap OpenAI spend before shipping a new feature?

level: principalimportance: should knowfreq 40%

basics

~20 s

Build a unit-cost model from a pilot: measured p50 and p95 input and output tokens per request, multiplied by the model's per-million rates and forecast volume. Then enforce it with per-project budgets, separate keys per feature for attribution, and alerts on tokens per request.

open as a page