How is an OpenAI API call billed, and why do long chats cost more?
answer
- Metered per token, not per request
- Two rates, and they differ
- Output side is the expensive one
- The API remembers nothing between calls
- Whole transcript re-sent, re-billed each turn
basics
~20 sOpenAI bills per token, quoted per million tokens, and output tokens cost several times more than input tokens. The API is stateless, so every turn resends the whole conversation as input — the input bill grows with each turn.
solid answer
~40 sBilling is per token in two buckets: input (the prompt you send) and output (what the model generates), each with its own per-million rate, and output is typically a few times more expensive than input. Every response carries a `usage` object with `prompt_tokens`, `completion_tokens` and `total_tokens`, which is the number you should log rather than guessing from character counts. The Chat Completions API keeps no server-side memory of the conversation, so to give the model context you resend the entire message history on every call. That makes turn 10 pay for turns 1 through 9 all over again, and cost grows roughly quadratically with conversation length. The usual controls are trimming or summarising old turns, capping the completion length, and picking a smaller model for cheap traffic.
go deeper
Be able to say plainly that billing is per token, that input and output have different prices with output higher, and that resending the conversation each turn is why long chats cost more.
Explain the usage object field by field and where the token count actually comes from, including framing tokens and tool schemas, and describe history trimming or summarisation as the standard control.
Show that you instrument cost per request in production from the usage object, including the streaming case, and that you treat prompt size as a recurring cost driver rather than a one-off design choice.
Own the economics: which traffic classes justify the flagship model, what the cost per served user is, and how prompt design, caching and history policy shift unit economics as volume grows.
## What is actually metered OpenAI does not bill per request, per character, or per second of compute. It bills per **token** — a sub-word chunk produced by the model's tokenizer. English text averages roughly 3–4 characters per token, so "approximately" is one token and a rare proper noun may be three or four. Code, JSON and non-Latin scripts tokenize less efficiently, which is why the same 1,000 characters can cost noticeably more in one language than another. Tokens are metered in two separate buckets with two separate prices: - **Input (prompt) tokens** — everything you send: system message, every prior user and assistant turn, tool/function JSON schemas, and the token cost of any images or files attached to the request. - **Output (completion) tokens** — everything the model generates, including tool-call arguments it emits, and, on reasoning models, the hidden reasoning tokens. Prices are published per **1 million tokens** per model, and output is consistently the more expensive side — commonly around three to five times the input rate for the same model. That asymmetry is the single most useful thing to remember when estimating cost: a long prompt with a short answer is usually cheap; a short prompt that produces pages of output is not. ## Where to read the real numbers Every non-streaming response includes a `usage` object: - `usage.prompt_tokens` — what you were charged for on the input side. - `usage.completion_tokens` — what you were charged for on the output side. - `usage.total_tokens` — their sum, a convenience field, not a separately-priced quantity. On newer models this object also carries breakdowns: `usage.prompt_tokens_details.cached_tokens` shows how much of the prompt was served from OpenAI's automatic prompt cache and therefore billed at a discounted input rate, and `usage.completion_tokens_details.reasoning_tokens` shows the hidden reasoning a reasoning model produced — billed as output even though you never see the text. Log these fields per request. Estimating cost from string lengths is a classic source of forecast error, because it ignores framing tokens, tool schemas, images and reasoning. ## Why the bill climbs across a conversation Chat Completions is **stateless**. The server does not remember the previous call; the only reason the model appears to remember turn 1 at turn 10 is that your client sent turn 1 again inside the `messages` array. So the input bill for turn *n* covers the whole transcript up to *n*: - turn 1 pays for the system prompt plus one user message; - turn 5 pays for the system prompt plus nine messages; - turn 20 pays for the system prompt plus thirty-nine messages. Summed over the session, input cost grows roughly with the square of the number of turns. A long-running assistant thread with a big system prompt and tool schemas can spend far more on re-sent history than on the answers it produces. ## The standard controls 1. **Trim the history.** Keep the system message plus the last N turns, or replace older turns with a short model-written summary. This is a product decision as much as a cost one — you are choosing what the assistant is allowed to forget. 2. **Cap the output.** Set a completion cap appropriate to the task; an unbounded cap invites the model to produce more expensive output than the feature needs. 3. **Right-size the model.** Classification, extraction and routing rarely need the flagship model; a smaller model in the same family is often an order of magnitude cheaper per token. 4. **Exploit caching.** A stable prefix (system prompt, tool schemas, a fixed document) that repeats across requests can be served from the automatic prompt cache at a reduced input rate — which is a reason to put the volatile part of the prompt last, not first. 5. **Batch the offline work.** Bulk jobs that tolerate a delay can go through the Batch API, which is billed at a discount relative to synchronous calls. ## The mental model Think of each request as: *cost = (input tokens × input rate) + (output tokens × output rate)*, where input tokens include everything you decided to re-send. Once you internalise that the transcript is re-billed on every turn, the design consequences follow naturally: prompts are a recurring cost, not a fixed one, and history management is a cost lever.
- What does the cached_tokens field inside prompt_tokens_details tell you about the bill?It reports how many of that request's input tokens were served from OpenAI's automatic prompt cache. Cached input is billed at a lower rate than fresh input, so a high cached_tokens share on a request with a large stable prefix means the repeated part of your prompt is costing less than list price. It is also a diagnostic: if it stays at zero on requests you expected to hit, something before the stable prefix is changing.
- A reasoning model returns a two-sentence answer but a huge bill. What happened?Reasoning models generate hidden reasoning tokens before the visible answer, and those are billed as output tokens at the output rate. Check usage.completion_tokens_details.reasoning_tokens: it is often many times the visible answer. It also counts against the completion cap you set, so a cap that is too low can consume the whole budget on reasoning and return an empty or truncated answer.
- How do you get usage numbers when you are streaming the response?A streamed Chat Completions call does not include a usage object by default. Set stream_options with include_usage true, and the API sends a final chunk carrying the usage totals after the content chunks. Without it, teams commonly end up with a cost blind spot on exactly their highest-traffic path, because streaming is what interactive features use.
saying these in an interview costs you the question
- Thinking the API remembers history server-side, so old turns are free
- Assuming input and output tokens cost the same rate
- Estimating cost from character or word counts instead of usage
- Believing pricing is per request rather than per token
- Ignoring that tool schemas and images consume input tokens