In the Anthropic Messages API usage object, what does input_tokens count?
answer
- Not just the newest turn
- Everything the server had to read
- Grows because nothing is stored
- Hidden reasoning still bills as output
- Cached reads live in other fields
basics
~20 sinput_tokens counts everything the server read for that one request — system prompt, tool definitions and the entire resent message history — not just the newest user turn. output_tokens counts what the model generated in that response.
solid answer
~40 sEvery Messages API response carries a `usage` object, and `input_tokens` is the billed size of the **whole request**: the top-level `system` prompt, any tool definitions, and every turn of the `messages` array you resent — not the newest message alone. `output_tokens` is everything the model generated this turn, including reasoning tokens, and is what `max_tokens` caps. Because the endpoint is stateless and you resend the accumulated history on each call, `input_tokens` grows turn by turn, so a long conversation's cumulative input cost grows roughly with the square of its length. Tokens served from a prompt cache are reported in their own fields rather than folded into `input_tokens`, so a naive sum of `input_tokens` alone will not reconcile against a bill when caching is in play.
code
python · 14 linesresponse = client.messages.create(
model="claude-opus-4-6",
max_tokens=1024,
system=SYSTEM_PROMPT,
messages=history,
)
log.info(
"llm_call",
model=response.model,
stop_reason=response.stop_reason,
input_tokens=response.usage.input_tokens, # system + tools + full history
output_tokens=response.usage.output_tokens, # generated this turn
)go deeper
Recall that every response carries a usage object with input_tokens and output_tokens, reported by the server rather than estimated by you.
Explain that input_tokens covers the system prompt, tool definitions and the entire resent history, and that output_tokens includes reasoning tokens and is what max_tokens caps.
Demonstrate the operating view: log usage per request, aggregate per conversation, recognise the quadratic input curve of a stateless endpoint, and know that cached reads are reported separately from input_tokens.
Own the economics — unit cost per session rather than per call, where context compaction pays for itself, and how usage telemetry feeds budget alerts and model-selection decisions across teams.
## Where the numbers come from Every response from `POST /v1/messages` includes a `usage` object. The two fields present on every call are: - **`input_tokens`** — how many tokens the server read to produce this response. - **`output_tokens`** — how many tokens the model generated in this response. They are reported *after* the fact by the server, using the model's own tokenizer, which makes them the authoritative figures. Any client-side estimate is an approximation; if you need a number before sending, the API offers a dedicated token-counting endpoint rather than asking you to guess. ## input_tokens is the whole request The single most common misreading is that `input_tokens` reflects the latest user message. It does not. It covers everything the model was given: - the top-level `system` prompt, - the tool definitions, if any, - **every turn in `messages`**, all the way back to the first one, - any non-text content those turns carry, converted to its token cost. This follows directly from the endpoint being stateless. There is no server-side conversation, so on turn twenty you are re-sending turns one through nineteen, and you are billed for reading them again. ## The quadratic curve Suppose each turn adds roughly `k` tokens to the history. Turn one reads `k`, turn two reads `2k`, turn three `3k`. After `n` turns the cumulative input read is proportional to `n²·k/2`. Output cost, by contrast, grows linearly. So in a long-running chat or agent loop, **input dominates**, and it dominates in a way that surprises anyone who budgeted linearly from a single-turn measurement. This is the number that justifies context management: trimming or summarising old turns, and marking a stable prefix so repeated reads of unchanged content are billed at a different rate. It is also why per-request usage should be aggregated per *conversation*, not just per call — a service whose average request looks cheap can still have expensive sessions. ## output_tokens `output_tokens` counts everything generated on this turn. Two things to internalise: - **Reasoning tokens are output tokens.** On models that produce them they are generated, billed, and counted against `max_tokens` even when their text is not shown to you. A response whose visible text is short but whose `output_tokens` is large is usually explained by this, not by a bug. - **It is what your ceiling caps.** `max_tokens` bounds this number, so `output_tokens == max_tokens` alongside `stop_reason: "max_tokens"` is the truncation signature. Output tokens are typically priced several times higher per token than input tokens, so the cheap-looking half of the request is often the expensive half per unit. ## Cached tokens are separate When prompt caching is in use, tokens served from cache are reported in their own `usage` fields; `input_tokens` reports the uncached remainder. Practical consequence: if you sum only `input_tokens` across a caching workload you will under-count what was actually read. Reconcile against all the token fields the object exposes, not the two you first learned. ## Using usage well - **Record it per request with the request id.** Usage plus stop reason plus latency is the minimum viable telemetry for an LLM call, and it is free — it comes back on every response. - **Aggregate by conversation and by route.** Per-call averages hide the quadratic tail. - **Convert to money at the edge.** Rates differ between input and output, and between models, so store raw counts and the model id and apply pricing downstream; storing a computed cost bakes in a rate that will change. - **Alert on shape, not just size.** A sudden jump in mean `input_tokens` usually means history trimming stopped working or a retrieval step started injecting far more context than intended — both are visible in this field long before they show up on an invoice. ## Streaming A streamed response reports the same accounting; the input side is known at the start of the stream and the final output count arrives with the terminal events, and SDK helpers that assemble a final message expose the completed `usage` object in the same place as a non-streamed call.
- Why does the input cost of a chat session grow faster than the number of turns?Because the endpoint is stateless: every request resends the accumulated history, so turn n reads roughly n times the per-turn content. Summed over a session that is quadratic in the turn count, while output cost stays linear. Long agent loops are therefore dominated by re-reading their own transcript, which is what context trimming and prefix reuse exist to attack.
- You see a short answer but a large output_tokens figure. What is the likely explanation?Reasoning tokens. On models that produce internal reasoning, those tokens are generated output — billed, counted in `output_tokens`, and charged against `max_tokens` — even when their text is not surfaced. It is expected behaviour rather than a defect; if the ratio is a problem, reduce how much reasoning you request rather than shrinking the ceiling.
- How would you estimate token cost before sending a request rather than after?Use the API's token-counting endpoint, which runs the same tokenizer the model uses against your exact request body. Character-count heuristics and other vendors' tokenizers both drift, and the drift is worst on the content that matters — code, non-Latin scripts and structured payloads. Measure with the real thing, then apply the model's published rates.
Each call is like re-faxing the entire case file to get one more paragraph of advice: you pay to have the whole file read again every single time.
saying these in an interview costs you the question
- Thinks input_tokens covers only the latest message
- Assumes usage accumulates across a conversation server-side
- Believes hidden reasoning tokens are not billed
- Estimates tokens by dividing characters by four
- Sums input_tokens alone when caching is enabled