skip to content

In the Anthropic Messages API usage object, what does input_tokens count?

level: seniorimportance: should knowfreq 45%

answer

  1. Not just the newest turn
  2. Everything the server had to read
  3. Grows because nothing is stored
  4. Hidden reasoning still bills as output
  5. Cached reads live in other fields

basics

~20 s

input_tokens counts everything the server read for that one request — system prompt, tool definitions and the entire resent message history — not just the newest user turn. output_tokens counts what the model generated in that response.

solid answer

~40 s

Every Messages API response carries a `usage` object, and `input_tokens` is the billed size of the **whole request**: the top-level `system` prompt, any tool definitions, and every turn of the `messages` array you resent — not the newest message alone. `output_tokens` is everything the model generated this turn, including reasoning tokens, and is what `max_tokens` caps. Because the endpoint is stateless and you resend the accumulated history on each call, `input_tokens` grows turn by turn, so a long conversation's cumulative input cost grows roughly with the square of its length. Tokens served from a prompt cache are reported in their own fields rather than folded into `input_tokens`, so a naive sum of `input_tokens` alone will not reconcile against a bill when caching is in play.

code

python · 14 lines
python
response = client.messages.create(
    model="claude-opus-4-6",
    max_tokens=1024,
    system=SYSTEM_PROMPT,
    messages=history,
)

log.info(
    "llm_call",
    model=response.model,
    stop_reason=response.stop_reason,
    input_tokens=response.usage.input_tokens,    # system + tools + full history
    output_tokens=response.usage.output_tokens,  # generated this turn
)

go deeper

for a junior

Recall that every response carries a usage object with input_tokens and output_tokens, reported by the server rather than estimated by you.

for a middle

Explain that input_tokens covers the system prompt, tool definitions and the entire resent history, and that output_tokens includes reasoning tokens and is what max_tokens caps.

for a senior

Demonstrate the operating view: log usage per request, aggregate per conversation, recognise the quadratic input curve of a stateless endpoint, and know that cached reads are reported separately from input_tokens.

for a principal

Own the economics — unit cost per session rather than per call, where context compaction pays for itself, and how usage telemetry feeds budget alerts and model-selection decisions across teams.

## Where the numbers come from Every response from `POST /v1/messages` includes a `usage` object. The two fields present on every call are: - **`input_tokens`** — how many tokens the server read to produce this response. - **`output_tokens`** — how many tokens the model generated in this response. They are reported *after* the fact by the server, using the model's own tokenizer, which makes them the authoritative figures. Any client-side estimate is an approximation; if you need a number before sending, the API offers a dedicated token-counting endpoint rather than asking you to guess. ## input_tokens is the whole request The single most common misreading is that `input_tokens` reflects the latest user message. It does not. It covers everything the model was given: - the top-level `system` prompt, - the tool definitions, if any, - **every turn in `messages`**, all the way back to the first one, - any non-text content those turns carry, converted to its token cost. This follows directly from the endpoint being stateless. There is no server-side conversation, so on turn twenty you are re-sending turns one through nineteen, and you are billed for reading them again. ## The quadratic curve Suppose each turn adds roughly `k` tokens to the history. Turn one reads `k`, turn two reads `2k`, turn three `3k`. After `n` turns the cumulative input read is proportional to `n²·k/2`. Output cost, by contrast, grows linearly. So in a long-running chat or agent loop, **input dominates**, and it dominates in a way that surprises anyone who budgeted linearly from a single-turn measurement. This is the number that justifies context management: trimming or summarising old turns, and marking a stable prefix so repeated reads of unchanged content are billed at a different rate. It is also why per-request usage should be aggregated per *conversation*, not just per call — a service whose average request looks cheap can still have expensive sessions. ## output_tokens `output_tokens` counts everything generated on this turn. Two things to internalise: - **Reasoning tokens are output tokens.** On models that produce them they are generated, billed, and counted against `max_tokens` even when their text is not shown to you. A response whose visible text is short but whose `output_tokens` is large is usually explained by this, not by a bug. - **It is what your ceiling caps.** `max_tokens` bounds this number, so `output_tokens == max_tokens` alongside `stop_reason: "max_tokens"` is the truncation signature. Output tokens are typically priced several times higher per token than input tokens, so the cheap-looking half of the request is often the expensive half per unit. ## Cached tokens are separate When prompt caching is in use, tokens served from cache are reported in their own `usage` fields; `input_tokens` reports the uncached remainder. Practical consequence: if you sum only `input_tokens` across a caching workload you will under-count what was actually read. Reconcile against all the token fields the object exposes, not the two you first learned. ## Using usage well - **Record it per request with the request id.** Usage plus stop reason plus latency is the minimum viable telemetry for an LLM call, and it is free — it comes back on every response. - **Aggregate by conversation and by route.** Per-call averages hide the quadratic tail. - **Convert to money at the edge.** Rates differ between input and output, and between models, so store raw counts and the model id and apply pricing downstream; storing a computed cost bakes in a rate that will change. - **Alert on shape, not just size.** A sudden jump in mean `input_tokens` usually means history trimming stopped working or a retrieval step started injecting far more context than intended — both are visible in this field long before they show up on an invoice. ## Streaming A streamed response reports the same accounting; the input side is known at the start of the stream and the final output count arrives with the terminal events, and SDK helpers that assemble a final message expose the completed `usage` object in the same place as a non-streamed call.

  • Why does the input cost of a chat session grow faster than the number of turns?
    Because the endpoint is stateless: every request resends the accumulated history, so turn n reads roughly n times the per-turn content. Summed over a session that is quadratic in the turn count, while output cost stays linear. Long agent loops are therefore dominated by re-reading their own transcript, which is what context trimming and prefix reuse exist to attack.
  • You see a short answer but a large output_tokens figure. What is the likely explanation?
    Reasoning tokens. On models that produce internal reasoning, those tokens are generated output — billed, counted in `output_tokens`, and charged against `max_tokens` — even when their text is not surfaced. It is expected behaviour rather than a defect; if the ratio is a problem, reduce how much reasoning you request rather than shrinking the ceiling.
  • How would you estimate token cost before sending a request rather than after?
    Use the API's token-counting endpoint, which runs the same tokenizer the model uses against your exact request body. Character-count heuristics and other vendors' tokenizers both drift, and the drift is worst on the content that matters — code, non-Latin scripts and structured payloads. Measure with the real thing, then apply the model's published rates.

Each call is like re-faxing the entire case file to get one more paragraph of advice: you pay to have the whole file read again every single time.

saying these in an interview costs you the question

  • Thinks input_tokens covers only the latest message
  • Assumes usage accumulates across a conversation server-side
  • Believes hidden reasoning tokens are not billed
  • Estimates tokens by dividing characters by four
  • Sums input_tokens alone when caching is enabled

context