skip to content

How does token pricing differ across Claude's Haiku, Sonnet, and Opus tiers?

level: middleimportance: must knowfreq 60%

answer

  1. Two rates per model, not one
  2. Output is the expensive side
  3. Rates climb up the tier ladder
  4. Token mix can outweigh tier choice
  5. Caching and batching discount the same model

basics

~20 s

Anthropic prices each Claude model per million tokens, with input and output billed at separate rates and output costing several times more than input. Rates climb from Haiku to Sonnet to Opus, so per-request cost depends on both the tier and the token mix.

solid answer

~50 s

Every Claude model has two published rates — one per million input tokens, one per million output tokens — and the output rate is the higher of the two, typically by around five times at each tier. The rates themselves climb up the tier ladder, so Opus input can cost several times Sonnet input, and Sonnet several times Haiku. The practical consequence is that request cost is `input_tokens x input_rate + output_tokens x output_rate`, and the token mix can matter more than the tier: a Haiku call carrying a 100K-token document can easily cost more than an Opus call with a 500-token prompt. That is why you profile actual per-request token counts from the `usage` figures the API returns before arguing about tiers. Prompt caching and the batch API discount the same model further, so "drop a tier" is only one of several cost levers.

code

python · 15 lines
python
# Rates are USD per million tokens; check Anthropic's pricing page for current values.
RATES = {
    "small": {"input": 1.00, "output": 5.00},
    "large": {"input": 5.00, "output": 25.00},
}


def request_cost(tier: str, input_tokens: int, output_tokens: int) -> float:
    r = RATES[tier]
    return (input_tokens * r["input"] + output_tokens * r["output"]) / 1_000_000


# Input-dominated call: the cheap tier is not automatically the cheap request.
print(request_cost("small", 100_000, 200))   # big prompt, tiny answer
print(request_cost("large", 500, 200))       # tiny prompt, tiny answer

go deeper

for a junior

Remember that input and output tokens are priced separately and that output is the expensive one, and that prices are quoted per million tokens rather than per request.

for a middle

Be able to compute a request's cost from the two rates and the two token counts, and explain why a large prompt on a small model can outprice a small prompt on a large one.

for a senior

Show that you profile real per-endpoint token usage before choosing a lever, and that caching, batching and turn-count reduction are all candidates alongside a tier downgrade.

for a principal

Own the unit-economics model: cost per resolved task, not cost per call, with tier choice, prompt size and loop length as tunable inputs and a quality eval as the constraint the optimisation runs against.

## The pricing shape Anthropic publishes, for each Claude model, a price per million input tokens and a separate, higher price per million output tokens. Two independent things move the number you pay: 1. **Which tier you call.** Rates rise from Haiku to Sonnet to Opus. The multiplier between adjacent tiers has historically been in the low single digits to roughly five-fold, but the exact figures change with every release and every price revision — quote the structure in an interview, not memorised numbers, and check the current pricing page before building a budget. 2. **How many tokens of each kind you spend.** Output is the expensive kind. At each tier the output rate has typically been about five times the input rate, so a change in response length moves your bill roughly five times harder than the same change in prompt length. ## Why the token mix often beats the tier The common mistake is to reason about cost as "Opus is expensive, Haiku is cheap" as if it were a per-request price. It is not. Work the arithmetic: - A retrieval-augmented call that stuffs 100,000 tokens of documents into the prompt and asks for a 200-token answer is **input-dominated**. Its cost is essentially the prompt, and dropping a tier saves you the tier ratio on that prompt — real, but the prompt itself is the problem. Retrieving fewer, better chunks may save more than the downgrade does. - An agentic loop that runs twenty turns, each producing a few thousand tokens of reasoning and tool arguments, is **output-dominated**. Here the tier ratio bites hardest, because it multiplies the already-expensive output rate. So the first move in any cost investigation is to measure, not to guess: read `usage.input_tokens` and `usage.output_tokens` off real responses, aggregate them per endpoint, and only then decide whether the lever is the tier, the prompt, or the number of turns. ## Extended thinking and hidden output Where extended thinking is enabled, the model's reasoning tokens are billed as output tokens and count against the response budget. A workload that looks input-heavy on paper can turn output-heavy once thinking is switched on, and the thinking budget becomes a first-class cost knob alongside the tier. If your bill jumped without a traffic change, a thinking-budget or prompt change is as likely a culprit as the model. ## The other cost levers on the same model Dropping a tier is not the only, or even the first, lever: - **Prompt caching** discounts repeated prefix content — long system prompts, tool definitions, a fixed document set — dramatically on cache hits, at the cost of a surcharge on the write. For a workload where every request shares a big prefix, this often beats a tier downgrade and costs you nothing in quality. - **The Message Batches API** processes asynchronous work at a discount relative to the same model called synchronously. If a job does not need an answer in real time — nightly enrichment, bulk classification, backfills — this is nearly free money. - **Prompt engineering for brevity.** Asking for structured, short outputs instead of prose cuts the expensive side of the bill directly. - **Turn count.** In agent loops, each turn resends the accumulated conversation, so cost grows superlinearly with turns. Cutting the loop from twelve turns to six can halve the bill more reliably than any model swap. ## Putting a number on a tier swap The honest way to evaluate a downgrade is a two-column estimate over real traffic: ``` monthly_cost(tier) = requests x (avg_in x in_rate(tier) + avg_out x out_rate(tier)) ``` Compute it for both tiers using your measured averages, then set the saving against the measured quality delta on a held-out eval set. If the saving is 15% of a small bill and the eval shows a two-point accuracy drop on a task that touches customers, the downgrade is a bad trade; if the saving is 70% of a large bill with no measurable delta, it is obvious. The number that makes the decision is a ratio of two measurements, never a vibe about which model is "better". ## Cascade arithmetic One subtlety catches teams that route small-first and escalate: an escalated request pays for **both** calls, plus the tokens spent re-sending the prompt to the larger model. If your escalation rate is high, the cascade costs more than calling the big model once. Cascades only pay when the cheap tier resolves the large majority of traffic and the check that decides escalation is itself cheap.

  • Your bill doubled but request volume is flat. How do you find the cause?
    Break the bill into input and output tokens per endpoint from the usage figures, and compare against the previous period. A prompt or retrieval change that grew the context shows up as input growth; a longer-answer instruction, more agent turns, or a raised thinking budget shows up as output growth. Model identifier changes are the third suspect. Only then look at the tier.
  • When is a tier downgrade the wrong first cost lever to pull?
    When the workload is input-dominated by a shared prefix — the same long system prompt, tool set or document pack on every call. Caching that prefix cuts the same cost without touching quality. Likewise for non-realtime bulk jobs, where the batch discount applies at any tier. Reach for the downgrade after the levers that cost you nothing in accuracy.
  • How does the cost of an agent loop grow with the number of turns?
    Faster than linearly. Each turn resends the whole accumulated conversation as input, so turn N pays for roughly N turns' worth of history plus its own output. Twelve turns cost far more than twice six. Trimming the loop, summarising history, or caching the stable prefix attacks the growth directly, whereas a tier swap only scales the whole curve down by a constant.

saying these in an interview costs you the question

  • Quotes one price per model instead of separate input and output rates
  • Thinks output and input tokens are billed identically
  • Assumes an Opus call always costs more than a Haiku call
  • Compares tiers on sticker price without measuring token counts
  • Forgets that an escalating cascade pays for both calls

context