skip to content

How does context window size vary across Claude's Haiku, Sonnet and Opus tiers?

level: middleimportance: should knowfreq 44%

answer

  1. Not a ladder like price is
  2. Same standard window across current tiers
  3. Long context arrived as an opt-in
  4. Output ceiling is a different limit
  5. Overflow fails, it does not truncate

basics

~20 s

Context window is not a tier ladder: the current Claude tiers share the same standard window, so moving from Haiku to Opus buys capability and not extra room. Extended long-context has been offered as a per-model opt-in with its own pricing, not as an Opus perk.

solid answer

~50 s

A common misconception is that context length scales with the tier the way price and capability do. It does not — as of mid-2026 the current Claude generation ships the same standard 200K-token window on Haiku, Sonnet and Opus alike, so "my documents don't fit, I'll upgrade to Opus" is a false inference. Where a larger window has been available, it has arrived as a **per-model opt-in beta** (an extended one-million-token window offered on Sonnet), gated behind a beta header and priced at a premium once a request exceeds the standard window. Two other limits are genuinely per-model and are what people often confuse with context: the maximum output tokens you may request in a single response, and the per-model token rate limits on your account. If the real problem is that the input does not fit, the fix is chunking, retrieval or summarisation — or the specific model that offers the extended window — not a tier upgrade.

go deeper

for a junior

Know that the context window covers the whole request plus the response, and that it is shared across the current Claude tiers rather than growing as you move up to Opus.

for a middle

Distinguish the context window from the per-model maximum output tokens and from rate limits, and know that overflow is rejected rather than silently truncated.

for a senior

Demonstrate that you solve a does-not-fit problem with retrieval, summarisation or map-reduce first, and reach for an extended-context model only when the task genuinely needs whole-corpus attention.

for a principal

Own the economics: long context bills at input rate on every turn and adds latency, so the architectural call is retrieval-plus-caching versus a premium long-context model, decided on measured cost per resolved task.

## The claim being tested The question is really a misconception probe. Price, latency and capability all move together up the Haiku → Sonnet → Opus ladder, so candidates extrapolate and assume context length does too. It does not. In the current Claude generation, all three tiers expose the same standard context window — 200,000 tokens as of mid-2026 — covering everything you send: the system prompt, tool definitions, the full message history, images, and the tokens the model generates in this turn. ## Two limits that are not the same thing **Context window** is the total budget for prompt plus response. Exceed it and the API rejects the request with a validation error rather than silently truncating — an important operational detail, because it means a runaway conversation fails loudly and you must implement your own trimming or summarisation policy. **Maximum output tokens** is a separate, genuinely per-model ceiling on how long one response may be. It is smaller than the window, it does vary between models, and it is the number that changes underneath you when you swap tiers. If a tier migration suddenly truncates your long generations, this ceiling — not the context window — is usually why. **Rate limits** are the third thing people conflate with capacity. Token-per-minute limits are allocated per model on an account, so they constrain throughput, not the size of any single request. ## Extended context Anthropic has offered an extended one-million-token context window as an opt-in beta on Sonnet, enabled with a beta header on the request and subject to premium pricing on requests that run past the standard 200K. Note the shape of that fact, because it is the opposite of the intuition being tested: the long-context option appeared on the **middle** tier, not the top one. Availability, the exact header value and the pricing tiers all change, so in an interview describe the mechanism — per-model opt-in, header-gated, premium priced beyond the standard window — and say you would confirm the current details in the model documentation rather than reciting a header string from memory. ## What to do when the input genuinely does not fit Upgrading the tier is not on this list, and saying so is the point of the answer: - **Retrieve instead of stuffing.** Chunk the corpus, embed it, and put only the relevant passages in the prompt. This is usually cheaper *and* more accurate than a giant context, because the model is not asked to locate a needle in a haystack. - **Summarise the history.** For long conversations and agent loops, periodically replace old turns with a compact summary. Without this, an agent's context grows monotonically until it hits the wall. - **Map-reduce the document.** Process sections independently on the small tier, then combine the section results in one call. - **Use the model that actually offers the bigger window**, accepting the premium pricing, when the task genuinely requires whole-corpus attention — a cross-reference over an entire codebase or contract set, where chunking would destroy the relationships you need. ## Why long context is not free even when it fits Even inside the window, a very large prompt costs real money at the input rate on every single turn, adds latency before the first token appears, and can dilute the model's attention on the instruction you actually care about. "It fits" is not the same as "it is a good idea". Teams that move from stuffing to retrieval usually see cost, latency and accuracy all improve at once, which is why the retrieval answer is the one interviewers are listening for. ## Interaction with caching When a large prefix is stable across requests — a fixed policy document, a tool catalogue, a system prompt — caching that prefix makes long context far more affordable than the raw input rate suggests. That changes the economics of a big-context design without changing the tier, and it is the standard pairing for a document-QA workload where every user question sits behind the same corpus.

  • What happens if your request exceeds the model's context window?
    The API rejects it with a validation error rather than silently dropping old messages. That makes overflow a caller responsibility: long-running conversations and agent loops need an explicit trimming or summarisation policy, plus a token estimate before sending, or they will start failing outright once history accumulates past the limit.
  • A tier migration suddenly truncates your long reports. What is the likely cause?
    The per-model maximum output tokens, not the context window. That ceiling varies between models, so a request that asked for a very long response on the old model may exceed what the new one permits, or your max_tokens value may now sit above the model's cap. Check the target model's output limit and adjust before shipping the swap.
  • When is a million-token context genuinely the right tool rather than retrieval?
    When the task requires relationships across the whole corpus that chunking would sever — cross-referencing every clause of a contract set, or reasoning over an entire codebase where the answer depends on distant call sites. For lookup-style questions retrieval usually wins on cost, latency and accuracy, so long context should be the exception you justify.

saying these in an interview costs you the question

  • Assumes Opus has a bigger context window than Haiku
  • Thinks exceeding the window truncates the oldest messages
  • Confuses maximum output tokens with the context window
  • Treats a fitting prompt as automatically a good design
  • Believes long-context pricing is the same as standard pricing

context