What do the Haiku, Sonnet, and Opus tiers mean in Claude's model line-up?
answer
- Three poetry names, one ladder
- Small, balanced, frontier
- Same endpoint, different model string
- Cost and latency climb together
- Task difficulty decides the gap
basics
~20 sHaiku, Sonnet and Opus are Anthropic's three Claude size tiers, ordered smallest to largest. Haiku is the fastest and cheapest, Opus is the most capable and most expensive, and Sonnet sits between them. All three are called through the same Messages API.
solid answer
~50 sAnthropic ships each Claude generation in three tiers named after poetry forms, which encode size: **Haiku** is the small, low-latency, low-cost tier; **Sonnet** is the balanced default; **Opus** is the frontier tier for the hardest reasoning and long agentic work. They are not different products with different APIs — you call all of them through the same Messages endpoint, and switching tiers is literally changing the `model` string on the request. What changes with the tier is capability on hard tasks, time-to-first-token and end-to-end latency, and the per-token price, which climbs steeply from Haiku to Opus. Because the interface is identical, the tier is a runtime choice you can make per task rather than a one-time architectural commitment, which is exactly how production systems use it: cheap extraction and classification on the small tier, the genuinely hard step on the large one.
go deeper
Know the ordering — Haiku smallest and cheapest, Sonnet in the middle, Opus largest and most capable — and be able to say that switching between them means changing the model name on the same API call.
Explain the three axes that move together (capability on hard tasks, latency, price per token) and name at least one thing that does not change with the tier, such as the API shape or the standard context window.
Show that you assign tiers per pipeline step rather than per application, and that you justify a downgrade with a held-out evaluation set instead of a spot check on a few prompts.
Own the framing that model tier is a runtime configuration knob, not an architectural commitment: keep the choice behind one config layer so cost, latency and quality can be retuned per workload as new tiers ship.
## The naming scheme Anthropic names Claude models after poetry forms of increasing length, and the ordering is deliberate: a haiku is short, a sonnet is medium, an opus is a large work. Every Claude generation has shipped some subset of these three tiers, so a full model name combines the tier with a generation number — the family reads as "Claude <tier> <version>". Interviewers use the tier names as shorthand for a position on the cost/capability ladder, so knowing the ordering (Haiku < Sonnet < Opus) is the floor. ## What actually differs between tiers **Capability.** The larger tiers are better on tasks with a long chain of dependent steps: multi-file code changes, ambiguous instructions, subtle reasoning, long agentic loops where one early mistake compounds. On easy, well-specified tasks — classify this ticket, pull these five fields out of this invoice, rewrite this sentence — the gap narrows sharply or disappears. This is the single most important nuance: the tiers are not "bad, okay, good", they are points on a curve whose spread depends entirely on task difficulty. **Latency.** Smaller models emit tokens faster and start responding sooner. For anything user-facing and interactive — autocomplete, inline suggestions, a chat that must feel instant — latency often decides the tier before capability does. **Price.** Pricing is per million tokens, quoted separately for input and output, and the rate climbs with the tier. The gap between the smallest and largest tier is typically several-fold, which is why a workload that runs millions of cheap calls is usually economically impossible on the top tier and trivially affordable on the small one. **Response-length ceilings.** The maximum number of output tokens you may request in one response varies by model, so swapping tiers can change the ceiling on how long a single answer can be. Check the ceiling for the specific model you are moving to rather than assuming it carries over. ## What does NOT differ The API surface is the same. Same endpoint, same request shape, same SDK client, same authentication, same streaming events, same tool-use and vision support in the current generation. That is a design decision with a practical consequence: a tier swap is a configuration change, not a rewrite. It also means you can A/B two tiers behind the same code path and compare outputs on real traffic. Context window is also not a tier ladder — the current tiers share the same standard window, so "I need to fit more text, therefore I need Opus" is a false inference. ## How the choice is actually made in production The naive pattern is to pick one model for the whole application. The professional pattern is to pick a tier **per step**, because most real pipelines are a mix of trivial and hard work. A support-triage system might use the small tier to classify the incoming message, the small tier again to extract the order id, and the large tier once to draft the reply to a genuinely complicated complaint. The cost profile of that system is dominated by the cheap calls and its quality is dominated by the one expensive call. The decision procedure that survives interview scrutiny is: 1. Start on the middle tier to get the task working at all, because it is the cheapest way to find out whether the task is even feasible. 2. Build a small held-out evaluation set of real inputs with known-good outputs. 3. Try the smaller tier against that set. If accuracy holds, downgrade — you have just cut cost and latency for free. 4. Only escalate to the top tier for steps where the eval demonstrably fails on the middle tier, and where the failure actually costs you something. Every arrow in that procedure is justified by measurement rather than intuition, which is exactly what an interviewer is listening for. ## Common traps Defaulting to the top tier "to be safe" is the most expensive mistake teams make, and it also makes the product slower, which users notice. The inverse mistake — forcing everything onto the small tier to save money — shows up as quiet quality regressions in the hard 5% of traffic that nobody eyeballs. Both mistakes come from treating the tier as a global constant instead of a per-task decision backed by an eval set. A second trap is assuming tier names carry across vendors. Haiku, Sonnet and Opus are Anthropic's naming only; other providers use entirely different schemes, and a "small model" from one vendor is not price- or capability-comparable to another's by name alone.
- If the API is identical across tiers, what has to change in your code to switch?Only the model identifier on the request. The client, authentication, message structure, streaming handling and tool definitions all stay the same. The two things worth re-checking are the maximum output tokens the target model allows, and your prompt — a prompt tuned for a larger model sometimes needs to be more explicit and more constrained to work well on a smaller one.
- When would you deliberately choose the smallest tier even though the largest is affordable for your volume?When latency is the product requirement. Interactive surfaces — inline completion, live suggestions, a chat that must respond within a few hundred milliseconds — are judged on responsiveness, and the small tier starts and finishes faster. Cost is not the only reason to go small; perceived speed frequently matters more to users than a marginal quality gain they cannot see.
- Does every Claude generation ship all three tiers at once?No. Tiers are released on their own schedule, so at any moment the newest Opus and the newest Haiku may belong to different generation numbers. That means comparing "Opus versus Haiku" is really comparing two specific model snapshots, and you should evaluate the exact model identifiers you plan to call rather than reasoning from the tier name alone.
saying these in an interview costs you the question
- Thinks Opus is always the correct default
- Assumes each tier has its own endpoint or SDK
- Believes the tiers differ only in speed
- Says a higher tier automatically gives a longer context window
- Treats the tier as one global choice for the whole app