How would you route a production workload across Claude tiers instead of one model?
answer
- Per step, not per application
- Mechanical work goes small
- Escalate on a cheap check
- Both calls billed when escalating
- Rate limits are per model
basics
~20 sSplit the pipeline into steps and assign a tier per step: small models for classification, extraction and routing, the large model for the one genuinely hard step. Prove each assignment with a held-out eval set, and cascade with escalation only when escalation stays rare.
solid answer
~50 sStart by decomposing the workload, because most pipelines are a mix of trivial and hard work and a single global model over-pays on the trivial part or under-performs on the hard part. Assign Haiku to the mechanical steps — classify the intent, extract fields, decide the route, check a guardrail — and reserve Sonnet or Opus for the step where a mistake is expensive. Where difficulty is not known in advance, run a **cascade**: attempt on the cheap tier, validate the result with a cheap deterministic check (schema parse, confidence threshold, self-verification), and escalate only on failure. Cascades only pay when escalation is rare, since an escalated request pays for both calls. Two operational details make this real: rate limits are allocated per model, so small-tier traffic does not eat the large tier's quota, and every routing decision must be justified by a held-out eval rather than intuition, with the model identifier held in config so the routing table is tunable without a deploy.
code
python · 24 linesimport os
from anthropic import Anthropic
client = Anthropic() # reads ANTHROPIC_API_KEY
# Pinned snapshot ids resolved from config, never hardcoded at call sites.
SMALL = os.environ["CLAUDE_SMALL_MODEL"]
LARGE = os.environ["CLAUDE_LARGE_MODEL"]
def ask(model: str, prompt: str) -> str:
resp = client.messages.create(
model=model,
max_tokens=512,
messages=[{"role": "user", "content": prompt}],
)
return resp.content[0].text
def answer(prompt: str, is_acceptable) -> str:
draft = ask(SMALL, prompt)
if is_acceptable(draft): # cheap deterministic check
return draft
return ask(LARGE, prompt) # escalated requests pay for both callsgo deeper
Know that the model is chosen per call, so one application can use several tiers, and that the cheap tier is normally enough for classification and extraction.
Explain static per-step assignment versus a cascade, and be able to state why an escalating cascade pays for both calls and therefore needs a low escalation rate.
Show the whole operational picture: a cheap deterministic escalation check, per-model rate-limit headroom, backoff on 429 and 529, and a held-out eval that justifies every downgrade.
Own routing as a governed system — roles mapped to pinned models in one config, evals in CI, and cost per resolved task plus tail latency as the metrics the routing table is optimised against.
## Why one model for the whole app is the wrong default A typical LLM feature is not one task. A support assistant classifies the message, extracts an order id, decides whether to search a knowledge base, drafts a reply, and checks the draft against policy. Four of those five steps are mechanical; one is hard. Running all five on the frontier tier means paying the frontier rate to answer "is this a billing question or a shipping question?", and it also makes the whole feature slower, because every step inherits the slowest model's latency. The senior move is to treat the model tier as a **per-step** parameter and to build the system so that parameter is configuration, not code. ## Static routing: assign tiers by step Where difficulty is known in advance, assign statically: - **Small tier**: classification, entity extraction, routing decisions, short summarisation, format conversion, guardrail checks, query rewriting for retrieval. High volume, well-specified, verifiable output. - **Middle tier**: the default for anything with real reasoning that is not the crown jewel — drafting, analysis over retrieved context, most tool-using loops. - **Top tier**: the step where a wrong answer costs money or trust, or the long-horizon agentic work where an early error compounds over many turns. This alone typically moves the majority of calls onto the cheap tier while leaving quality where it matters, because in most systems the mechanical steps outnumber the hard one by an order of magnitude. ## Dynamic routing: cascade with escalation When difficulty varies per input, try small first and escalate. The design hinges entirely on the **escalation check**, which must be cheap and reliable: - **Structural validation** — does the output parse as the required schema, contain the required fields, cite a real document id? Free, deterministic, catches a lot. - **Business-rule validation** — does the extracted total match the invoice line items? Free and highly reliable where it applies. - **Confidence signals** — an explicit self-rating or an abstain token the small model is prompted to emit when unsure. Cheap but softer; calibrate it against labelled data. - **A second cheap model as judge.** Adds a call, so only worth it when the large tier is much more expensive than two small calls. Avoid a check that itself needs the large model — that defeats the purpose entirely. ## The cascade arithmetic Let `e` be the escalation rate, `c_s` the small-tier cost and `c_l` the large-tier cost. The cascade costs `c_s + e x c_l` versus `c_l` for calling large directly. It only wins while `c_s + e x c_l < c_l`, i.e. `e < 1 - c_s/c_l`. If the cheap tier costs a fifth of the expensive one, escalation must stay below roughly 80% — usually easy — but if the tiers are close in price, the margin evaporates fast. Escalation also costs **latency**: escalated requests pay both round trips, so the p99 gets worse even as the mean improves. For interactive surfaces, measure the tail, not the average. ## Operational plumbing **Rate limits are per model.** Each model has its own token-per-minute allocation on the account, so shifting bulk work to the small tier frees headroom on the large tier rather than competing with it. This is a real capacity argument for routing, independent of cost. **Fallbacks on overload.** Transient 429 rate-limit and 529 overloaded responses should be retried with exponential backoff and jitter. Whether to *fall back to a different tier* on overload is a product decision: silently degrading quality can be worse than a slightly slower correct answer, so make the fallback explicit, log it, and alarm on its rate. **One routing module.** Every call site should ask for a role ("fast", "balanced", "deep") and a central config should resolve that to a pinned model identifier. This keeps migrations to one change and lets you shift a role to a new tier without redeploying every consumer. ## Proving the routing table None of this is credible without evaluation. Build a held-out set of real inputs with known-good outputs per step. To downgrade a step, run both tiers over the set and compare task accuracy, not vibes. Keep the set in CI so a later prompt change cannot silently invalidate a downgrade decision made months earlier. Track, in production: cost per resolved task, escalation rate, per-step latency, and the quality metric you can actually measure (schema-valid rate, human-override rate, thumbs-down rate). Those four numbers are what turn routing from an opinion into an operations discipline. ## Failure modes to name Over-escalation that quietly makes the cascade more expensive than the simple design; a validation check so lenient that bad small-tier answers pass through; routing rules scattered across call sites so nobody can say which model served a request; and no eval, which means the first quality regression is discovered by a customer.
- How do you decide whether the escalation check should be a rule or another model call?Prefer a rule whenever the output is verifiable — schema parse, required fields, arithmetic that must reconcile, a citation that must resolve. Rules are free, deterministic and auditable. Use a model-based judge only for genuinely subjective quality, and price it in: two small calls plus an occasional large one must still beat calling large once.
- What do you monitor to know a cascade is still paying off?Escalation rate above all, since the break-even is a direct function of it, plus cost per resolved task, p99 latency including escalated round trips, and the rate at which escalated answers actually differ from the cheap one. A rising escalation rate after a prompt or traffic change is the signal to re-evaluate the split or retune the small-tier prompt.
- Is falling back to a smaller tier on a 529 overloaded response a good idea?Only when a degraded answer genuinely beats a delayed one, and only when it is explicit. Retry with exponential backoff and jitter first. If you do downgrade, log it, expose it in the response metadata for downstream consumers, and alarm on the rate — an unnoticed silent downgrade is a quality regression that no eval will catch because it only fires under load.
- How do you keep tier assignments from rotting as models change?Hold model identifiers in one config layer keyed by role rather than at call sites, and keep the per-step eval sets in CI. When a new model ships, you change one mapping and re-run the evals to see which steps can move. Without that, tier assignments become archaeology and nobody dares touch them.
saying these in an interview costs you the question
- Picks one model for the entire application
- Escalates on every request, so both calls are always billed
- Uses the expensive model to check the cheap model's output
- Downgrades a step on a spot check instead of an eval set
- Assumes all models share one account-wide rate-limit bucket