When does a small-first LLM model cascade stop saving money?
answer
- Escalated requests pay both tiers
- Break-even is a rate, not a ratio
- The check has a price too
- Tail latency moves the wrong way
- Cheap wrong answers are invisible in cost
basics
~20 sA cascade only pays when the cheap tier answers enough traffic outright. Escalated requests are billed for both tiers plus the escalation check, so savings disappear once the escalation rate approaches the point where the small tier's spend no longer offsets the large tier's.
solid answer
~50 sA cascade sends every request to a small model first, checks the output, and re-asks a larger model only when the check fails. Per request you pay `small + check + escalation_rate x large`, against a baseline of `large` for always using the big model. Solving that, the break-even escalation rate is roughly `1 - small/large` (minus whatever the check costs). If the small model is 5% of the price, you can escalate ~95% of traffic before you lose money on tokens — so cascades rarely fail on token arithmetic alone. They fail three other ways: the escalation check is itself an LLM call and can cost more than the tier it guards; escalated requests pay both latencies serially, so p99 gets worse exactly on the hard requests; and requests the check wrongly passes ship a cheap wrong answer, which is a quality cost the spreadsheet does not show.
code
python · 13 linesdef cascade_cost(small, verify, large, escalation_rate):
return small + verify + escalation_rate * large
def break_even_rate(small, verify, large):
return max(0.0, (large - small - verify) / large)
SMALL, VERIFY, LARGE = 0.05, 0.01, 1.00
print(cascade_cost(SMALL, VERIFY, LARGE, 0.30)) # 0.36 vs 1.00 baseline
print(break_even_rate(SMALL, VERIFY, LARGE)) # 0.94
# A judge-model check that costs 40% of the large model wrecks it:
print(break_even_rate(SMALL, 0.40, LARGE)) # 0.55go deeper
Know the shape: try a cheap model first, check its answer, re-ask a stronger model when the check fails. Be able to say that escalated requests are paid for twice.
Be ready to write the cost formula on the whiteboard and derive the break-even escalation rate from the price ratio. Explain why the verification step's cost belongs in that formula.
Show the production judgment: cascades usually win on tokens and lose on tail latency and silent false passes. Say how you would measure blended cost, end-task quality by segment, and split latency.
Own the framing that a cascade is one point on a cost-quality-latency surface, not a default. Argue when the pattern is the wrong shape entirely and what organizational cost two model tiers add to prompts, evals and upgrades.
## What a cascade is A cascade is a serial routing pattern: the request goes to the cheapest capable model first, its output is inspected by some check, and only if the check fails is the request re-sent to a stronger, more expensive model. It is *reactive* — the system learns the request was hard by trying it. That distinguishes it from a predictive router, which classifies difficulty up front and dispatches once. A concrete shape: a tax-prep assistant answers standard W-2 questions with a small model, validates that the answer cites a real form field and produces a numeric result that reconciles with the filing data, and escalates multi-state K-1 returns — where the check fails or the input is flagged as out of scope — to a reasoning model. ## The cost arithmetic Let `c_s` be the small-model cost per request, `c_l` the large-model cost, `c_v` the cost of the verification step, and `e` the fraction of traffic that escalates. Cascade cost per request is: `c_s + c_v + e * c_l` The baseline you are comparing against is `c_l` (always use the large model). The cascade wins while `c_s + c_v + e * c_l < c_l`, i.e. while `e < 1 - (c_s + c_v) / c_l` If the small model costs 5% of the large one and the check is free, break-even sits around a 95% escalation rate. That is a startling number the first time you compute it, and it is the single most useful thing to be able to say about cascade economics in an interview: **on tokens alone, cascades are extremely hard to lose.** Even a mediocre small tier that only handles a third of traffic is saving real money. ## Where cascades actually lose **The check is not free.** If your escalation trigger is "ask a judge model whether the small answer is good", you have added a third call to every request. A judge on a mid-tier model can easily cost more than the small tier it is guarding, and it pushes `(c_s + c_v)/c_l` toward 1, collapsing the break-even rate. Cheap deterministic checks — schema validation, a recomputed total, a database lookup that confirms an entity exists, a compile or test run — keep `c_v` near zero and are what make cascades viable. **Latency is charged twice on the tail.** Escalated requests pay small-tier latency, plus check latency, plus large-tier latency, in sequence. The requests that escalate are exactly the hard ones that were already slow, so a cascade tends to make p99 worse even while it makes mean cost better. If your SLO is a tail SLO, a cascade can be a regression you deployed as an optimization. Mitigations: run the tiers speculatively in parallel for classes known to be hard and cancel the loser (costs tokens, buys tail latency), or route those classes predictively so they skip the small tier entirely. **False passes are a silent quality cost.** The check's job is to catch bad small-tier answers. Every one it misses ships a cheaper, worse answer than the baseline would have produced. This never appears in the cost dashboard, so it must appear in the eval: measure end-task quality of the cascade as a whole, not just the escalation rate. In a regulated domain a single wrong answer can cost more than a month of the savings. **Operational surface doubles.** Two models means two prompts, two output formats to keep aligned, two sets of eval results, and two upgrade paths. When either provider ships a new version, the routing threshold and the quality comparison both have to be re-tuned. ## What to measure Four numbers describe a cascade: escalation rate, blended cost per request against the all-large baseline, end-task quality of the cascade versus all-large, and p50/p99 latency split by escalated versus not. Track quality by segment too — a cascade that is neutral in aggregate can be badly degraded on one important request class while the easy majority masks it. ## When the pattern is wrong If almost every request is hard, skip the small tier — you are paying an extra call for nothing and hurting the tail. If almost every request is easy, do not build a cascade either; ship the small model with a narrow guardrail. Cascades pay in the middle, on mixed traffic with a cheap, trustworthy check, and where the cost of an occasional cheap wrong answer is bounded.
- If the small tier handles 70% of traffic, why can p99 latency still get worse?Because the 30% that escalate pay small-tier latency plus check latency plus large-tier latency, serially — and those are the same hard requests that already sat in the tail. The cascade adds fixed overhead to precisely the slowest bucket. If a tail SLO matters, route those classes predictively so they skip the small tier, or run both tiers in parallel and cancel the large call when the check passes.
- When would you run both tiers in parallel instead of sequentially?When latency matters more than tokens: fire the small and large calls together, return the small answer if the check passes, otherwise the large one. You pay large-model tokens on every request, so the cost saving mostly evaporates, but the escalated path no longer serializes two calls. It is a reasonable trade on an interactive path with a strict tail SLO and modest traffic.
- How would you decide whether the escalation check should be a model call or a deterministic rule?Price it against the tier it guards. A judge call that costs a meaningful fraction of the large model destroys the break-even rate, so start with deterministic checks — schema validation, recomputed arithmetic, an entity lookup, a test run. Reach for a model-based check only when no programmatic signal exists, and then use the smallest model that reaches acceptable agreement with human labels on your own data.
saying these in an interview costs you the question
- Assumes a cascade always cuts cost because the small model is cheaper
- Forgets that escalated requests are billed for both tiers
- Treats the escalation check as free when it is a model call
- Reports cost savings without measuring end-task quality
- Claims a cascade cannot lower quality below the large model