skip to content

How would you train a difficulty classifier to route requests before the LLM call?

level: seniorimportance: should knowfreq 38%

answer

  1. Decide before spending anything
  2. Labels come from running both tiers
  3. The router must be nearly free
  4. Errors are asymmetric, so is the threshold
  5. Keep an exploration slice alive

basics

~20 s

Mine labels from logged traffic where both tiers ran and a verifier judged the cheap answer, train a small cheap model over request features, and tune its threshold on the cost-quality curve rather than on accuracy. Keep a random exploration slice so labels stay fresh.

solid answer

~50 s

A predictive router decides before spending anything, so it avoids the double latency a cascade pays on hard requests. The hard part is labels: run both tiers on a sample of real traffic in shadow, and label a request "needs the large model" when the small tier's output fails a verifier or materially diverges from the large tier's answer under a rubric. Features are cheap and pre-call — an embedding of the request plus structured metadata like document type, length, entity counts, customer segment. A logistic regression, gradient-boosted tree or small fine-tuned encoder runs in single-digit milliseconds for effectively nothing, which matters because a router that costs a model call has defeated its own purpose. Tune the threshold against asymmetric costs: a missed escalation ships a wrong answer, an unnecessary one only costs tokens. And keep 1-5% of traffic on both tiers permanently, or the label supply dies and drift goes unseen.

go deeper

for a junior

Know the idea: a cheap classifier looks at the request first and picks which model should answer, instead of trying the small model and retrying.

for a middle

Explain where the training labels come from — running both tiers on sampled traffic and judging the cheap answer — and why the router itself has to be far cheaper than the call it replaces.

for a senior

Show that you tune the threshold against asymmetric error costs, measure precision and recall per request class, and keep a random exploration slice so the router's own decisions don't censor its future training data.

for a principal

Own the lifecycle: who re-labels after a model upgrade, how seasonality and traffic drift are detected, and whether the maintenance burden of a learned router is justified over a rule table plus a verifier.

## Predictive routing versus reactive cascading A cascade discovers difficulty by trying: cheap model, check, escalate. A predictive router decides up front from the request alone. The trade is clean. The cascade needs no training data and adapts automatically when the small model improves, but it pays small-tier latency and tokens on every hard request. The router pays nothing on hard requests and keeps the tail clean, but it needs labels, it can be wrong in ways nothing downstream catches, and it goes stale. Most mature systems run both: a classifier that sends the obviously-hard classes straight to the strong tier, and a verifier-driven cascade behind the cheap tier for everything else. Consider a tax-prep assistant with a classifier trained on eight thousand past filings that predicts escalation before the small model is called: multi-state K-1 filings skip the cheap tier entirely, standard W-2 questions go to it with a verifier behind them. ## Where the labels come from This is the question interviewers are actually probing. The only trustworthy label source is your own traffic run through both tiers: 1. Sample production requests — stratified across request classes, not uniformly, so rare hard classes are represented. 2. Run both the small and large tier on each, offline or in shadow. 3. Label positive ("needed the large model") when the small tier's output fails a deterministic verifier, or when a rubric-based comparison judges the large tier's answer materially better. Label negative when the small answer is acceptable. 4. Have humans adjudicate a slice of the automated labels to check the labeling function itself is sound. What does *not* work as a label source: the small model's self-reported confidence (it is generated text, not a measurement), public difficulty ratings for superficially similar tasks (they measure a different distribution), and raw proxies like request length used alone. Length is a useful *feature*; it is not a label. ## Features and model choice The router must be nearly free, or its cost eats the saving it exists to produce. That rules out using an LLM as the router for high-volume traffic. What works: an embedding of the request text, concatenated with structured metadata you already hold — document or form type, token count, count of distinct sub-questions, entity or account counts, customer segment, source channel, whether a previous attempt failed, retrieval score if the request is retrieval-backed. On top of that, a logistic regression, a gradient-boosted tree, or a small fine-tuned encoder. All run in milliseconds at negligible cost, and the linear and tree models have the extra benefit of being inspectable, which matters when someone asks why a filing was routed to the expensive path. ## Choosing the threshold Do not optimize accuracy. The two errors have different prices: - **Under-routing** (predicted easy, actually hard): the small tier ships a wrong answer. Cost = the business cost of that error, which in a regulated domain can be very large and is not on any cost dashboard. - **Over-routing** (predicted hard, actually easy): you paid the large tier unnecessarily. Cost = the price difference, bounded and visible. So sweep the threshold, and for each value plot blended cost per request against end-task quality. That sweep *is* the frontier for this router; pick the point that clears your quality floor most cheaply. Report escalation precision and recall separately, and report them per request class — a router that is excellent in aggregate is often blind on the one class that matters, because that class is a small fraction of the training data. ## Drift and the feedback loop Two things move under a deployed router. **Input drift.** The traffic mix shifts — seasonality is the obvious case, where a filing-deadline surge changes both volume and the composition of requests. A router trained on off-season traffic can be badly miscalibrated in peak season. **Label drift from model upgrades.** Every label encodes "the small model of that date could not handle this". Upgrade the small model and a chunk of your positive labels become wrong; the router keeps escalating requests the cheap tier now handles fine, and you silently pay for a saving you already earned. Any model change on either tier invalidates the training set and demands re-labeling. **The feedback loop is the subtle failure.** Once the router sends a class straight to the large tier, you never again observe whether the small tier could have handled it — the router's own decisions censor the data that would correct it. The fix is deliberate exploration: keep a random slice of traffic, typically a few percent, running through both tiers regardless of the prediction. That slice is your continuous label supply, your drift detector, and your evidence when someone proposes retiring the router. ## What to monitor Routed-fraction per class, blended cost against an all-large baseline, end-task quality by class, router precision and recall recomputed weekly on the exploration slice, and the age of the training set. When the last one grows without the others being re-checked, the router has quietly become a hard-coded rule table.

  • Why can't the router itself be an LLM call?
    On high-volume traffic it usually can't, because the router runs on 100% of requests. An LLM classifier can easily cost a meaningful fraction of the tier it is protecting, wiping out the saving, and it adds latency and nondeterminism ahead of every call. It becomes defensible only on low-volume, high-value traffic where a wrong route is expensive and per-request cost barely matters.
  • How does a model upgrade on the cheap tier affect an existing router?
    It invalidates the labels. Every positive example encoded "the old small model failed this", so after the upgrade the router keeps sending to the expensive tier work the new cheap model now handles — you pay for a saving you already earned. Re-label from the exploration slice after any upgrade on either tier, re-sweep the threshold, and treat the router's training-set age as a monitored metric.
  • How would you combine a predictive router with a verifier-based cascade?
    Use the router as a fast pre-filter for classes it is confident about: obviously-hard classes skip the cheap tier and avoid the double-latency penalty, obviously-easy ones go straight to it. Everything in the uncertain middle takes the cascade path — cheap model plus deterministic verifier — so the verifier catches what the router could not predict. The cascade also supplies labels that keep the router current.

saying these in an interview costs you the question

  • Trains the router on the small model's self-reported confidence
  • Optimizes router accuracy instead of the cost-quality trade-off
  • Uses an LLM call as the router on high-volume traffic
  • Never re-labels after upgrading either model tier
  • Routes all traffic by prediction, leaving no exploration slice

context