On the quality-cost frontier, how do you choose an operating point for an LLM feature?
answer
- Keep only non-dominated configurations
- The bar comes from the cost of being wrong
- Ratios pick cheap mediocrity
- One point per class, not per system
- Latency turns the curve into a surface
basics
~20 sSweep routing thresholds and effort levels, plot end-task quality against cost per request, keep only the non-dominated configurations, then let the business set a quality floor or a budget ceiling and take the cheapest or best point that satisfies it — per request class, not globally.
solid answer
~50 sBuild the frontier empirically: run your eval suite across every routing configuration you can actually ship — escalation thresholds, effort levels, model pairings — and plot quality against blended cost per request. Discard anything another configuration beats on both axes; what remains is the frontier. Choosing on it is a business decision, not a maths one. Where a wrong answer is expensive — regulated advice, irreversible actions — set a quality floor from that cost and take the cheapest point clearing it with margin. Where output is a draft a human reviews, set a budget ceiling and take the best quality inside it. Maximizing a quality-per-dollar ratio is the common mistake: it picks a cheap mediocre point when a floor exists. Two refinements matter: the frontier differs per request class, so per-class points dominate any single global one, and latency is a third axis that can make an otherwise attractive point unshippable.
go deeper
Know that cheaper models are not simply worse — each configuration trades quality against cost, and the job is to pick a point rather than to always take the strongest model.
Be able to explain domination: a configuration beaten on both quality and cost is discarded, and the survivors form the curve you actually choose from.
Show how you build the frontier from your own eval suite, carry p99 alongside every point, and derive the quality floor from the cost of a wrong answer rather than picking by ratio.
Own the segmentation and the cadence: per-class frontiers because volume and risk sit in different classes, marginal cost per point of quality stated in terms finance can act on, and a named owner who re-sweeps after every model or pricing change.
## What the frontier is Every routing decision — which model, which effort level, what escalation threshold, whether a semantic cache tier serves the request at all — produces a (quality, cost) pair for your feature. Plot them and most points are *dominated*: some other configuration is both better and cheaper. Throw those away. What remains is the Pareto frontier, the set of configurations where buying more quality strictly costs more money. Choosing an operating point means picking one of those, and the whole value of drawing the frontier is that it turns an argument about opinions into a choice between measured options. ## Building it The frontier has to be measured on *your* task, with your eval suite, on your traffic distribution. Sweep the parameters you can actually change in production — escalation threshold, effort level per class, model pairing, cache similarity threshold — and for each configuration record end-task quality (the score your eval suite produces, not a proxy), blended cost per request, and latency at p50 and p99. Run enough repetitions per configuration that the quality difference between neighbouring points is bigger than the run-to-run noise; otherwise you will "choose" a point that is indistinguishable from its neighbour. Two things make a frontier misleading. First, evaluating on a dataset that does not match production traffic mix — the cheap configurations look far better than they are because the hard segment is underrepresented. Second, measuring quality as an average when the distribution is what matters: a configuration with the same mean quality but a fat tail of catastrophic answers is not equivalent. ## Choosing a point The frontier does not tell you where to sit; the cost of being wrong does. **Quality-floor mode.** Where a bad answer is expensive — advice with regulatory exposure, an irreversible action, anything a customer acts on without review — derive a floor from that cost and take the cheapest frontier point that clears it, with margin for the fact that production quality is usually a little worse than eval quality. Below the floor, cost savings are not savings; they are deferred liability. **Budget-ceiling mode.** Where output is a draft a human edits, or an internal summary, fix an acceptable cost per request from unit economics and take the best quality inside it. **What not to do:** maximize quality per dollar. That ratio is maximized by cheap mediocrity whenever a floor exists, and it is insensitive to exactly the thing that matters — whether the point clears the bar. Ratios are for comparing marginal moves, not for selecting an operating point. ## Marginal accuracy per dollar The useful ratio is the *marginal* one: what does the next increment of spend buy? Frontiers are strongly concave. Moving from a weak configuration to a competent one often costs little and buys a lot; the last few points of quality cost superlinearly, because they come from routing the hardest residual traffic to the most expensive tier at the highest effort. Compute the slope between adjacent frontier points and you can say, concretely, "raising effort from low to high for amended filings costs 2.4x on that class and buys 1.9 points of accuracy" — a sentence a finance partner can act on, unlike "the expensive model is better". ## The frontier is per segment A single global operating point is almost always dominated by a set of per-class points. Split traffic into classes with genuinely different difficulty and different error costs, draw a frontier per class, and choose per class. The routine, high-volume class usually sits far down the cheap end with no measurable quality loss; the rare, high-stakes class sits at the expensive end, and because it is rare its cost barely registers. That asymmetry — most volume in the cheap class, most risk in the rare one — is why segmented routing beats global model selection, and it is the core argument for building a router at all. ## Latency is a third axis Quality and cost make a curve; add latency and it is a surface. A configuration that is attractive on the (quality, cost) plane can be unshippable because escalation doubles the tail or because high effort adds tens of seconds of thinking before anything is visible. Carry p99 alongside every frontier point, and for interactive paths treat the latency budget as a hard constraint that filters the frontier before you choose on it. ## The point is a parameter, not a decision The frontier moves under you. New model versions shift both axes, prices change, your traffic mix drifts, and a prompt improvement can lift the cheap tier enough to redraw everything. Treat the operating point as something with an owner and a re-tuning cadence, re-swept after every model or pricing change, with the sweep results stored so the choice is auditable. Frontier configurations chosen offline still have to be validated on live traffic before they become the default — that rollout machinery is a separate concern from the frontier analysis itself, but skipping it turns a measured choice back into a guess.
- Why is maximizing quality per dollar a bad selection rule?Because it is indifferent to whether the point clears the bar. A configuration at 70% accuracy for a tenth of the price wins on ratio and fails a compliance floor of 90%. Ratios are the right tool for comparing marginal moves between adjacent frontier points — what the next dollar buys — but selection should be driven by a floor derived from the cost of a wrong answer, or by a budget ceiling where errors are cheap.
- How do you keep the frontier from being an artefact of your eval set?Sample the eval set from real production traffic and stratify it so the hard segments are represented in proportion to their risk, not their volume. Run enough repetitions that differences between neighbouring configurations exceed run-to-run noise. Report quality distributions rather than only means, since a fat tail of catastrophic answers matters more than a decimal of average score, and re-draw the set as traffic drifts.
- What changes when a new model version ships at a lower price?Both axes move, so the whole frontier is redrawn and the current operating point may become dominated — the same quality is now available for less, or more quality is available at the same spend. Re-sweep thresholds and effort levels against the eval suite rather than assuming the old configuration transfers, since quality at each effort level and token consumption at each level both shift between versions.
saying these in an interview costs you the question
- Picks the configuration with the best quality-per-dollar ratio
- Chooses one global operating point for all request classes
- Compares configurations on a benchmark instead of the product's own eval
- Ignores p99 latency when selecting a frontier point
- Treats the chosen point as permanent across model and price changes