Why are logprobs and self-reported confidence weak escalation triggers in a cascade?
answer
- Fluency is not correctness
- The number is generated, not measured
- Averaging buries the decisive tokens
- Post-training flattens calibration
- Prefer evidence outside the generation
basics
~20 sToken log-probabilities measure how fluent the wording is, not whether the content is right, and instruction-tuned models are poorly calibrated. Self-reported confidence is just more generated text — systematically high and easily swayed by prompt phrasing. Neither tracks correctness well enough to gate escalation.
solid answer
~50 sThe escalation trigger is the whole design of a cascade, and the two easiest signals are the two weakest. Token log-probabilities score the model's certainty about the *next token*, which is dominated by syntax and common phrasing; a confidently-worded fabrication scores high, and averaging over a long answer buries the few tokens that carried the claim. Post-training also flattens calibration, and many reasoning models do not expose usable logprobs at all. Self-reported confidence is worse: the number is sampled like any other token, clusters at 0.8-0.95, and moves when you reword the prompt rather than when the answer changes. What works is evidence external to the generation — schema validation, a recomputed total, a lookup that confirms the cited entity exists, a test or compile run — plus structural request features and, where nothing programmatic exists, a small dedicated verifier validated against human labels.
go deeper
Know that a model sounding confident and a model being right are different things, and that a self-reported confidence number is generated text rather than a measurement.
Explain what a token log-probability actually scores, why averaging it over an answer dilutes the informative tokens, and why preference tuning damages calibration.
Demonstrate the production instinct: replace self-assessment with external evidence — schema checks, recomputed values, grounding against retrieved passages — and tune the escalation threshold on labeled traffic with precision and recall reported separately.
Own the argument that the escalation gate is the system's real quality control, needs its own eval, ownership and re-validation cadence, and that any routing scheme resting on model self-report is an unmonitored risk in a regulated domain.
## The trigger is the design In a small-first cascade every interesting property — cost, quality, tail latency — is determined by one function: given the small model's output, do we accept it or escalate? Get that function right and a cheap tier is nearly free quality. Get it wrong in the permissive direction and you ship cheap wrong answers; wrong in the conservative direction and you escalate everything, paying for two tiers to get one model's quality. ## Why token log-probabilities disappoint A log-probability is the model's assigned likelihood for the token it emitted, given everything before it. That quantity is overwhelmingly driven by *linguistic* predictability. In "the deduction is limited to", the next tokens are near-certain regardless of whether the number that follows is correct. Aggregate the sequence and the mass of easy structural tokens drowns the handful of tokens that carried the actual claim; the answer that invents a plausible figure in fluent prose scores as confidently as the correct one. Three further problems. First, calibration: a base model's probabilities are reasonably calibrated on next-token prediction, but preference tuning pushes models toward assertive, well-formed output and flattens the relationship between probability and correctness. Second, availability: extended-thinking models often do not expose token probabilities over the hidden reasoning, so the signal you can see covers only the visible answer. Third, comparability: thresholds tuned on one model version do not transfer to the next, so the signal quietly decalibrates on every upgrade. Log-probabilities are not useless. They are a reasonable *cheap prior* over a narrow, fixed-format task — classification into a small label set, where the probability sits on a single decisive token — and they are almost worthless as a general "is this answer right" score over free prose. ## Why self-reported confidence is worse Asking the model to append "confidence: 0.87" produces text, not introspection. There is no internal uncertainty estimate being read out; the model is predicting what a confident-sounding assistant writes. In practice the numbers pile up in a narrow high band, barely separate right from wrong answers, and shift with surface changes to the prompt — telling the model to be careful moves the number without moving the accuracy. Coarse self-assessment ("I could not find this in the provided documents") is more useful than a fake decimal, because it is a claim about the *input* the model can actually observe rather than a claim about its own correctness. ## Signals that do work The reliable triggers put evidence outside the generation: - **Deterministic verifiers.** Does the output parse against the schema? Does the recomputed total match the line items? Does the cited form, ticket or account id exist in the system of record? Do the generated tests compile and pass? These are cheap, unambiguous, and they fail loudly. - **Grounding checks.** For a retrieval-backed answer, check that every claim maps to a retrieved passage. Failure to ground is a strong escalation trigger and a strong hallucination alarm at the same time. - **Request-side structural features.** Input length, presence of rare entities, number of distinct sub-questions, low retrieval scores, an unusual document type. These are available *before* the small model runs, so they let you skip the cheap tier entirely on requests it will obviously fail. - **Disagreement across independent attempts.** Sampling the small model more than once and escalating on divergence is a genuine signal, though it multiplies the cheap tier's cost and belongs to the broader family of self-consistency techniques rather than to routing specifically. - **A small dedicated verifier.** When nothing programmatic exists, a purpose-built checker prompt on a cheap model, validated against human labels on your own traffic, beats asking the generator to grade itself — it is a fresh pass over the answer without the generator's commitment to its own output. ## Tuning and measuring the trigger Whatever the signal, treat the threshold as a tuned parameter, not a constant. Build a labeled set from real traffic — small-tier output plus a human or strong-model judgment of whether it was acceptable — and sweep the threshold to get escalation precision (of the requests we escalated, how many genuinely needed it) and escalation recall (of the requests that needed the large model, how many did we catch). Recall is usually the metric that matters, because a missed escalation ships a wrong answer while a spurious escalation only costs money. Re-validate after every model change on either tier. A small-model upgrade shifts which requests it can handle; a large-model upgrade shifts what escalation is worth. And log the signal value alongside the outcome in production so the labeled set keeps refreshing instead of aging into the traffic mix of six months ago.
- Which is more costly for a cascade: low escalation precision or low escalation recall?Usually low recall. A missed escalation ships a wrong answer to a user, which is a quality and sometimes a compliance failure that no dashboard shows. Low precision only means you escalated requests the small tier could have handled — that costs tokens and some latency, both bounded and visible. So tune the threshold toward recall, and revisit only if the escalation rate climbs far enough to erase the savings.
- How would you validate a small verifier model before trusting it as the escalation gate?Build a few hundred labeled examples of small-tier output judged acceptable or not by humans, run the verifier at temperature 0, and measure agreement against those labels — not against the generator. Check both error directions separately, and check for the classic biases: preferring longer answers, and rating output more favourably when it came from the same model family. Re-run the check on every model or prompt change.
- Are token log-probabilities ever a defensible routing signal?Yes, in narrow, fixed-output tasks where the decision rests on one or two tokens — routing a support ticket into six categories, or a yes/no extraction. There the probability sits on the decisive token instead of being diluted across prose, and a margin between the top two labels is a usable uncertainty proxy. Calibrate the threshold on your own data and re-calibrate after any model upgrade.
saying these in an interview costs you the question
- Treats a high log-probability as evidence the answer is factually correct
- Trusts a model's self-reported confidence score as a measured probability
- Averages logprobs over a long answer and calls it certainty
- Sets an escalation threshold once and never re-validates after model upgrades
- Assumes a model can reliably detect its own hallucinations