skip to content

How do you set the confidence threshold that sends an extracted field to human review?

level: seniorimportance: should knowfreq 47%

answer

  1. it is an operating point, not a constant
  2. two numbers trade off directly
  3. cost of error differs per field
  4. validations beat model scores
  5. audit what you never reviewed

basics

~20 s

Tune it against a labelled sample rather than intuition: sweep the cutoff, plot residual error of auto-posted fields against review volume, and set it per field at the point where the remaining error meets that field's business tolerance.

solid answer

~50 s

Treat it as an operating point on a curve, chosen per field. Build a labelled sample from real traffic, stratified by document class and scan quality. For each candidate cutoff, compute two numbers: the share of fields sent to review, and the error rate among the fields that were auto-posted. Those two trade off directly, and the business — not the model — decides where to sit: a handwritten weight on a customs entry and a tick-box for "hazardous" have wildly different costs of being wrong, so they get different thresholds. Confidence itself should be a composite, not one model score: classical OCR word confidence where you have it, cross-field validations like a container check digit or line items summing to the stated total, and disagreement between two extractors. Finally, audit a random sample of *auto-posted* documents continuously — that is the only way to see the residual error rate drift when input quality or the model changes.

go deeper

for a junior

Know that low-confidence extractions should go to a person rather than straight into the downstream system, and that the cutoff is decided from measured accuracy rather than picked as a round number.

for a middle

Explain the tradeoff explicitly: raising the cutoff increases review volume, lowering it increases error among auto-posted fields. Describe sweeping the threshold on a labelled sample and reporting both numbers at each candidate point.

for a senior

Demonstrate production practice: per-field thresholds driven by cost of error, a composite confidence built from validations and grounding rather than one score, and continuous random audits of the auto-posted stream to catch drift.

for a principal

Own the operating point as a business decision with a stated automation rate and residual error rate per field, backed by the labelling budget it requires. Be ready to argue for improving upstream capture instead of buying more review headcount.

## The threshold is an operating point, not a constant Every extraction pipeline that touches money or goods ends up with a review tier. The engineering question is never "is the model confident" in the abstract; it is where to put the cut between auto-post and human review, and that is a choice on a curve you have to measure. The curve has two axes. As the threshold rises, more fields route to review — that is labour cost and latency. As it falls, more fields auto-post — and the error rate among auto-posted fields rises, because you are admitting progressively weaker evidence. You cannot optimise both. What you can do is quantify the exchange rate and let the business pick the point. ## Building the sample you tune against The measurement needs ground truth: a labelled set drawn from real traffic, not from the clean documents someone had handy. Stratify it by the dimensions that actually move accuracy — document class, source channel (born-digital versus fax versus phone photo), vendor template, and language. A few hundred documents per meaningful stratum is usually enough to place a threshold; the strata matter more than the total. Then sweep. For each cutoff, compute review rate and auto-post error rate per field. The result is often surprising and always useful: some fields are safe to auto-post at almost any threshold, while one or two carry nearly all the residual error and would need a cutoff so high that nothing auto-posts. Those fields should simply be routed to review unconditionally, or dropped from automation until the upstream capture improves. ## Per-field thresholds, because the cost of error is not uniform A single global threshold is the most common mistake. On an arrival notice, getting a shipper's address slightly wrong is an annoyance someone fixes later; getting a container number wrong misroutes a physical box, and getting a dangerous-goods checkbox wrong is a regulatory event. The threshold should reflect the cost of a wrong value passing, which means per-field, sometimes per-field-per-document-class. The same logic applies to routing granularity. Sometimes one weak field should send only that field for review, with the rest posting; sometimes a weak field on a legally-signed document means the whole document is held. Both are legitimate; decide deliberately rather than inheriting whatever the framework does. ## Confidence should be a composite Relying on a single model-emitted score is fragile, and on the model tiers that read messy documents best, that score often does not exist. Build a composite from independent signals: - **Recognition confidence** from a classical OCR pass, where you run one. It is genuinely informative about smudged, skewed, or low-contrast glyphs. - **Validations from the domain.** These are the strongest signals available because they do not depend on the model at all: a shipping-container identifier carries a check digit, line-item amounts must sum to the stated total, a date must fall inside a plausible window relative to the sailing date, an enum value must be in the enum. A failed validation should override a high confidence score outright. - **Grounding checks.** A field whose value cannot be located in the page's recognised text, or whose box falls on whitespace, is a hallucination candidate regardless of what score accompanied it. - **Agreement between extractors.** Running a second, cheaper model — or the same model on a re-rendered page — and comparing gives a disagreement signal that correlates well with error. It costs a second pass, so it is usually reserved for high-stakes fields. A note on calibration: raw model scores are rarely calibrated probabilities, so do not read 0.9 as "90% likely correct". What you need is that the score *ranks* correctness usefully; the mapping from score to actual error rate comes from your labelled sweep, not from the number's face value. ## Keeping it honest in production A threshold tuned once decays. Input quality shifts when a customer swaps scanners, a carrier redesigns a template, or a model version changes. Two practices keep it honest. First, **sample the auto-posted stream**. Pull a small random slice of documents that were never reviewed, have a human check them, and track the measured error rate over time. Without this you are blind to exactly the population you care about — reviewed documents tell you nothing about what slipped through. Second, **capture reviewer corrections as labels**. Every correction is a labelled example, and it costs nothing extra to store. Over months that stream becomes your refreshed evaluation set, your drift detector, and — if you go that way — training data. Beware its bias though: corrections come only from documents that were routed to review, so this stream over-represents hard cases and cannot replace the random audit. The outcome to aim for is a stated, measured operating point per field: "printed container numbers auto-post with a measured residual error of 0.3%, handwritten weights and the hazardous checkbox always go to review, and 11% of documents touch a human." That is an answer an operations lead can act on, which is what the question is really asking for.

  • Why is auditing a random sample of auto-posted documents essential, given you already review the low-confidence ones?
    Because reviewed documents tell you nothing about the population that bypassed review. The residual error rate among auto-posted fields is the number the business is exposed to, and it is only observable by sampling that stream. It is also the first place drift shows up when scan quality or a model version changes.
  • Reviewer corrections accumulate for free. Why not just use them as your evaluation set?
    They are biased by construction: corrections exist only for documents the threshold routed to review, so the set over-represents hard cases and contains no examples of confident errors that slipped through. Use them for drift signals and training candidates, but keep an unbiased random sample as the set you tune the threshold on.
  • A model reports 0.95 confidence on a field. Can you treat that as a 5% error rate?
    No. Model-reported scores are rarely calibrated probabilities. What matters is whether the score ranks correct extractions above incorrect ones; the actual error rate at each score band has to be measured on your labelled sample. Publish the measured mapping and tune against that, never against the score's face value.

saying these in an interview costs you the question

  • Using one global threshold for every field regardless of error cost
  • Treating a model confidence score as a calibrated probability
  • Setting the cutoff by intuition or by whatever review capacity exists
  • Never auditing documents that auto-posted without review
  • Ignoring domain validations like check digits and totals as signals
  • Tuning once and assuming the operating point holds as inputs drift

context