skip to content

In document extraction, when do you pick classic OCR over a document VLM or a frontier model?

level: middleimportance: must knowfreq 65%

answer

  1. three tiers, not two
  2. cost, determinism, and structure pull different ways
  3. glyph readers versus layout readers versus reasoners
  4. fluent wrong answers versus visible failures
  5. choose per document class, prove on labelled data

basics

~20 s

Classic OCR wins on clean, fixed layouts at high volume: cheap, fast, deterministic, and it returns per-word boxes and confidence scores. Purpose-built document VLMs handle messy layout and tables; frontier models are for open-ended reasoning over the page.

solid answer

~50 s

Treat it as three tiers, chosen per document class rather than per company. **Classic OCR engines** are the cheapest and most predictable option and give per-word boxes and confidences for free, but they only read glyphs — structure has to come from layout rules you write and maintain, which shatter when a vendor moves a field. **Purpose-built document VLMs** — the Mistral OCR / olmOCR / dots.ocr / Granite-Docling class, several small enough to self-host — read the page and emit layout-aware output (markdown or HTML with tables and blocks) at a per-page cost far below a frontier model; they do most bulk enterprise work as of mid-2026. **Frontier multimodal models** earn their price only when the task needs open-ended reasoning over the page — an ambiguous clause, a judgement about which of three dates is the arrival date. A good probe for separating them is one document: a line-item table that breaks across pages with a repeated header.

go deeper

for a junior

Know that OCR turns pixels into characters and that newer document models also recover layout and tables. Be able to say that clean fixed forms suit a cheap OCR engine and messy varied layouts suit a document model.

for a middle

Explain the three tiers and what each returns: glyphs with boxes and confidences, layout-aware markdown or HTML, and reasoning over the page. Name the cost, latency, and determinism differences and give a document that separates them.

for a senior

Demonstrate that you pick per document class from measured numbers on a stratified labelled sample, and that you match controls to failure mode — thresholds for visible OCR failures, grounding and validation for silent VLM hallucinations.

for a principal

Own the portfolio decision: a mixed deployment with self-hosted tiers for bulk and a frontier tier reserved for exceptions, plus the data-residency, vendor-lock-in and re-benchmarking cadence that comes with a field where the model lineup turns over every few months.

## Three tiers, not a binary The old framing was "OCR or an LLM". By mid-2026 the honest framing has three tiers, and the interesting engineering is deciding which document class goes to which. **Tier 1 — classical OCR engines** (Tesseract, cloud OCR services). These recognise glyphs and return words with bounding boxes and per-word confidence scores. They are cheap enough to be effectively free at volume, run in milliseconds to low seconds per page, are fully deterministic, and can be self-hosted with no data leaving your network. What they do *not* do is understand the document: a bill of lading is, to them, a bag of positioned words. Structure comes from code you write — anchor on the label "Container No.", take the text to its right, within this y-band. Those rules work brilliantly on a fixed form and break the day a carrier reflows their template. **Tier 2 — purpose-built document VLMs.** A distinct class of models trained specifically for document conversion: Mistral OCR, olmOCR, dots.ocr, PaddleOCR-VL, GOT-OCR 2.0, DeepSeek-OCR, Granite-Docling. They take a page image and emit layout-aware text — markdown or HTML with headings, reading order, and tables reconstructed as tables. Several are small (low single-digit billions of parameters) and openly licensed, so they self-host on modest GPUs. Their per-page cost sits far below a frontier model while their layout fidelity sits far above classical OCR. This tier now carries most bulk enterprise document work. **Tier 3 — general frontier multimodal models.** These read the page *and* reason about it in the same call: apply a policy, resolve which of three dates is the one you asked for, notice that a clause contradicts a field. They are the most expensive per page and the slowest, and they typically do not give you the per-word geometry that tier 1 hands over for free. ## The probe that separates them Take a multi-page line-item table with a header repeated on every sheet — a freight arrival notice is the classic case. Classical OCR returns correct words in the wrong shape: no notion that the table continues, and the repeated header appears as data. A document VLM usually reconstructs the table on each page and marks the header, but you still need logic to detect the repeat and merge the continuation rows. A frontier model can be *told* the header repeats and asked to emit one merged table — and will often do it — but you are paying reasoning prices for a formatting problem, and it may quietly renumber or drop a row under length pressure. That single document tells you more than a benchmark score, because it exercises the exact capability the tiers differ on: carrying structure across a page break rather than reading glyphs. ## The failure modes differ, and that changes your controls Classical OCR fails *visibly*: a smudged digit becomes a wrong character with a low confidence score attached, and your rule finds no anchor and returns nothing. VLM tiers fail *silently and plausibly*: a model that cannot read a fax-degraded container number may emit a well-formed container number that was never on the page. The mitigations therefore differ. Tier 1 needs confidence thresholds and anchor-miss alarms; tiers 2 and 3 need schema constraints, grounding back to the page, cross-field validation, and sampling — because a fluent wrong answer passes every syntactic check you have. ## Choosing per class, and measuring it A freight forwarder does not have "a document problem" — it has six of them. Born-digital invoices from three named carriers are a tier-1 problem forever. Faxed, skewed bills of lading with variable layouts are a tier-2 problem. A one-off dispute where someone must read a hand-annotated clause and decide whether it changes liability is a tier-3 problem, at ten documents a month where the price is irrelevant. Decide with numbers, on a labelled sample drawn from real traffic and stratified by document class and quality. Measure character error rate for raw reading, a table-structure metric such as TEDS for reconstructed tables, and — the one that actually matters — end-to-end field-level accuracy on the fields you post downstream, since a pipeline can be mediocre at reading prose while being perfect on the twelve fields you care about. Then put cost and latency per page next to those numbers. The frequent result is a mixed deployment: tier 1 for the predictable bulk, tier 2 for the messy majority, tier 3 reserved for exceptions and for the fields where a mistake is expensive.

  • You measured a document VLM at 99% character accuracy. Why might that number still mislead you?
    Character accuracy averages over the whole page, where most characters are boilerplate. The 1% that is wrong can be concentrated exactly in the dense, low-contrast fields you extract — container numbers, weights, amounts. Always report field-level accuracy on the fields you post downstream, sliced by document class and scan quality, alongside the character metric.
  • What do you lose by moving from a classical OCR engine to a document VLM?
    Chiefly two things: calibrated per-word confidence and cheap determinism. Classical engines hand you a score per word that you can threshold directly, and the same input yields the same output every run. VLM tiers give you fluent output with no native confidence signal, so you have to manufacture one from validation rules, cross-model agreement, or sampling.
  • When is a frontier model the right choice despite the price?
    When the task is genuinely reasoning, not transcription — resolving which of several dates satisfies a policy, judging whether an annotation modifies a clause, handling a long tail of layouts no rule set can cover — and when volume is low enough that per-page cost does not dominate. It is also a reasonable bootstrap for labelling a gold set that a cheaper tier is then measured against.

saying these in an interview costs you the question

  • Assuming a frontier model beats classical OCR on every document
  • Ignoring that VLM tiers hallucinate plausible field values silently
  • Benchmarking on a handful of clean samples instead of stratified labelled data
  • Choosing one tier for the whole company rather than per document class
  • Expecting per-word confidence scores from a VLM tier
  • Reporting character accuracy when field-level accuracy is what ships

context