skip to content

In document extraction, why demand a bounding box and page number per extracted field?

level: middleimportance: must knowfreq 56%

answer

  1. a string alone cannot be checked
  2. evidence travels with the value
  3. click the field, land on the pixels
  4. a box over whitespace is an invented value
  5. null is a legitimate answer

basics

~20 s

Grounding turns an unverifiable string into an auditable claim. A clerk jumps straight to the pixels behind "container MSKU4412345", and a box landing on blank space exposes an invented value before it reaches the downstream system.

solid answer

~60 s

Two things make an extraction usable in production: a **fixed schema** and **provenance**. The schema pins the output to a known set of fields with types, enums and an explicit null for "not present on this document", so a missing field is reported rather than invented. Provenance — page number, bounding box, and ideally the literal source text the value came from — is what makes each field checkable. It buys three things at once. Review gets fast: the reviewer clicks the field and lands on the pixels, instead of scanning seven pages. Hallucinations become detectable: a value whose box covers whitespace, or whose box text does not match the value, is almost certainly generated rather than read. And you get an audit trail, which is what a customs or finance auditor actually asks for. One caution: general-purpose models are unreliable at emitting precise pixel coordinates, so the robust pattern is to keep OCR word boxes and derive each field's box by matching the extracted value back to those words.

code

json · 26 lines
json
{
  "document_type": "bill_of_lading",
  "fields": {
    "container_number": {
      "value": "MSKU4412345",
      "source_text": "MSKU 4412345",
      "page": 2,
      "bbox": [412, 318, 596, 344],
      "confidence": 0.94
    },
    "gross_weight_kg": {
      "value": 12500.0,
      "source_text": "12,500.00 KG",
      "page": 2,
      "bbox": [412, 352, 604, 378],
      "confidence": 0.71
    },
    "notify_party": {
      "value": null,
      "source_text": null,
      "page": null,
      "bbox": null,
      "confidence": null
    }
  }
}

go deeper

for a junior

Be able to say that an extracted field should come back with where it was found — page and a box — so a person can check it, and that a missing field should be reported as null rather than guessed.

for a middle

Explain the two halves of the contract: a typed schema with explicit nulls, and provenance per field. Describe the concrete checks grounding enables — box on ink, box text matching the value — and why they catch hallucinations cheaply.

for a senior

Show the robust implementation: keep OCR word boxes and derive field boxes by matching values back, rather than trusting model-emitted coordinates. Add cross-field invariants like check digits and totals, and treat unlocatable values as automatic quarantine.

for a principal

Own provenance as a compliance and evolvability asset: it is the audit answer to "why did you post this figure", and the diff that makes a model upgrade measurable. Weigh its output-size and pipeline cost against the alternative of shipping unverifiable extractions.

## The output contract, not just the output An extraction pipeline that returns free-form prose about a document is not an extraction pipeline. What downstream systems need is a record: named fields, typed values, and enough evidence to defend each one. That contract has two halves. The **schema** fixes what may appear: field names, types (a date is a date, a weight is a number with a unit), enumerations for closed sets like incoterms or container types, and — critically — an explicit representation of absence. A pipeline without an explicit null teaches the model that every field must be filled, which is precisely how you get a plausible consignee on a document that never named one. "Return null when the document does not state it" is one of the highest-value instructions in the whole system. The **provenance** half attaches, to every field, where the value came from: page number, a bounding box in page coordinates, the literal text spanned by that box, and whatever confidence signal the tier provides. ## What grounding buys **Review throughput.** This is the immediate operational win. A clerk verifying a bill of lading with twenty fields across seven scanned pages spends most of their time *finding* the value, not judging it. With a box, the UI highlights the region and the clerk answers yes or no in a second. On a queue of thousands of documents a day, that difference decides whether human review is affordable at all. **Hallucination detection.** The model tiers that read messy documents best are also the ones that fail most plausibly: asked for a container number on a fax-degraded page, a model may emit a well-formed container number that appears nowhere. Grounding gives you cheap, automatic checks. Does the box fall on ink or on whitespace? Does the text inside the box match the value that was returned, allowing for normalisation? Does the box sit on the page the field is supposed to be on? A field that fails these is quarantined without a human ever looking at it. **Audit and dispute.** In freight, customs and finance, someone will eventually ask why your system posted a particular weight. "The model said so" is not an answer. "Page 3, this rectangle, this text, extracted at this time by this model version" is. That record also makes regressions tractable: when accuracy drops after a model change, you can diff not just the values but where they were read from. **Evaluation.** Boxes let you score *reading the right thing* separately from *reading it correctly*. A pipeline that gets the right value from the wrong region is one layout change away from silently breaking, and only provenance exposes that. ## How to get boxes you can trust Here the practical detail matters. General multimodal models are good at reading text and unreliable at emitting exact pixel coordinates — coordinates drift, or arrive in a normalised space you have to guess at. Purpose-built document models in the OCR-VLM tier are far better, because grounded layout output is what they were trained to produce. The most robust pattern does not ask the model for geometry at all. Run a pass that yields word-level boxes — a classical OCR engine, or a document model that emits blocks with coordinates — and keep them. Then take each extracted field value and match it back against those words to derive its box. String matching with normalisation (whitespace, case, punctuation, thousands separators) resolves most fields; fuzzy matching handles the rest. The match itself becomes a signal: a field whose value cannot be located anywhere in the page's recognised text is a hallucination candidate, flagged automatically. ## Schema design choices that pay off Keep the schema flat and shallow where the domain allows; deeply nested structures raise both the error rate and the difficulty of reviewing a single field. Represent repeated structures — line items — as arrays of objects, each item carrying its own provenance, so a reviewer can approve nineteen rows and correct one. Normalise at the boundary, but store the raw string too: "12,500.00 KG" is what the page said, 12500.0 with unit KG is what the system needs, and disagreements between them are a useful alarm. Add cross-field invariants where the domain provides them — a total that must equal the sum of line items, an identifier with a check digit — because a validation failure is a far stronger signal than any model-reported confidence. ## What it costs Grounding is not free. Emitting boxes and source spans adds output size and, in the matching approach, a second pass to maintain. The honest framing is that you pay a modest engineering and runtime cost to convert an unverifiable output into a reviewable one — and on documents that move money or goods, an unverifiable output was never shippable in the first place.

  • The model returns the right value but a box that is off by fifty pixels. Is that a problem?
    Yes, though a lesser one. It breaks the review workflow — the highlight lands on the wrong field and the clerk stops trusting the overlay — and it weakens hallucination detection, because box-versus-value checks start producing false alarms. It is also the symptom of asking a model for coordinates directly; deriving boxes by matching values back to OCR word boxes avoids it.
  • Why insist on an explicit null rather than an empty string for a field the document does not contain?
    Because empty string and "absent" are different facts downstream, and a schema without a sanctioned way to say "not present" pressures the model to fill every field. Explicit null makes absence a first-class, countable outcome — you can monitor how often a field goes missing per vendor, which is often the earliest signal that a template changed.
  • How do cross-field validations compare with model confidence as a quality signal?
    They are stronger, because they are independent of the model. A container identifier's check digit, or line items summing to the stated total, either holds or does not — no calibration required. Model confidence is a soft ranking signal at best. Use validations as hard gates and confidence to prioritise what a human looks at first.

saying these in an interview costs you the question

  • Returning bare field values with no page or coordinates
  • Letting the model fill every field rather than allowing null
  • Asking a general multimodal model for exact pixel coordinates and trusting them
  • Treating a fluent, well-formatted value as evidence it was on the page
  • Discarding the raw source string after normalising the value

context