skip to content

A vendor invoice PDF passed human review — how can extracted text still carry a directive the reviewer never saw?

level: juniorimportance: must knowfreq 68%

answer

  1. two artefacts, not one
  2. which one did the human read?
  3. rendering hides, extraction recovers
  4. reviewed as displayed, consumed as extracted

basics

~20 s

The reviewer read the document as rendered; extraction deliberately recovers what rendering omits — form-field values, annotation contents, off-canvas or notice-sized runs. Those are two different texts, and only the extracted one reaches the model.

solid answer

~40 s

Because the reviewer and the model were handed different texts. A rendered page shows what the layout engine puts on screen; an extraction stage exists precisely to recover machine-readable content the layout does not display — form-field values, annotation contents, alternative text, runs positioned off the canvas or set at notice size. A vendor who authors the file chooses which side of that line a span falls on. The review was real, but it was a review of the rendering; nothing in it inspected the string the model actually received. The gap is not a bug in either component: thorough extraction and selective rendering are both doing their jobs, and the party submitting the invoice is simply the one who noticed they disagree.

go deeper

for a junior

Be ready to state that a rendered page and an extracted text are two different artefacts, and that a human review attaches to only one of them. Naming a couple of container kinds is enough; the separation is the point.

for a middle

An interviewer expects you to explain why the two behaviours diverge on purpose — layout displays selectively, extraction recovers thoroughly — and to say that neither component is misbehaving.

for a senior

Show the habit of asking which artefact a claim attaches to before accepting it, and of naming the collectable evidence: what the page displays beside what extraction emitted, per source field.

for a principal

Own the framing that review coverage is defined by artefact, not by document. Be able to say plainly which stage in a pipeline nobody currently reads, and what a sign-off on the other stage can honestly be claimed to mean.

## The claim being corrected The usual answer to *how did that reach the model* is **the document was reviewed before it was indexed**. That answer is not false — a person really did open the file. It is *aimed at the wrong artefact*. The document was reviewed **as rendered**, and the extraction stage exists specifically to recover what rendering leaves out. The distance between those two behaviours is the entire surface. ## Four things called *the document* In an accounts-payable flow where a counterparty uploads an invoice and a model reads the fields an approver signs off on, at least four distinct objects get called the document: - **the bytes** somebody uploaded; - **the rendered pages** a reviewer opens in a viewer; - **the text an extractor emitted** from those bytes; - **the field summary** the approval screen shows. A control, a review or a claim always attaches to exactly one of these. Saying *the document was checked* without saying which one is the mistake, and in practice the human checked the second while the model consumed the third. ## Why the two texts differ by design Rendering is **selective**. A layout engine decides what goes on a page, at what size, in what order, inside a visible canvas. Content that is present in the file but not placed under an eye — a value stored in an interactive form field, the body of an annotation, alternative text attached to a graphic, a run positioned outside the visible area or set at a size a person skims past — simply does not appear. Extraction is **thorough**. Its job is the opposite: recover everything machine-readable so that downstream stages do not miss a value the page happened to lay out awkwardly. A good extractor deliberately reaches into containers the page does not show, because in ordinary documents those containers hold real data. Both behaviours are correct. Neither is a defect. And the author of the file gets to choose which side of the line a given span sits on. ## What the reviewer's approval actually establishes That a person looked at a rendering and found nothing objectionable in it. It establishes nothing about the extracted string, because the reviewer never saw that string. This is the same directional discipline you apply everywhere in this domain: a passing screen proves the artefact it inspected scored below a threshold, not that a different artefact is clean. ## Why this is not about obfuscation A span in a non-displayed container is not encrypted, encoded or disguised. It is plain text sitting in a place a viewer does not paint. A machine reading the bytes can see it perfectly well; that is exactly why the extractor emits it. The asymmetry being exploited is **human visibility versus machine recovery**, not secrecy. Candidates who reach for encoding tricks here have misread the mechanism. ## Where the class does and does not apply The construction depends on the reviewer's artefact and the model's input being two different texts. In a pipeline where the model is given only a rasterised page image, non-displayed containers never reach it and this whole family is inert — a different family applies there instead. Equally, if the string an approver signs is literally the extracted text, the span is on screen and the construction has no hiding place. The interesting question in any real deployment is therefore not *is there hidden text* but *which artefact did the human actually read, and which one did the model get*. ## What an interviewer is scoring Not a list of places to hide things. They want to hear you separate the artefacts, name which stage saw which one, and refuse to let *it was reviewed* stand as an unqualified claim.

  • Does it matter that the span would have been visible in a different viewer?
    Visibility is a property of a rendering, not of the file. Some viewers paint annotation bodies or form-field values that others collapse, so the same file is a different document in two applications. The rendering that mattered is the one the reviewer actually used, and the file's author can only ever bet on typical viewer behaviour — which is part of what the construction costs them.
  • A byte-level scan of the upload would find that text. Why does the class still work?
    It would, and that is the point: nothing is hidden from a machine here. The gap is between machine reading and human reading. A scan that flags every non-displayed field flags most ordinary business documents too, because those containers legitimately hold data, so the signal it produces is weak and the carrier stays cheap to author.
  • How is this different from text that is only created downstream by optical recognition?
    There the text is genuinely absent from the uploaded bytes and a later stage manufactures it, so anything inspecting the upload examined an artefact that did not contain it. Here the text is in the bytes the whole time — a byte scan sees it, a person opening the page does not. Different stage, different gap, different evidence.

A printed page and the file behind it are different objects. The print shows what the layout chose to show; the file still holds everything the layout skipped, and a machine reading the file gets all of it.

saying these in an interview costs you the question

  • Says the document was reviewed, therefore the extracted text is clean
  • Assumes extracted text is the same string the page displays
  • Treats non-displayed text as encrypted or obfuscated
  • Concludes an empty-looking viewer means the file holds nothing
  • Blames the extractor for a bug when thoroughness is its job

context