An approved invoice record's extracted fields differ from the PDF page — what does the approval log prove?
answer
- a click is a decision, not a comparison
- which artefact was on their screen?
- the approval is the effect, not the evidence
- compare emitted against displayed, per field
basics
~20 sIt proves a decision was recorded against a record by an actor at a time. It does not prove which artefact was on screen; in this flow the approver usually reads the extracted summary, not the page.
solid answer
~40 sThe log proves a decision was recorded against a record id, by an actor, at a time — and nothing about which artefact that person had in front of them. In an accounts-payable flow the approver typically reads the extracted field summary rather than opening the pages, so the click can mean the extracted values were checked against nothing. When triaging, the evidence you can actually collect is the mismatch itself: what the page displays beside what extraction emitted, and which source field each emitted value came from. Treat the approval as the thing that made the outcome durable rather than as evidence that the two artefacts agreed. The finding is not *a model was tricked*; it is that two different texts were treated as one and a signature landed on the wrong one.
code
text · 10 linesvendor-invoice.pdf
render -> what the page paints: line items, totals, remit-to block
extract -> what the stage emitted: page text
+ form-field values
+ annotation contents <- [directive span elided]
approve -> recorded decision: actor = ap-approver
record = INV-0042
artefact on screen = extracted field summary
outcome = approved
...go deeper
Be able to say that a recorded approval shows who decided and when, and that it does not record which version of a document the person was looking at.
Explain why an approver reading an extracted summary is downstream of the divergence, so their sign-off cannot surface a page-versus-extraction mismatch.
Show how you would build the finding: a per-field comparison of emitted value, source container and displayed page content, with an honest statement of what reproduced and how often.
Own the distinction between evidence and effect. The approval is what makes the outcome durable and drives severity, while proving nothing about coverage — be ready to hold both of those in one sentence to an owner who wants only the first half.
## The triage question Somebody files a report: the vendor name and remit-to values on an approved invoice record do not match what the PDF's pages display, and a directive-shaped run sits in a container the pages never paint. The process owner's first response is that the record was approved by a named person, so a human was in the loop. You have to say precisely what that approval does and does not establish. ## What a record proves, item by item | Artefact | What it establishes | What it does not | | --- | --- | --- | | Approval row | A decision was recorded against a record id, by an actor, at a time | Which text the actor read, or that they compared anything | | Extraction output | Which string the model received, and from which source field | That the string appears on any page | | Rendered page | What a viewer paints for a reader | What the extractor emitted from the same bytes | | Model's own account of its reasoning | Text the model produced afterwards | A reliable record of what influenced it | The habit worth demonstrating is refusing to let one row stand in for a different one. This is the same discipline as elsewhere in the domain: a passing screen proves the inspected artefact scored below a threshold; a tool-call log proves which call ran, not who chose the argument values; an approval proves a decision was captured, not that a comparison happened. ## Why the approval matters anyway Not as evidence — as **effect**. The construction's payoff was never the model's output on its own; it was a sanctioned downstream record. The approval is what converts an extraction discrepancy into an authorised payment nobody overrode, and it is why the impact section of the finding is not hypothetical. So the click is central to severity and irrelevant to whether the two artefacts agreed. Candidates frequently invert exactly this. ## Building a finding somebody will act on What you can put in front of an owner is a per-field comparison: the value the extractor emitted, the source container it came from, and what the rendered page shows in the corresponding place. That comparison is reproducible, does not depend on the model behaving the same way twice, and does not require anybody to remember what they looked at three weeks ago. If your report leans on the approver's memory or on the model reproducing an output, it will be argued away. ## Reproduction honesty A model is probabilistic, so the same file may extract the same way most times and differently occasionally. A single reproduction proves the construction worked once against one deployment. That is worth stating in the finding rather than being caught by it — and it is one reason the *artefact mismatch* is the better evidence than the *model output*: the mismatch between emitted and displayed content is a deterministic property of the file and the pipeline, not of one generation. ## The one-line answer The log records that somebody signed. Which document they read is not in it, and in this pipeline the likeliest answer is: not the one the model got.
- The owner says a human was in the loop. What is the accurate correction?That a human was in the loop over the extracted summary, which is downstream of the discrepancy, not over the artefact where the two texts diverge. Being in the loop describes a position in a workflow, not coverage. State which artefact the person reads today, and note that the loop as built cannot surface a difference between the page and the emitted text because it never puts both in one view.
- The construction reproduced once in five attempts. Is that a finding?Yes, but you must report it as what it is. The extraction discrepancy — a container emitted into the model's input that no page paints — is deterministic and reproduces every time; only the model's downstream behaviour is variable. Lead with the deterministic half, state the observed rate for the rest, and do not let one lucky reproduction be described as reliable.
saying these in an interview costs you the question
- Reads an approval row as proof the human compared page and fields
- Treats the model's own explanation of its reasoning as evidence
- Says a human was in the loop without naming which artefact they read
- Builds the finding on a model output rather than the artefact mismatch
- Reports one successful reproduction as a reliable result