skip to content

How would you prove page-image retrieval beats an OCR-then-embed baseline on figure-heavy decks?

level: seniorimportance: should knowfreq 35%

answer

  1. same corpus, same queries, two indexes
  2. nDCG@5 and recall at your context budget
  3. public benchmark shortlists, own set decides
  4. slice by page type
  5. report cost beside quality

basics

~20 s

Run both pipelines over the same pages with the same queries and compare ranking metrics such as nDCG@5 and recall@k. Use a ViDoRe-style visual-document benchmark for a sanity check, but decide on a labelled set built from your own corpus, sliced by page type.

solid answer

~50 s

I would treat it as a controlled comparison, not a demo. Same corpus, same queries, same k, two indexes: one built by extracting text and embedding it, one built from the rendered page images. Report nDCG@5 and recall@k with confidence intervals, plus index size, ingest cost and p95 latency, because a recall win that triples the storage bill is a decision, not a victory. ViDoRe is the public benchmark for exactly this comparison and is worth running as a sanity check on model choice, but public numbers do not transfer: my slides, my query distribution, my failure modes. So I build a labelled set from real user questions and mark, for each, which page actually answers it. The critical move is **slicing** — text-dense pages, dense multi-panel slides, chart-only pages, scanned or handwritten pages. The aggregate almost always understates the win, because page-image retrieval's advantage is concentrated in the slices where extraction produced nothing.

go deeper

for a junior

Know that you compare the two pipelines on the same queries and pages using ranking metrics like recall@k, rather than judging by a few example searches.

for a middle

Be ready to name nDCG@5 and recall@k, explain that k is set by how many page images fit in the generator's context, and mention ViDoRe as the public suite for this comparison.

for a senior

Show the experimental discipline: everything downstream held constant, an in-house labelled set mined from real questions, per-page-type slices, and cost reported alongside quality.

for a principal

Own the decision the evaluation feeds: whether a slice-level win justifies a routed two-index architecture, how the query set is versioned against future model swaps, and what you would tell the budget holder.

## Why this comparison needs discipline Both pipelines will retrieve *something* for every query, and both will look plausible in a demo. The difference lives in the tail: queries whose answer sits in a figure that produced no extractable text. If your evaluation set is drawn from whatever questions were easy to write, it will over-sample text-answerable queries and the two pipelines will look equivalent. The evaluation design is therefore the whole exercise. ## The two pipelines, held constant The **baseline**: render or parse each page, extract its text, embed the text, index it. The **candidate**: render each page as an image and index it with a page-image embedding model or a late-interaction model. Everything downstream must be identical — same query set, same k, same tie-breaking, same hardware for latency numbers. If you change the generator model at the same time you change the retriever, you have measured nothing. ## Metrics Retrieval is a ranking problem, so use ranking metrics. **nDCG@k** rewards putting the right page high and is the standard headline for visual-document retrieval; ViDoRe reports nDCG@5. **Recall@k** answers the question that actually matters downstream: is the answering page inside the k pages I can afford to put into the generator's context? If your budget allows five page images, recall@5 is the number that decides whether the system can possibly answer. **MRR** is fine when exactly one page is correct. Report quality *next to* cost: index bytes per 1,000 pages, ingest cost, and p95 query latency. A pipeline that wins nDCG by four points and costs fifty times the index is a tradeoff to argue, not a result to celebrate. ## The public benchmark and its limits ViDoRe (the Visual Document Retrieval benchmark) exists precisely to score retrieval over page images against text-extraction baselines across figure-heavy, table-heavy and multilingual documents, and it is what the ColPali family and the unified page-embedding models report on. Use it to shortlist models — it will tell you quickly whether a candidate is in the right class. Do not use it to decide. Public benchmarks suffer contamination and saturation, their query distribution is not yours, and leaderboard position is increasingly a fine distinction between models that behave differently on real corpora. In 2026 the top of ViDoRe is crowded; the gaps that matter are on your data. ## Building the in-house set Mine real questions — support tickets, search logs, the questions the medical-affairs team actually asks — rather than inventing them, because invented queries drift toward the vocabulary that appears in the text. Aim for a few hundred queries; that is usually enough to separate pipelines that differ meaningfully, and small enough that a human can label honestly. Label by identifying the page(s) that genuinely answer each query. Have a second annotator label an overlapping subset and check agreement — if two people disagree about which page answers a question, the metric is measuring annotation noise. ## Slicing is where the answer appears Tag every query or page with a type: text-dense prose, dense multi-panel slide, chart-only, table-heavy, scanned or handwritten, non-English. Then report per slice. The expected pattern: near-parity on text-dense pages, where a good text pipeline is genuinely competitive and much cheaper; a large gap on chart-only and handwritten pages, where extraction returns a title and nothing else. A single aggregate number averages those together and can easily show a modest overall win that hides a decisive one on the slice you care about. It also tells you the cheapest architecture: route by page type and pay for the expensive index only where it earns its keep. ## Diagnosing rather than scoring Beyond metrics, read the failures. For each query the baseline missed, look at what its extraction actually produced for the correct page — usually a title, a footnote and nothing else, which is the concrete artifact that convinces a sceptical stakeholder far better than a decimal point. For each query the image pipeline missed, check whether the detail was simply too small for the encoder's patch resolution, which is a ceiling no scoring change will lift. ## Guarding the result Fix the query set and version it, so later model swaps are compared against the same yardstick. Re-check periodically for drift as the corpus grows and users' questions change. And state your assumptions in the write-up — extraction quality, k, hardware — because the most common way this comparison goes wrong is that someone tuned one side and not the other.

  • Why is recall@k more decision-relevant than nDCG here?
    Because k is set by the generator's context budget. If I can afford to put five page images in front of the model, then recall@5 is a hard gate: below it the system cannot answer at all, no matter how good the ranking is above. nDCG tells me how well the ordering is working, which matters for user-facing result lists, but recall@k is what predicts end-to-end answer quality in a RAG setting.
  • Your aggregate shows only a two-point nDCG win. Do you ship the more expensive pipeline?
    Not on that number alone. I would look at the slices: if the two points are the average of parity on text pages and a twenty-point win on chart-only pages, and chart queries are what users actually ask, the aggregate is misleading. The likely answer is a routed system — cheap index by default, expensive index over the figure-heavy subset — which captures most of the win at a fraction of the cost.
  • How many labelled queries do you need?
    A few hundred is usually enough to separate pipelines that differ meaningfully, and it keeps labelling honest enough to trust. What matters more than raw count is coverage: every page-type slice needs enough queries to carry its own confidence interval, otherwise the slice-level story you build the decision on is noise. I would rather have 300 well-distributed queries than 3,000 drawn from one document type.
  • What confounder most often invalidates this comparison?
    Tuning one side and not the other. A weak extraction step, a mismatched text embedding model, or a different k on the baseline manufactures a win for the image pipeline that will not survive contact with a sceptic. I hold everything downstream identical, use a genuinely competitive text pipeline as the baseline, and state extraction quality explicitly in the write-up.

saying these in an interview costs you the question

  • Deciding from public leaderboard position alone
  • Reporting one aggregate number with no page-type slices
  • Writing evaluation queries from the extracted text
  • Comparing quality without reporting index size or latency
  • Changing retriever and generator in the same experiment

context