How would you prove page-image retrieval beats an OCR-then-embed baseline on figure-heavy decks?
answer
- same corpus, same queries, two indexes
- nDCG@5 and recall at your context budget
- public benchmark shortlists, own set decides
- slice by page type
- report cost beside quality
basics
~20 sRun both pipelines over the same pages with the same queries and compare ranking metrics such as nDCG@5 and recall@k. Use a ViDoRe-style visual-document benchmark for a sanity check, but decide on a labelled set built from your own corpus, sliced by page type.
solid answer
~50 sI would treat it as a controlled comparison, not a demo. Same corpus, same queries, same k, two indexes: one built by extracting text and embedding it, one built from the rendered page images. Report nDCG@5 and recall@k with confidence intervals, plus index size, ingest cost and p95 latency, because a recall win that triples the storage bill is a decision, not a victory. ViDoRe is the public benchmark for exactly this comparison and is worth running as a sanity check on model choice, but public numbers do not transfer: my slides, my query distribution, my failure modes. So I build a labelled set from real user questions and mark, for each, which page actually answers it. The critical move is **slicing** — text-dense pages, dense multi-panel slides, chart-only pages, scanned or handwritten pages. The aggregate almost always understates the win, because page-image retrieval's advantage is concentrated in the slices where extraction produced nothing.
go deeper
Know that you compare the two pipelines on the same queries and pages using ranking metrics like recall@k, rather than judging by a few example searches.
Be ready to name nDCG@5 and recall@k, explain that k is set by how many page images fit in the generator's context, and mention ViDoRe as the public suite for this comparison.
Show the experimental discipline: everything downstream held constant, an in-house labelled set mined from real questions, per-page-type slices, and cost reported alongside quality.
Own the decision the evaluation feeds: whether a slice-level win justifies a routed two-index architecture, how the query set is versioned against future model swaps, and what you would tell the budget holder.
## Why this comparison needs discipline Both pipelines will retrieve *something* for every query, and both will look plausible in a demo. The difference lives in the tail: queries whose answer sits in a figure that produced no extractable text. If your evaluation set is drawn from whatever questions were easy to write, it will over-sample text-answerable queries and the two pipelines will look equivalent. The evaluation design is therefore the whole exercise. ## The two pipelines, held constant The **baseline**: render or parse each page, extract its text, embed the text, index it. The **candidate**: render each page as an image and index it with a page-image embedding model or a late-interaction model. Everything downstream must be identical — same query set, same k, same tie-breaking, same hardware for latency numbers. If you change the generator model at the same time you change the retriever, you have measured nothing. ## Metrics Retrieval is a ranking problem, so use ranking metrics. **nDCG@k** rewards putting the right page high and is the standard headline for visual-document retrieval; ViDoRe reports nDCG@5. **Recall@k** answers the question that actually matters downstream: is the answering page inside the k pages I can afford to put into the generator's context? If your budget allows five page images, recall@5 is the number that decides whether the system can possibly answer. **MRR** is fine when exactly one page is correct. Report quality *next to* cost: index bytes per 1,000 pages, ingest cost, and p95 query latency. A pipeline that wins nDCG by four points and costs fifty times the index is a tradeoff to argue, not a result to celebrate. ## The public benchmark and its limits ViDoRe (the Visual Document Retrieval benchmark) exists precisely to score retrieval over page images against text-extraction baselines across figure-heavy, table-heavy and multilingual documents, and it is what the ColPali family and the unified page-embedding models report on. Use it to shortlist models — it will tell you quickly whether a candidate is in the right class. Do not use it to decide. Public benchmarks suffer contamination and saturation, their query distribution is not yours, and leaderboard position is increasingly a fine distinction between models that behave differently on real corpora. In 2026 the top of ViDoRe is crowded; the gaps that matter are on your data. ## Building the in-house set Mine real questions — support tickets, search logs, the questions the medical-affairs team actually asks — rather than inventing them, because invented queries drift toward the vocabulary that appears in the text. Aim for a few hundred queries; that is usually enough to separate pipelines that differ meaningfully, and small enough that a human can label honestly. Label by identifying the page(s) that genuinely answer each query. Have a second annotator label an overlapping subset and check agreement — if two people disagree about which page answers a question, the metric is measuring annotation noise. ## Slicing is where the answer appears Tag every query or page with a type: text-dense prose, dense multi-panel slide, chart-only, table-heavy, scanned or handwritten, non-English. Then report per slice. The expected pattern: near-parity on text-dense pages, where a good text pipeline is genuinely competitive and much cheaper; a large gap on chart-only and handwritten pages, where extraction returns a title and nothing else. A single aggregate number averages those together and can easily show a modest overall win that hides a decisive one on the slice you care about. It also tells you the cheapest architecture: route by page type and pay for the expensive index only where it earns its keep. ## Diagnosing rather than scoring Beyond metrics, read the failures. For each query the baseline missed, look at what its extraction actually produced for the correct page — usually a title, a footnote and nothing else, which is the concrete artifact that convinces a sceptical stakeholder far better than a decimal point. For each query the image pipeline missed, check whether the detail was simply too small for the encoder's patch resolution, which is a ceiling no scoring change will lift. ## Guarding the result Fix the query set and version it, so later model swaps are compared against the same yardstick. Re-check periodically for drift as the corpus grows and users' questions change. And state your assumptions in the write-up — extraction quality, k, hardware — because the most common way this comparison goes wrong is that someone tuned one side and not the other.
- Why is recall@k more decision-relevant than nDCG here?Because k is set by the generator's context budget. If I can afford to put five page images in front of the model, then recall@5 is a hard gate: below it the system cannot answer at all, no matter how good the ranking is above. nDCG tells me how well the ordering is working, which matters for user-facing result lists, but recall@k is what predicts end-to-end answer quality in a RAG setting.
- Your aggregate shows only a two-point nDCG win. Do you ship the more expensive pipeline?Not on that number alone. I would look at the slices: if the two points are the average of parity on text pages and a twenty-point win on chart-only pages, and chart queries are what users actually ask, the aggregate is misleading. The likely answer is a routed system — cheap index by default, expensive index over the figure-heavy subset — which captures most of the win at a fraction of the cost.
- How many labelled queries do you need?A few hundred is usually enough to separate pipelines that differ meaningfully, and it keeps labelling honest enough to trust. What matters more than raw count is coverage: every page-type slice needs enough queries to carry its own confidence interval, otherwise the slice-level story you build the decision on is noise. I would rather have 300 well-distributed queries than 3,000 drawn from one document type.
- What confounder most often invalidates this comparison?Tuning one side and not the other. A weak extraction step, a mismatched text embedding model, or a different k on the baseline manufactures a win for the image pipeline that will not survive contact with a sceptic. I hold everything downstream identical, use a genuinely competitive text pipeline as the baseline, and state extraction quality explicitly in the write-up.
saying these in an interview costs you the question
- Deciding from public leaderboard position alone
- Reporting one aggregate number with no page-type slices
- Writing evaluation queries from the extracted text
- Comparing quality without reporting index size or latency
- Changing retriever and generator in the same experiment