Caption-and-index, page-image embeddings, or late interaction — how do you pick for a figure-heavy slide corpus?
answer
- recall against index size
- what the ingest step throws away
- one vector per page versus a thousand
- pooling sits between the extremes
- measure on your own corpus first
basics
~20 sThree designs trade recall for cost: captioning is cheapest but loses whatever the caption omits; one vector per page image is cheap and strong on most queries; late interaction wins on visually hard queries at a far larger index.
solid answer
~50 sI frame it as a recall-versus-storage decision on a corpus I have measured, not a preference. **Caption-and-index** runs each page through a vision model, writes a text description, and embeds that text — cheap, reuses text infrastructure, but the caption is a lossy summary written before anyone knew the query, so anything it omits is permanently unretrievable. **Unified page-image embeddings** (Cohere Embed v4, voyage-multimodal-3.x) take the rendered page straight in and return one vector, so the index is the same size as a text index and nothing is thrown away at ingest. **Late interaction** (ColPali, ColQwen2.5, ColNomic) keeps roughly a thousand patch vectors per page and scores with MaxSim — the best recall on figure-heavy material and the largest index by two orders of magnitude. My default is single-vector first, then measure: if the queries that fail are visually hard, pay for late interaction on that slice.
code
python · 4 linespages = 40_000
multi = pages * 1030 * 128 * 2 # ~1030 patch vectors per page, fp16
single = pages * 1536 * 4 # one 1536-d fp32 vector per page
print(round(multi / 1e9, 2), "GB vs", round(single / 1e9, 3), "GB")go deeper
Know the three names and the one-line difference: describe the page in text, embed the page image as one vector, or keep many per-page vectors. Say plainly that captioning discards detail before anyone asks a question.
Be ready to explain why a single vector struggles on a dense multi-panel slide, and to give rough index-size arithmetic for the multi-vector option versus the single-vector one.
Show how you would decide empirically: a labelled set from the real corpus, recall at the k you can afford, and cost reported next to quality. Name pooling and quantisation as the middle ground rather than picking an extreme.
Own the tradeoff across corpus sizes and lifetimes — what re-embedding 40,000 pages costs when models change, whether to route by query type, and how you would justify a two-order-of-magnitude index to whoever pays for it.
## The problem this decision exists to solve Take a medical-affairs team with 40,000 conference slides. A typical answer is a Kaplan-Meier survival curve with a hand-drawn annotation, sitting on a slide whose only extractable text is a title and a footnote. A text-only pipeline indexes the title and the footnote, so the page is effectively invisible to any query about the shape of the curve. Multimodal retrieval exists because the information is *in the pixels*, and every architecture below is a different answer to "how much of the page do we keep, and at what price?" ## Design 1 — caption-and-index Render each page, ask a vision-language model to describe it, store the description, embed the description with an ordinary text model. Retrieval is then plain text retrieval and every tool you already own keeps working: keyword hybrid, rerankers, filters, existing dashboards. The fatal property is that captioning is **lossy compression performed before the query is known**. A caption that says "survival curve comparing two arms" cannot answer "which slide shows the crossover at month 18?" Adding detail helps only until the caption itself becomes long, expensive to generate, and diluted when embedded into a single vector. Ingest cost is also real: one VLM call per page, times 40,000, times every re-caption when you change models. It is still the right choice when pages are text-dominant, when you need the caption text for other purposes anyway, or when your retrieval stack cannot store image-derived vectors at all. ## Design 2 — unified single-vector page embeddings A newer class of embedding models ingests the page image directly (optionally with any text) and emits **one** vector in a space shared with text queries. Index size per page is identical to a text chunk, latency is one nearest-neighbour lookup, and there is no captioning step to go stale. Because the model saw the pixels, layout and figure content influence the vector in ways a caption never captures. The limit is representational: one vector must summarise a whole dense page. On a slide with four unrelated panels, the single vector is a blend, and a query targeting one panel competes with the other three. Some of these models expose Matryoshka-style truncatable dimensions, so you can trade a little accuracy for a smaller index without changing models. ## Design 3 — late interaction over page images ColPali-style models keep the per-patch vectors the vision encoder already produced — on the order of a thousand small vectors per page — and score a query by summing, over each query token, its best match against any patch (MaxSim). No OCR anywhere in the pipeline. This preserves fine-grained, localised evidence, which is exactly why it dominates on figure-heavy documents. The bill is the point of the question. A rough arithmetic: 40,000 pages × ~1,000 patch vectors × 128 dims × 2 bytes ≈ 10 GB of raw vectors, against ~0.25 GB for one 1,536-dim float vector per page. Scoring is also heavier — you normally run an approximate first stage to fetch candidates, then MaxSim-rerank only those, which adds a stage to operate and tune. ## The pooling compromise The two extremes are not the only options. Patch vectors can be **pooled** — clustered or averaged down from ~1,000 to a few dozen per page — which recovers most of the storage while keeping much of the localisation benefit. Dimension reduction and binary or scalar quantisation cut it further. In practice the interesting engineering is here, not in the binary choice: pooled multi-vector plus quantisation often lands within a point or two of full late interaction at a fraction of the index. ## How to actually decide Build a small labelled set from *your* corpus and your users' real questions, then measure recall at the k you can afford to put in the generator's context. Report cost alongside quality: index GB, ingest cost per 10,000 pages, and p95 query latency. Segment the results — if late interaction only wins on the 20% of queries that are chart-shaped, route by query type or index only the figure-heavy subset that way. Be honest that this is contested ground in 2026: the models on both sides are improving fast, and the right answer for a 5,000-page corpus (take the recall, the storage is free) is not the right answer for 50 million pages. ## Failure signals Captioning is failing when users ask about details that never appear in your captions. Single-vector is failing when the correct page is retrieved at rank 30 rather than rank 3 on dense multi-panel slides. Late interaction is failing operationally, not qualitatively — index growth, rebuild time and rerank latency are what force the retreat.
- Your late-interaction index has grown past what the cluster can hold. What do you cut first?Pooling before anything else: cluster the ~1,000 patch vectors per page down to a few dozen, which typically costs a small amount of recall for a large storage win. Then quantise — binary or scalar — and reduce dimensionality. Only after those would I fall back to single-vector embeddings, and I would do it per-collection so the visually hard corpora keep the expensive representation.
- Would you ever run two of these architectures side by side?Yes, and it is often the pragmatic answer. Run cheap single-vector retrieval over everything as the default path, and maintain a late-interaction index over the subset that measurably needs it — the figure-heavy decks. You can also fuse: take candidates from the cheap index and rerank them with the multi-vector model, which bounds the expensive work to a few dozen pages per query.
- Does caption-and-index have any advantage the other two cannot match?It produces human-readable text as a by-product, which is genuinely useful: captions can be keyword-searched, filtered, shown in results, audited, and reused for summaries or accessibility. Neither embedding-based design gives you an artifact a person can read. If explainability of the index matters as much as recall, captioning earns its place — often alongside, not instead of, an image-based index.
Captioning is filing a photograph by writing a sentence about it; a page embedding is filing the photograph itself; late interaction is filing every square inch separately so any detail can be matched.
saying these in an interview costs you the question
- Assuming captioning is lossless because the model is strong
- Choosing late interaction without pricing the index
- Treating storage cost as fixed rather than tunable by pooling
- Deciding from a public leaderboard instead of the team's own corpus
- Believing one vector per page cannot beat captions