How does ColPali-style late interaction score a page image against a text query?
answer
- ColBERT's idea moved to page images
- many vectors per page, not one
- best-matching patch per query token
- sum of maxima, not an average
- no OCR anywhere in the pipeline
basics
~20 sThe page image is stored as many patch vectors instead of one. Each query token vector is matched against every patch, the best match per token is kept, and those maxima are summed — the MaxSim score. No OCR is involved.
solid answer
~50 sLate interaction borrows ColBERT's scoring and applies it to page images. At index time the page is rendered, passed through a vision-language encoder, and the **per-patch** vectors are kept — typically around a thousand small vectors for one page — rather than pooled into a single embedding. At query time the text query is encoded into per-token vectors in the same space. The score is `MaxSim`: for each query token, take the highest cosine similarity against any patch of that page, then sum those maxima across query tokens. "Late" means query and document never meet until scoring time, so documents stay precomputable while the interaction stays fine-grained. The practical consequence is that a single query term can latch onto one small region — an axis label, a bar, an annotation — which is exactly why the family dominates on figure-heavy pages that OCR would flatten.
code
python · 11 linesimport numpy as np
rng = np.random.default_rng(0)
q = rng.random((12, 128)) # 12 query token vectors
p = rng.random((1030, 128)) # ~1030 patch vectors for one page image
q /= np.linalg.norm(q, axis=1, keepdims=True)
p /= np.linalg.norm(p, axis=1, keepdims=True)
sims = q @ p.T # 12 x 1030 cosine similarities
score = sims.max(axis=1).sum() # MaxSim: best patch per query token, summed
print(round(float(score), 3))go deeper
Recall that the page is stored as many small vectors rather than one, and that the score keeps the best-matching patch for each query word. Say clearly that no text extraction step is involved.
Be ready to write MaxSim out: max over patches per query token, summed over tokens, and explain why max rather than mean is the load-bearing choice.
Explain the two-stage serving pattern that makes it affordable, the storage mitigations (pooling, quantisation), and the corpora where a plain text pipeline still wins on cost.
Own the position of late interaction on the no-interaction to full-interaction spectrum, and argue when the precomputability it preserves is worth the index it costs relative to a cross-encoder rerank.
## Where the term comes from Retrieval architectures sit on a spectrum of *when* the query meets the document. **No interaction** (bi-encoders) encodes each side into one vector independently and compares with a dot product — fast, but everything about a document is squeezed into one point. **Full interaction** (cross-encoders) feeds query and document through a model together — accurate, but you cannot precompute anything, so it only works as a reranker over a handful of candidates. **Late interaction** sits between: both sides are encoded independently, so documents are still precomputed and indexed, but each side keeps *many* vectors and the comparison happens between all of them at scoring time. ColBERT introduced this for text; ColPali carried it to page images, and ColQwen2.5, ColNomic and ColSmolVLM continue the family. ## What is stored per page The page is rendered as an image — no OCR, no layout parser, no text extraction anywhere in the pipeline. A vision-language backbone splits it into patches and produces one embedding per patch, which is what the vision tower naturally emits before pooling. A ColPali-style model projects those to a small dimensionality (128 is typical) and stores all of them. For a full page that is on the order of 1,000 vectors. Everything the encoder saw about layout and figures is preserved as a spatially indexed bag of embeddings. ## The MaxSim score For query token vectors q1..qn and page patch vectors p1..pm, the score is: `score(Q, D) = sum over i of ( max over j of cosine(qi, pj) )` Read it as: every query token independently shops the page for its single best-matching region, and the page's score is the total of those best matches. Two properties follow. First, it is **max, not mean**, so a page is not penalised for being mostly irrelevant. A slide where 950 of 1,000 patches are whitespace and 50 patches are the survival curve still scores highly for a query about the curve. Mean pooling would drown that signal — which is precisely what a single-vector embedding does implicitly. Second, it is a **sum over query tokens**, so every part of the query must find *something*. A multi-term query like "crossover point in the Kaplan-Meier curve" only scores well if each term finds a home somewhere on the page. This gives late interaction its term-coverage behaviour, closer to a keyword system than a dense bi-encoder, while still matching semantically rather than by string. MaxSim is asymmetric by construction: it maximises over document patches for each query token, not the other way round. That asymmetry is intentional — the query is short and every token matters; the page is long and most of it is irrelevant. ## Costs and how systems absorb them Naively, scoring one page is an n×m similarity matrix. Across millions of pages that is not viable as a first stage, so production systems run two stages: an approximate step retrieves candidate pages (for example by searching over the individual patch vectors and gathering their parent pages, or over a pooled representation), then full MaxSim reranks only those few hundred candidates. Storage is the other cost — hundreds to thousands of vectors per page — mitigated by pooling patches down to a few dozen and by binary or scalar quantisation of each vector. ## Why it works so well on documents The alternative pipeline extracts text and embeds it. On a slide whose meaning lives in a plot, a hand-drawn annotation, or the alignment of a table, extraction produces either nothing or a mangled linearisation. Late interaction never asks "what does this page say in words?" — it asks "does any region of this page look like what the query is asking for?", which is a question the pixels can answer. It also removes the entire OCR stage, which is usually the most brittle and most expensive part of a document pipeline. ## Where it does not help On clean, text-dense, single-column documents, a good text pipeline is competitive and vastly cheaper. Late interaction also does not reason: it retrieves the right page, but a model still has to read the page to produce an answer. And it inherits its ceiling from the vision encoder's resolution — detail too small for the patch grid to resolve is not recoverable by clever scoring.
- Why max over patches rather than mean?Because relevance on a page is local. A slide is mostly whitespace and unrelated panels; averaging similarity over every patch would dilute the one region that actually answers the query until it is indistinguishable from noise. Taking the maximum lets a small, highly relevant region carry the page, which is the whole reason the architecture beats single-vector pooling on figure-heavy documents.
- How do you keep MaxSim tractable over millions of pages?You do not run it as the first stage. Systems retrieve candidates approximately — searching the individual patch vectors and gathering their parent pages, or searching a pooled per-page representation — and then apply full MaxSim only to a few hundred candidates as a rerank. Storage is separately reduced by pooling patches down to a few dozen per page and quantising the vectors.
- Does this replace the vision-language model that answers the question?No. Late interaction only ranks pages; it produces no answer. The retrieved page images still go into a vision-language model's context so it can read the chart and respond. The two are complementary — retrieval decides which pages the generator is allowed to see, and the generator decides what they mean.
Instead of summarising a poster in one sentence and comparing sentences, you let each word of the question wander the poster and point at the one spot that matches it best, then add up how well each word did.
saying these in an interview costs you the question
- Saying late interaction runs OCR before embedding
- Describing the score as an average similarity across patches
- Confusing it with a cross-encoder that reads query and page together
- Assuming one vector per page is stored
- Claiming it answers the question rather than ranking pages