skip to content

Should retrieved slide images or their extracted text go into the model's context, and what does that cost?

level: middleimportance: should knowfreq 48%

answer

  1. retrieval and generation are separate choices
  2. do not transcribe away what you retrieved visually
  3. a page image costs far more than its text
  4. context budget sets how many pages
  5. hybrid: text header plus page image

basics

~20 s

Pass the page images when the answer lives in a figure — extracted text drops exactly the content that made retrieval worth doing. The cost is tokens: a full page image runs on the order of a thousand-plus tokens, so the context budget caps how many pages you can send.

solid answer

~50 s

Retrieval and generation are separate decisions, and a common mistake is retrieving by page image and then handing the model a text transcription of that page. If a Kaplan-Meier curve with a hand-drawn annotation is the answer, the transcription contains a title and a footnote and the model will hedge or hallucinate. Pass the images. The constraint is arithmetic: a rendered page costs on the order of a thousand or more input tokens depending on resolution and provider, versus a couple of hundred for its text, so a five-page context that was trivially affordable in text becomes a real per-query cost in images. That is why recall@k matters — k is set by what you can afford. Practical shape: send the top few pages as images, include cheap text (title, page number, source) alongside each so the model can cite, and fall back to text-only for text-dense pages where nothing visual is at stake.

go deeper

for a junior

Recall that if the answer is in a chart, the model needs to see the chart, not a text transcription of the slide. Say that images cost far more input tokens than text.

for a middle

Be ready to explain how the token cost of page images converts the context budget into a hard cap on k, and why that makes recall at that k the retrieval metric worth optimising.

for a senior

Show the hybrid payload patterns — text header plus image, downsampling by rank, progressive fetch on demand — and how you would verify citations when the evidence is pixels rather than quotable text.

for a principal

Own the cost and reliability envelope: per-query cost at production traffic, resolution floors validated against the query mix, and whether an image-evidence system can meet the grounding-audit requirements the business has committed to.

## The mistake this question exists to catch A team builds a page-image retrieval system, is pleased that it finds the right slide, and then — because the rest of their stack is text — extracts the slide's text and puts *that* in the prompt. They have paid for visual retrieval and thrown the visual evidence away at the last step. The model receives "Overall Survival, Study 301, n=428" and is asked where the curves cross. It cannot know, so it either refuses or invents a month. Retrieval decides *which* pages the generator may see. Generation decides *what* it sees of them. Those are two independent choices and both must be made deliberately. ## When images are the right payload Pass the rendered page when the answer is spatial or graphical: reading a value off a plot, comparing bar heights, following an annotation, interpreting a diagram's arrows, reading a table whose meaning depends on its alignment, or handling scanned and handwritten material. In all of these, any linearisation into text either loses the information or fabricates a structure that was not there. ## When text is the right payload Pass text when the page is text-dense prose and the answer is a sentence on it. Text is roughly an order of magnitude cheaper per page, it is exactly quotable so citations are verifiable character-for-character, and it is trivially cacheable, diffable and loggable. Sending an image of a paragraph is paying image prices for something a text encoder handles better. The honest default is therefore *conditional*: route on page type, which you already know from the same slicing you used to evaluate retrieval. ## The arithmetic that sets k Image cost scales with the area the model must encode, so a full-page render is expensive relative to a text chunk — order-of-magnitude, a page image lands in the low thousands of input tokens at usable resolution while its extracted text is a few hundred. Exact numbers depend on provider and resolution, so measure with a token count on your real pages rather than assuming. The consequence is that context budget converts directly into k. If you can spend, say, 10,000 input tokens on evidence, that is a handful of page images or dozens of text chunks. This is why the retrieval metric to optimise is recall at *that* k: a system whose correct page ranks fourth is useless if you can only afford three images. It also means the cost conversation and the retrieval-quality conversation are the same conversation. Second-order effects matter too. Images generally cannot be reused across queries as cheaply as a shared text prefix, latency grows with pixels processed, and a long tail of pages in one request degrades the model's attention to any single one — sending eight mediocre pages is usually worse than sending three good ones. ## Hybrid payloads The strongest practical pattern sends both, asymmetrically. For each retrieved page, include a small text header — document title, page number, date, source URL — and then the page image. The header is nearly free, gives the model something exact to cite, and lets it distinguish pages that look alike. The image carries the evidence. A second hybrid is downsampling by rank: full resolution for the top page, lower for the next two. Quality degrades gracefully with rank while cost drops sharply, and the top page is usually the one the answer comes from. A third is progressive: send text first, and let the model request the page image for a specific page when the text is insufficient. This pays image cost only on the queries that need it, at the price of a second round trip and a more complex loop. ## Failure modes to watch **Silent truncation.** Push too many images and something gets dropped — often without an obvious error — and the model answers from what survived. Count tokens before sending, do not discover the limit by hitting it. **Unverifiable citations.** A model reading an image will paraphrase numbers it read off a chart. Those cannot be string-matched against a source, so grounding checks that work for text do not transfer. Ask for page-level attribution — "which page did this come from?" — and verify at page granularity rather than quote granularity. **Resolution below the evidence.** If you downsample to save tokens and the axis labels become unreadable, the model will still answer, and confidently. Small-text and fine-detail reading is a known weak spot, so validate that the resolution you chose actually preserves what your queries ask about. **Cost surprise.** Image-heavy contexts change the per-query cost profile enough that a system that was cheap in staging is expensive at production traffic. Model the cost per query at your real k before launch, not after.

  • How do you keep citations verifiable when the evidence is an image?
    Attribute at page granularity rather than quote granularity. Include a small text header with each image — document, page number, date — and require the model to name which page each claim came from. You cannot string-match a number the model read off a chart against a source text, so the check becomes: does that page, when a human or a second model looks at it, support the claim?
  • You can afford three page images per query but recall@3 is only 70%. What do you do?
    Two levers. Improve ranking so the answering page moves into the top three — a reranking stage, or the more expensive retrieval architecture on the figure-heavy slice. Or change the payload economics: downsample lower-ranked pages, send a text header plus a low-resolution image for ranks two and three, or go progressive and fetch a full-resolution page only when the model asks. I would measure both against end-to-end answer accuracy, not recall alone.
  • Is there a case for sending both the image and the full extracted text of the same page?
    Occasionally, for pages that are genuinely mixed — a dense table beside a chart — where the text gives exactly quotable values and the image gives the structure. But it roughly doubles the cost of that page for a modest gain, so I would treat it as a targeted exception justified by evaluation on that page type, not a default.

Retrieving the right page then sending only its text is like finding the right photograph in an archive and then reading the caption to someone over the phone.

saying these in an interview costs you the question

  • Retrieving by page image, then sending only extracted text
  • Assuming an image costs about the same as a text chunk
  • Downsampling until axis labels are unreadable and trusting the answer
  • Sending as many pages as fit rather than the few that are relevant
  • Expecting quote-level grounding checks to work on image evidence

context