skip to content

Multimodal AI

You will learn how AI systems handle images, documents, speech, audio, and video alongside text — from how vision-language models fuse modalities to the practical APIs and retrieval patterns that power multimodal products. Interviewers probe this because real applications stopped being text-only, and they want engineers who can reason about the trade-offs of each modality, not just call an endpoint.

on this pageshow

explore

questions

page 1 of 2

Why does copying text from a scanned PDF return nothing, but not from a born-digital PDF?

level: juniorimportance: must knowfreq 72%

answer

  1. one file, two very different payloads
  2. characters versus pixels
  3. test each page, not each file
  4. scanner and fax pages carry no characters
  5. OCR creates text that was never there

basics

~20 s

A born-digital PDF stores real text objects — character codes with fonts and positions — so extraction just reads them out. A scanned page stores only a photograph of the paper, so no characters exist until OCR creates them.

solid answer

~50 s

PDF is a container for page-painting instructions, not a text format. When a program produces the file — an ERP printing an invoice, a word processor exporting a report — the page carries a **text layer**: character codes, their fonts, and the coordinates where each run is drawn. Extracting from that is a parse: instant, free, and exact. When a page arrives through a scanner, a fax gateway or a phone camera, the PDF holds a single raster image; the characters exist only as pixels, so any text has to be *recognised* by OCR, which costs money, adds latency and introduces errors. Real corpora mix both inside one file — a born-digital cover page stapled to scanned annexes — so the working rule is to test each **page** for a usable text layer and send only the pixel pages down the recognition path.

code

python · 11 lines
python
import fitz  # PyMuPDF

doc = fitz.open("bill_of_lading.pdf")
for page in doc:
    text = page.get_text("text").strip()
    has_full_page_image = any(
        page.get_image_bbox(img).get_area() > 0.7 * page.rect.get_area()
        for img in page.get_images(full=True)
    )
    route = "parse" if len(text) > 50 and not has_full_page_image else "ocr"
    print(page.number, route)

go deeper

for a junior

Be able to say plainly that a scanned page is an image with no characters in it, and that OCR is what turns those pixels into text. Knowing to check for a text layer before reaching for OCR is the expected answer.

for a middle

Explain the mechanics: a text layer holds character codes with fonts and coordinates, so extraction is a parse, while a scan needs recognition that costs money and has an error rate. Describe a concrete per-page detection heuristic.

for a senior

Show that you route per page, not per file, and that you instrument the split because it drives cost. Mention the searchable-PDF trap — an embedded text layer that is really untrusted OCR output — and how you would validate it against a labelled sample.

for a principal

Own the economics: the parsed-versus-recognised ratio is a budget line that moves when customers change how they submit documents. Be ready to argue for pushing born-digital submission upstream, and to set the policy that no page is recognised twice.

## What a PDF actually stores A PDF describes how to paint a page. For a file produced by software, that description includes text-showing operations: character codes, the font each code belongs to, and the coordinates where each run of text is placed. Collectively this is the *text layer*. Pulling text out of such a page is a parse of data that is already present — it is exact, costs nothing but CPU, and comes with per-run positions you can reuse for layout work. A page that entered the file through a scanner, a fax gateway, or a phone camera contains none of that. The whole page is one large raster image — a grid of pixels — wrapped in a PDF container. There are no character codes anywhere in the file. Selecting text in a viewer selects nothing because there is nothing to select. ## Two paths with very different economics The text-layer path is a parse: microseconds per page, no per-page fee, no recognition errors, and deterministic across reruns. The raster path is recognition: you render the page to an image (200–300 DPI is the usual working range), then run an OCR engine or a document vision-language model over it. That path costs real money per page, adds latency, and has a non-zero character error rate — the standard metric is CER, character error rate, and it climbs sharply with skew, low contrast, fax compression artefacts, and handwriting. The practical consequence: in a freight forwarder's pipeline that ingests bills of lading and arrival notices, the born-digital half is essentially free and perfectly accurate, while the faxed half carries the entire extraction budget and nearly all of the review workload. Knowing which half a page belongs to *before* you spend anything is the first design decision in any document pipeline. ## Both kinds live in the same corpus, and often the same file Documents get merged. A clerk prints a system-generated cover sheet, staples the carrier's faxed annex behind it, and scans the staple line. The result is one PDF where page 1 has a perfect text layer and pages 2–7 are images. Deciding per file therefore produces one of two failures: you OCR pages that did not need it (paying and adding errors), or you skip pages that had no text at all (silently dropping half the shipment's line items). Detection belongs at page granularity. ## How to detect a usable text layer The cheap heuristic is to extract text per page and look at how much came back relative to how much ink is on the page: a page with a handful of characters but a full-page image object is a scan. Refinements worth having are checking whether a single image covers most of the page area, and guarding against legitimately sparse pages (a separator sheet, a signature page) so they are not misrouted as failures. One important special case: **searchable PDFs**. When someone has already run OCR, the tool typically writes the recognised text back into the file as an invisible text layer positioned under the image. Those pages *do* extract text — but that text is somebody else's OCR output, inherited errors and all. Treat it as untrusted: sample a labelled subset and measure its accuracy before deciding whether to trust it or re-recognise the pixels. ## Traps inside a genuine text layer A text layer being present does not make it clean. Fonts embedded without a Unicode mapping extract as garbage or private-use codepoints even though the page renders correctly on screen. Kerning and justification split a single word into several runs. Ligatures and soft hyphens survive into your output. And crucially, the order of drawing operations is whatever the producing software emitted — it is not a reading order, which is why multi-column pages can extract as interleaved nonsense even when every character is exact. ## What this means for the pipeline Route per page into two lanes. The text-layer lane is cheap and exact but still needs layout work to recover structure. The raster lane renders and recognises. A hybrid is legitimate and common: you sometimes rasterise a *born-digital* page on purpose, because a layout-aware model reading the rendered image reconstructs a dense table better than the raw run order does — you are trading a little cost for structure you could not otherwise get. Whatever you choose, instrument the split. The ratio of parsed pages to recognised pages is your cost model, and it drifts as customers change how they send documents.

  • A born-digital page extracts every character correctly, yet the words come out jumbled. What is happening?
    The text layer records drawing operations in whatever order the producing software emitted them, not in reading order. On a multi-column or boxed layout, that order interleaves regions. The characters are exact; the sequence is meaningless. Recovering order requires layout analysis over the run positions rather than trusting the extraction sequence.
  • How would you decide a page needs OCR without misclassifying a legitimately near-empty page?
    Combine signals instead of thresholding character count alone: how many characters came back, whether a single image object covers most of the page area, and how much of the page is non-white ink. A separator sheet has little text and little ink; a scan has little text and a full-page image. That pair separates them.
  • A vendor sends 'searchable PDFs' that already contain a text layer. Do you trust it?
    Not by default. That layer is usually a previous OCR pass, so it carries someone else's recognition errors and may be misaligned with the pixels. Sample a labelled subset, measure character and field-level accuracy, and only then decide between trusting it, re-recognising the image, or using it as a second opinion.

A born-digital PDF is a typed document; a scanned PDF is a photograph of one. You can search the first immediately; the second needs someone to read it back into words first.

saying these in an interview costs you the question

  • Assuming every PDF contains extractable text
  • Running OCR on born-digital pages that already parse exactly
  • Believing OCR improves text that was already exact
  • Deciding text-layer vs scan per file instead of per page
  • Treating the extraction sequence as the reading order
  • Trusting a searchable PDF's embedded OCR without measuring it

context

open as a page

How does a text-to-image diffusion model turn random noise into an image?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Generation starts from pure random noise. The model repeatedly predicts how much noise is present and removes a little of it, conditioned on the text prompt at every step. After a fixed number of steps the noise has been shaped into an image.

open as a page

In a vision-language model, why can one image cost more tokens than a page of text?

level: juniorimportance: must knowfreq 72%

basics

~20 s

The vision encoder cuts the image into a grid of fixed-size patches and emits roughly one token per patch, so token count grows with image area. A 1024x1024 image at patch size 16 is 64x64 = 4096 patches.

open as a page

In document extraction, why demand a bounding box and page number per extracted field?

level: middleimportance: must knowfreq 56%

basics

~20 s

Grounding turns an unverifiable string into an auditable claim. A clerk jumps straight to the pixels behind "container MSKU4412345", and a box landing on blank space exposes an invented value before it reaches the downstream system.

open as a page

In document extraction, when do you pick classic OCR over a document VLM or a frontier model?

level: middleimportance: must knowfreq 65%

basics

~20 s

Classic OCR wins on clean, fixed layouts at high volume: cheap, fast, deterministic, and it returns per-word boxes and confidence scores. Purpose-built document VLMs handle messy layout and tables; frontier models are for open-ended reasoning over the page.

open as a page

How does ColPali-style late interaction score a page image against a text query?

level: middleimportance: must knowfreq 55%

basics

~20 s

The page image is stored as many patch vectors instead of one. Each query token vector is matched against every patch, the best match per token is kept, and those maxima are summed — the MaxSim score. No OCR is involved.

open as a page

In diffusion image generation, what does raising the guidance scale trade away?

level: middleimportance: must knowfreq 62%

basics

~20 s

Guidance scale trades diversity and naturalness for prompt adherence. Low values give loose, varied, plausible images that may ignore parts of the brief; high values track the prompt harder but flatten variety and, past a point, produce over-saturated, contorted output.

open as a page

How do image detail levels change what a vision model sees and what each image costs?

level: middleimportance: must knowfreq 57%

basics

~20 s

A low detail setting collapses the image to a small fixed token budget — cheap and fast, but fine text and precise layout are gone. Higher settings preserve resolution and cost several times more, because image tokens scale with the pixels actually processed.

open as a page

When sending images to a vision model, how do you choose between base64, a URL, and an uploaded file?

level: middleimportance: must knowfreq 62%

basics

~20 s

Inline base64 suits one-off images but inflates the request by roughly a third. A URL keeps requests small when the image is already hosted and reachable by the provider. An uploaded file reference wins when the same image is sent many times across requests.

open as a page

Why choose streaming speech-to-text over batch transcription for a live dictation product?

level: middleimportance: must knowfreq 70%

basics

~20 s

Streaming returns interim text within a few hundred milliseconds so the speaker can see and correct errors as they talk. Batch waits for the whole recording but decodes with full right-hand context, so its final transcript is usually more accurate.

open as a page

How does frame-sampling rate drive the token cost of sending video to a multimodal model?

level: middleimportance: must knowfreq 72%

basics

~20 s

Each sampled frame becomes a few hundred image tokens, so cost scales with duration times frames per second. At roughly 258 tokens a frame, one frame per second across an 8-hour shift is about 7.4 million tokens, far past any context window. Rate and per-frame resolution are the levers.

open as a page

In a vision-language model, what do the encoder, connector and decoder each do?

level: middleimportance: must knowfreq 68%

basics

~20 s

Three parts: a vision encoder (a Vision Transformer, usually SigLIP-class) turns the image into patch embeddings, a connector maps those embeddings into the decoder's embedding space, and the language decoder consumes them alongside ordinary text tokens.

open as a page

In image generation, when do you use mask inpainting instead of conversational editing?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Use a mask when a region of an approved photograph must be guaranteed byte-identical outside the edit — legal, brand or evidentiary constraints. Use conversational multi-reference editing when the change is global or identity must carry across several outputs, accepting that the whole frame is regenerated.

open as a page

What kinds of image questions do vision models still get wrong, and how do you design around it?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Counting repeated objects, reading dense small text, precise spatial relations, and pulling exact values off unlabelled charts remain unreliable — and the model states wrong answers fluently rather than abstaining. Design around it by cropping, decomposing into enumerable steps, and requiring ranges instead of false precision.

open as a page

How should a voice agent handle barge-in when a customer talks over its spoken reply?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Stop playback within a couple of hundred milliseconds, keep the microphone open the whole time so nothing is missed, and truncate the conversation history to the words the customer actually heard — not the full reply the model generated.

open as a page

Why can a 6% word error rate still be unusable for a clinical dictation product?

level: seniorimportance: must knowfreq 55%

basics

~20 s

A headline word error rate averages over all words, so it is dominated by common speech. Drug names, dosages and negations are rare in the corpus but carry nearly all the risk, and errors there can run several times the headline rate while barely moving it.

open as a page

Why do video models get event ordering and duration wrong, and how do you mitigate it?

level: seniorimportance: must knowfreq 54%

basics

~20 s

Sampled stills carry appearance, not continuity. The model sees a pallet on the floor and a forklift reversing but not the motion between them, so it infers a plausible order rather than observing one, and it has no clock unless the frames are timestamped.

open as a page

How do you ask a vision model to compare two screenshots in one request?

level: juniorimportance: should knowfreq 44%

basics

~20 s

Put both images in the same message, each preceded by a short text label such as "Before:" and "After:", then ask the comparison question. Images are positional, so labelling them in text is what lets the model refer to each one reliably.

open as a page

How do modern text-to-speech systems let you control a voice's tone and pacing?

level: juniorimportance: should knowfreq 35%

basics

~20 s

Mainly through plain-language style instructions — "warm, unhurried, slight smile" — plus inline tags in the script and phonetic spellings for tricky words. Newer voice models have moved away from the heavy XML-style markup that older engines required.

open as a page

What do you gain by using a native video input API instead of extracting frames yourself?

level: juniorimportance: should knowfreq 56%

basics

~20 s

A native video API takes the file plus a frame-rate hint and handles decoding, sampling, timestamps and the audio track for you, so the model can answer with times. Hand-extracted frames sent as an image list give you full control but arrive as anonymous stills.

open as a page

Should retrieved slide images or their extracted text go into the model's context, and what does that cost?

level: middleimportance: should knowfreq 48%

basics

~20 s

Pass the page images when the answer lives in a figure — extracted text drops exactly the content that made retrieval worth doing. The cost is tokens: a full page image runs on the order of a thousand-plus tokens, so the context budget caps how many pages you can send.

open as a page

In a CLIP-style joint image-text space, what is the modality gap and how does it bite retrieval?

level: middleimportance: should knowfreq 38%

basics

~20 s

Image and text embeddings occupy separate regions of the shared space rather than mixing. Image-image similarities therefore run systematically higher than image-text ones, so a single threshold or a mixed candidate pool quietly favours whichever modality matches the query's own type.

open as a page

Why do latent diffusion models denoise in a compressed latent space instead of pixels?

level: middleimportance: should knowfreq 48%

basics

~20 s

Denoising every pixel for dozens of steps is far too expensive. Latent diffusion encodes the image to a much smaller tensor with an autoencoder, runs the whole denoising loop there, and decodes once at the end — cutting compute and memory enough to run on a single consumer GPU.

open as a page

How does SigLIP's sigmoid loss differ from CLIP's softmax contrastive objective?

level: middleimportance: should knowfreq 42%

basics

~20 s

CLIP normalizes similarity over every other pair in the batch with a softmax, so the loss depends on batch composition and needs large synchronized batches. SigLIP scores each image-text pair independently with a pairwise sigmoid loss, removing that global dependence.

open as a page

A two-column customs form extracts as interleaved gibberish — how do you fix the reading order?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Extraction emitted characters in drawing or naive scan order, not reading order. Fix it upstream with layout analysis: detect columns and blocks first, order text within each region, then read regions in sequence — never by post-processing a flat text dump.

open as a page

How do you set the confidence threshold that sends an extracted field to human review?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Tune it against a labelled sample rather than intuition: sweep the cutoff, plot residual error of auto-posted fields against review volume, and set it per field at the point where the remaining error meets that field's business tolerance.

open as a page

How would you prove page-image retrieval beats an OCR-then-embed baseline on figure-heavy decks?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Run both pipelines over the same pages with the same queries and compare ranking metrics such as nDCG@5 and recall@k. Use a ViDoRe-style visual-document benchmark for a sanity check, but decide on a labelled set built from your own corpus, sliced by page type.

open as a page

In an open-weights image pipeline, when do you use ControlNet versus a LoRA?

level: seniorimportance: should knowfreq 40%

basics

~20 s

ControlNet imposes structure on one generation from a control image — depth, edges, pose, segmentation — and needs no training. A LoRA is a small trained adapter that carries a consistent style or subject across every generation. Structure per image, look across images; often both together.

open as a page

Why did image models move from U-Net DDPM to DiT backbones with flow matching?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Transformer backbones scale with data and compute the way language models do and handle text-image attention better, while flow matching learns a near-straight path from noise to data instead of a long stochastic denoising chain — so comparable quality arrives in far fewer sampling steps.

open as a page

What is thinking with images, and when does it beat pre-cropping a screenshot yourself?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Thinking with images means the model crops, zooms and re-inspects regions of an image during its own reasoning instead of answering from one downscaled glance. It beats manual pre-cropping when you do not know in advance where the answer lives.

open as a page

showing 1–30 of 38