skip to content

Multimodal AI

You will learn how AI systems handle images, documents, speech, audio, and video alongside text — from how vision-language models fuse modalities to the practical APIs and retrieval patterns that power multimodal products. Interviewers probe this because real applications stopped being text-only, and they want engineers who can reason about the trade-offs of each modality, not just call an endpoint.

on this pageshow

explore

questions

page 2 of 2

What makes speaker diarization fail on a three-way consultation with an interpreter?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Diarization answers "who spoke when" by clustering voice segments, so it breaks where turn boundaries blur: overlapping speech, very short back-channels, and a third speaker who alternates languages and register. Errors concentrate at turn edges, and one bad cluster split mislabels the whole recording.

open as a page

When does uniform frame sampling fail, and how would you sample video adaptively instead?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Uniform sampling spends the same budget on empty footage as on the moment that matters, so short events fall between samples while idle aisles are described dozens of times. Adaptive sampling gates frames on change: dense while something moves, sparse while nothing does.

open as a page

When would you choose a Q-Former connector over a simple MLP projection in a VLM?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Choose query-based resampling when image tokens must be bounded — a Q-Former compresses any patch count into a fixed number of learned queries, capping context cost at the price of discarded detail. MLP projection keeps one token per patch and preserves detail.

open as a page

Why did VLMs move from fixed 336x336 resizing to native-resolution tiling?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Squashing a large image into one small square destroys anything smaller than a patch — small objects, fine text, chart labels. Tiling encodes the image as several full-resolution tiles plus a downscaled global view, preserving detail at the cost of far more tokens.

open as a page

Caption-and-index, page-image embeddings, or late interaction — how do you pick for a figure-heavy slide corpus?

level: principalimportance: should knowfreq 42%

basics

~20 s

Three designs trade recall for cost: captioning is cheapest but loses whatever the caption omits; one vector per page image is cheap and strong on most queries; late interaction wins on visually hard queries at a far larger index.

open as a page

For a tool-heavy voice agent, would you use a cascaded STT-LLM-TTS stack or speech-to-speech?

level: principalimportance: should knowfreq 48%

basics

~20 s

Cascaded remains the production default for tool-heavy agents: you get a text transcript to log, evaluate and guard, and each stage is independently swappable. Speech-to-speech wins on latency and prosody because nothing is flattened to text, but it is harder to inspect and constrain.

open as a page

How would you summarize a 3-hour video into chapters when it exceeds the context window?

level: principalimportance: should knowfreq 40%

basics

~20 s

Process it hierarchically: split into overlapping windows, summarize each with timestamps, then summarize the summaries into chapters. The reduce pass sees only the intermediate text, so detail and cross-window continuity have to be deliberately carried through it.

open as a page

How do you build audio understanding when the signal is non-speech, like machine noise?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

Transcription pipelines produce nothing here, because there are no words. Either send the audio to a model that accepts raw sound and can describe or classify events, or train a dedicated classifier on spectrograms when the categories are narrow and the accuracy bar is high.

open as a page

showing 31–38 of 38