Multimodal AI
You will learn how AI systems handle images, documents, speech, audio, and video alongside text — from how vision-language models fuse modalities to the practical APIs and retrieval patterns that power multimodal products. Interviewers probe this because real applications stopped being text-only, and they want engineers who can reason about the trade-offs of each modality, not just call an endpoint.
on this pageshowhide
explore
- Vision-Language Models5 questions
- Image Understanding in Practice5 questions
- Document AI5 questions
- Image Generation6 questions
- Speech-to-Text & TTS6 questions
- Video & Audio Understanding6 questions
- Multimodal Embeddings & RAG5 questions
questions
page 2 of 2What makes speaker diarization fail on a three-way consultation with an interpreter?
basics
~20 sDiarization answers "who spoke when" by clustering voice segments, so it breaks where turn boundaries blur: overlapping speech, very short back-channels, and a third speaker who alternates languages and register. Errors concentrate at turn edges, and one bad cluster split mislabels the whole recording.
When does uniform frame sampling fail, and how would you sample video adaptively instead?
basics
~20 sUniform sampling spends the same budget on empty footage as on the moment that matters, so short events fall between samples while idle aisles are described dozens of times. Adaptive sampling gates frames on change: dense while something moves, sparse while nothing does.
When would you choose a Q-Former connector over a simple MLP projection in a VLM?
basics
~20 sChoose query-based resampling when image tokens must be bounded — a Q-Former compresses any patch count into a fixed number of learned queries, capping context cost at the price of discarded detail. MLP projection keeps one token per patch and preserves detail.
Why did VLMs move from fixed 336x336 resizing to native-resolution tiling?
basics
~20 sSquashing a large image into one small square destroys anything smaller than a patch — small objects, fine text, chart labels. Tiling encodes the image as several full-resolution tiles plus a downscaled global view, preserving detail at the cost of far more tokens.
Caption-and-index, page-image embeddings, or late interaction — how do you pick for a figure-heavy slide corpus?
basics
~20 sThree designs trade recall for cost: captioning is cheapest but loses whatever the caption omits; one vector per page image is cheap and strong on most queries; late interaction wins on visually hard queries at a far larger index.
For a tool-heavy voice agent, would you use a cascaded STT-LLM-TTS stack or speech-to-speech?
basics
~20 sCascaded remains the production default for tool-heavy agents: you get a text transcript to log, evaluate and guard, and each stage is independently swappable. Speech-to-speech wins on latency and prosody because nothing is flattened to text, but it is harder to inspect and constrain.
How would you summarize a 3-hour video into chapters when it exceeds the context window?
basics
~20 sProcess it hierarchically: split into overlapping windows, summarize each with timestamps, then summarize the summaries into chapters. The reduce pass sees only the intermediate text, so detail and cross-window continuity have to be deliberately carried through it.
How do you build audio understanding when the signal is non-speech, like machine noise?
basics
~20 sTranscription pipelines produce nothing here, because there are no words. Either send the audio to a model that accepts raw sound and can describe or classify events, or train a dedicated classifier on spectrograms when the categories are narrow and the accuracy bar is high.
showing 31–38 of 38