In multimodal moderation, how do you catch harm that lives only in the image?
answer
- clean words, unsafe picture
- OCR helps, but is not enough
- separate checks miss the pairing
- one call sees both modalities
- sample frames, never trust the thumbnail
basics
~20 sScore the image, not just the words around it. Text-only classifiers pass a clean caption over an unsafe picture, so the pipeline needs image-capable moderation — either a joint text-plus-image call or a dedicated image classifier — plus sampled frames for video rather than one thumbnail.
solid answer
~50 sA text-only moderation pass on a post with a harmless caption and an unsafe image returns clean, because it never saw the image — the harm is invisible to it by construction. The fix is an image-capable check in the pipeline. Two shapes exist: a **joint** call that scores text and image together, as `omni-moderation-latest` and Llama Guard 4 do, or a **dedicated image classifier** such as ShieldGemma 2 run alongside the text check. Joint scoring is what catches compositional harm, where each modality is innocuous but the pairing is not — a benign photo plus a caption that recontextualises it as a threat. Independent unimodal checks structurally cannot see that, because neither one has both halves. Video needs sampled frames, not just the uploaded thumbnail, since the thumbnail is chosen by the uploader. Screen before the item enters distribution, not after it has been recommended.
go deeper
Be able to say that a text moderation check never sees the picture, so any platform carrying images needs an image-capable classifier in the pipeline as well.
Explain the three shapes of image-carrying harm — harm in the image, text rendered inside the image, and harm that exists only in the caption-image pairing — and why the third one forces joint scoring.
Show judgment about cost and placement in a real pipeline: cheap unimodal passes first, joint scoring on the ambiguous slice, sampled frames plus transcribed audio for video, and screening before the item enters distribution.
Own the budget and risk framing — per-item multimodal scoring at platform scale is a real cost line, so decide which surfaces and creator tiers justify dense sampling and what residual leakage rate the business accepts.
## Why the modality gap is a real production hole A moderation pipeline built for text quietly assumes that everything worth judging is in the words. On a platform that carries images or video that assumption fails on the first upload. A clean caption over an unsafe photograph scores clean — not because the classifier is wrong but because the harmful bytes were never passed to it. This is one of the easiest moderation layers to misplace, and it is usually misplaced by omission rather than by a bad decision. ## Three shapes of image-carrying harm **Harm entirely in the image.** The picture is sexual, graphic, or otherwise violating; the caption is a weather remark. Any image-capable classifier catches this. A text pipeline catches none of it. **Harm in text rendered inside the image.** Slurs, harassment, or prohibited claims typed onto a picture. This defeats a text pipeline that reads only the caption field. Running OCR and feeding the extracted string into the text classifier helps, and is worth doing, but it is a supplement rather than a substitute: OCR misses stylised type, and it tells you nothing about the imagery the words sit on. **Compositional harm.** Each modality is fine and the pairing is not — an ordinary photograph of a person alongside a caption that recontextualises it as a threat, an accusation, or sexualised content. This is the case that decides your architecture, because two independent unimodal classifiers structurally cannot detect it. Neither one holds both halves, so neither one can see the relationship. Only a check that receives the text and the image together has the information required. ## Joint scoring versus parallel unimodal checks The available tooling reflects both patterns. Hosted multimodal moderation endpoints accept text and image inputs in one call and return category scores for the item as a whole; open models such as Llama Guard 4 take image plus text against a hazard taxonomy; and image-only classifiers such as ShieldGemma 2 are built to screen pictures for sexually explicit, violent, and dangerous content. Cloud content-safety services expose separate text-analysis and image-analysis operations that you compose yourself. A reasonable production layout is layered. Run the cheap unimodal checks first, because they dispose of the obvious cases: an image classifier that returns high severity means you are done, no joint reasoning required. Send to the joint or reasoning check the items where the unimodal signals are individually unremarkable but the pairing warrants judgment — which is also the slice where a human reviewer would need to see both halves anyway. ## Video is not one image Video tempts an obvious shortcut: screen the thumbnail. The thumbnail is the one frame the uploader chose, which makes it the least representative frame in the file. Real pipelines sample frames — at a fixed interval, at scene changes, or both — and score the sample, accepting that a rare offending frame between samples can slip through. They also handle the audio track, usually by transcribing it and running the transcript through the text classifier, while remembering that a transcript preserves the words and drops the tone, so a threatening delivery of neutral words reads as clean. Cost drives most of these choices. Scoring an image costs more than scoring a sentence, and scoring thirty sampled frames per video costs thirty times that. On a platform taking millions of uploads a day this is a budget item, not a footnote, and it pushes teams toward cheap classifiers on the full stream with expensive multimodal reasoning reserved for a narrow slice. ## Placement matters as much as capability Getting the modality right and the position wrong still leaks. Image screening belongs before the item enters distribution — before it is indexed, recommended, or served to anyone but its author. A pipeline that publishes on upload and moderates asynchronously has already exposed the content to whatever audience the first minutes of distribution reach, and a takedown afterwards does not undo that. When the moderation call is slow enough that inline screening hurts the upload experience, the usual compromise is to admit the item to a restricted state — visible to its creator, excluded from recommendation and search — until the check clears. ## What to say in an interview The strong answer names the gap concretely, distinguishes the three shapes of harm above, explains why compositional harm forces joint scoring rather than two parallel checks, and treats video and audio as separate sampling problems rather than as images with extra steps. The weak answer is "run an image classifier too" with no account of composition, sampling, or where in the request path the check sits.
- Why is running a text classifier and an image classifier separately insufficient?Because neither one holds both halves of the item. Compositional harm — an innocuous photograph paired with a caption that turns it into a threat or an accusation — exists only in the relationship between the modalities. Two independent classifiers each see a clean input and each return clean, so the composition is invisible by construction. Joint scoring is the only check with the information required.
- How would you moderate a two-minute video without paying to score every frame?Sample rather than exhaustively score: frames at a fixed interval plus frames at scene changes, and transcribe the audio track into the text classifier. Never rely on the uploader-chosen thumbnail. Accept explicitly that a rare offending frame between samples can pass, and tune sampling density by the creator's risk tier rather than uniformly.
- Does running OCR over an image and checking the extracted text replace image classification?No. OCR catches slurs and prohibited claims typed onto a picture, which a caption-only pipeline misses, so it is worth adding. But it says nothing about the imagery itself, and it degrades on stylised or low-contrast type. Treat it as a supplementary text signal layered on top of an image classifier, never as a substitute for one.
saying these in an interview costs you the question
- Assumes a clean caption means the whole post is safe
- Runs text and image classifiers and never combines the verdicts
- Screens only the uploader-chosen thumbnail of a video
- Thinks OCR of embedded text replaces image classification
- Publishes on upload and moderates images asynchronously afterwards