How does multimodality work in Spring AI — how do you send an image (or audio) to a chat model, and what are the constraints?
answer
- Media = MimeType + data (Resource/byte[]/URL)
- .user(u -> u.text(...).media(IMAGE_PNG, res))
- only vision/audio-capable models (GPT-4o, Claude, Gemini)
- multiple media per message
- input multimodality != ImageModel generation
basics
~20 sMultimodality means the model accepts more than text — e.g. images or audio — as input. In Spring AI you attach a Media object (MimeType + data) to the user message: ChatClient.prompt().user(u -> u.text("...").media(MimeTypeUtils.IMAGE_PNG, resource)). Only vision/audio-capable models (e.g. GPT-4o, Claude) support it.
solid answer
~50 sA multimodal model can take non-text inputs alongside text. Spring AI models this with the Media abstraction: a Media carries a MimeType (IMAGE_PNG, IMAGE_JPEG, audio types, etc.) plus the payload as a Resource, byte[], or URL. You attach media to the user message via the ChatClient fluent API — .user(u -> u.text("Describe this image").media(MimeTypeUtils.IMAGE_PNG, imageResource)) — or by constructing a UserMessage with a List<Media>. Spring encodes it (typically base64 or a URL) in the provider's expected format. Support is per-model: only vision-capable models (OpenAI GPT-4o, Anthropic Claude, Google Gemini, some Ollama models like LLaVA) accept images, and audio input needs an audio-capable model/variant. Output multimodality (image/audio generation) uses separate models (ImageModel, speech/transcription models), not the chat Media input path. Constraints: model must support the modality and MIME type, size/token limits apply, and cost is higher for large media.
code
java · 9 linesvar image = new ClassPathResource("/receipt.png");
String result = ChatClient.create(chatModel) // must be a vision-capable model
.prompt()
.user(u -> u
.text("Extract the total amount and date from this receipt.")
.media(MimeTypeUtils.IMAGE_PNG, image))
.call()
.content();go deeper
Know you attach a Media (mime + data) to the user message and it needs a capable model.
Can wire .media(...) with the right MimeType and Resource, including multiple images.
Distinguishes input multimodality from generation models, and knows support/token/cost constraints.
Plans model selection and cost around media token usage, and separates chat-media pipelines from ImageModel/audio-transcription pipelines architecturally.
## What multimodality means **Multimodality** is a model's ability to accept (and sometimes produce) content beyond plain text — most commonly **images**, and increasingly **audio**, PDFs, or video. For chat, the relevant case here is passing image/audio *inputs* to a chat model so it can describe, analyze, or reason over them. ## The Media abstraction Spring AI represents a non-text attachment as `org.springframework.ai.content.Media` (formerly under the model message package). A `Media` has: - a **MimeType** — e.g. `MimeTypeUtils.IMAGE_PNG`, `IMAGE_JPEG`, or an audio type — telling the provider how to interpret the bytes; - the **data** — supplied as a Spring `Resource` (e.g. `ClassPathResource`, `UrlResource`), a `byte[]`, or a `URL`. Spring converts this into whatever wire format the provider expects (base64-inlined data or a URL reference). ## Attaching media via ChatClient The fluent API exposes `.media(...)` on the user message builder: ```java var image = new ClassPathResource("/photo.png"); String description = ChatClient.create(chatModel) .prompt() .user(u -> u.text("What is in this picture?") .media(MimeTypeUtils.IMAGE_PNG, image)) .call() .content(); ``` Equivalently, build a `UserMessage` with media directly: ```java var userMessage = UserMessage.builder() .text("Describe these images") .media(List.of(new Media(MimeTypeUtils.IMAGE_PNG, image))) .build(); chatModel.call(new Prompt(userMessage)); ``` You can attach **multiple** media items and mix text with several images in one message. ## Audio input The same `Media` mechanism carries audio for models that accept audio *input* (e.g. OpenAI's audio-capable chat variant): you attach the audio bytes with the appropriate MIME type. This is distinct from **speech-to-text transcription** and **text-to-speech**, which in Spring AI are separate dedicated models (`OpenAiAudioTranscriptionModel`, `OpenAiAudioSpeechModel`) — not the chat Media path. ## Model support is the hard constraint Multimodal input only works if the **underlying model supports it**: - Images: OpenAI GPT-4o / GPT-4-vision, Anthropic Claude 3+, Google Gemini, Ollama vision models (e.g. LLaVA). - Audio input: specific audio-capable variants. Sending an image to a text-only model fails or is ignored. Spring AI does not add vision to a model that lacks it — it just packages the input. ## Output vs input multimodality This leaf is about **input** multimodality (image/audio in). Generating images or audio is a different pipeline: `ImageModel` (e.g. `OpenAiImageModel`) for image generation and speech models for audio out — they are separate abstractions, not `ChatClient` Media. ## Gotchas - **Wrong or missing MimeType** — the provider may reject the payload; always set the correct one. - **Size/token limits** — images consume many tokens; large or many images raise cost and can exceed context limits. - **URL vs inline** — some providers fetch a URL; others require inlined base64. Spring handles encoding, but network-fetched URLs must be reachable by the provider. - **Model mismatch** — the single most common failure is pointing at a non-vision model. - **Don't confuse** chat image *input* with the separate `ImageModel` for image *generation*.
- What happens if you attach an image to a text-only chat model?It won't work — the provider either rejects the request or ignores the media, because Spring AI only packages the input; it can't grant vision to a model that doesn't support it. You must target a vision-capable model like GPT-4o or Claude 3+.
- How is passing an image to a chat model different from generating an image?Image input uses the ChatClient Media API on a vision-capable chat model to analyze/describe the picture. Image generation is a separate abstraction — ImageModel (e.g. OpenAiImageModel) — with its own prompt/response types; the two don't share the Media-input path.
saying these in an interview costs you the question
- Thinking Spring AI adds vision to any model regardless of provider support
- Confusing chat image input (Media) with ImageModel image generation
- Forgetting to set the correct MimeType
- Assuming all image bytes are free tokens-wise