skip to content

Spring AI

Spring AI: the chat client over portable model abstractions, prompts and structured output, embeddings and vector stores, RAG, tool calling, chat memory and MCP. Interviewers ask about it because it is now a routine feature request rather than a research project.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

What is Spring AI's ChatClient, and how do you make a basic call to a model with it?

level: juniorimportance: must knowfreq 70%

answer

  1. prompt().user().call().content()
  2. fluent client over ChatModel
  3. create() vs builder()
  4. content / chatResponse / entity
  5. call() blocks

basics

~10 s

ChatClient is Spring AI's fluent client for talking to an LLM. You write chatClient.prompt().user("...").call().content() to send a user message and get the reply back as a String.

solid answer

~40 s

ChatClient is the high-level, fluent API in Spring AI for calling a chat model (an LLM). You build it once from an auto-configured ChatModel bean via ChatClient.create(chatModel) or ChatClient.builder(chatModel).build(). To send a request you chain: prompt() starts a request spec, user("...") sets the user message (optionally system("...") for a system prompt), then call() executes synchronously. From the call spec you extract the result: content() gives the reply text as a String, chatResponse() gives the full ChatResponse (tokens, metadata, finish reason), and entity(MyType.class) maps the reply into a typed object (structured output). It mirrors the ergonomics of WebClient/RestClient, so the fluent chain reads left-to-right from building the prompt to extracting the answer.

code

java · 18 lines
java
@Service
class AssistantService {
    private final ChatClient chatClient;

    AssistantService(ChatModel chatModel) {           // auto-configured bean
        this.chatClient = ChatClient.builder(chatModel)
            .defaultSystem("You are a helpful travel guide.")
            .build();
    }

    String tips(String city) {
        return chatClient.prompt()
            .user(u -> u.text("Name three things to do in {city}.")
                        .param("city", city))
            .call()
            .content();                               // String reply
    }
}

go deeper

for a junior

Know the happy-path chain prompt().user().call().content() and that it returns the reply text.

for a middle

Distinguish content() vs chatResponse() vs entity(); know it wraps a ChatModel and is built via builder() with defaults.

for a senior

Explain thread-safety/reuse, structured output via entity(), and that the fluent spec is the per-request state.

for a principal

Frame ChatClient as the app-facing seam that keeps business code provider-agnostic and testable (mock the ChatModel or client).

**Spring AI** is Spring's integration library for building AI/LLM features. An **LLM (large language model)** is a text-generation model like OpenAI GPT, Anthropic Claude, or a local Ollama model. Spring AI gives you two layers to call one: - **ChatModel** — the low-level portable interface (one method conceptually: take a Prompt, return a ChatResponse). Spring Boot auto-configures a ChatModel bean for whichever provider starter is on the classpath. - **ChatClient** — a **fluent (builder-style) client** layered on top of ChatModel that removes boilerplate: assembling messages, setting options, attaching advisors, and extracting the result. **Creating a ChatClient.** Inject the auto-configured ChatModel and build once (typically in a @Configuration or @Service): ```java ChatClient chatClient = ChatClient.create(chatModel); // quick ChatClient chatClient = ChatClient.builder(chatModel) // configurable .defaultSystem("You are a terse assistant.") .build(); ``` ChatClient is thread-safe and meant to be reused as a bean. **Making a call — the fluent chain:** 1. `prompt()` — begins a *request spec*. Overloads: `prompt(String)` sets the user text directly, or `prompt(Prompt)` passes a pre-built Prompt object. 2. `.system("...")` / `.user("...")` — set the **system message** (instructions/persona) and **user message** (the actual question). You can also pass a lambda to add parameters for template substitution. 3. `.call()` — executes the request **synchronously (blocking)** and returns a *call response spec*. 4. Extract the result: - `.content()` → `String`, just the reply text. - `.chatResponse()` → `ChatResponse`, the full envelope: generations, `ChatResponseMetadata` (token usage, model name, finish reason). - `.entity(Class<T>)` / `.entity(ParameterizedTypeReference<T>)` → maps the reply to a typed object using structured-output converters (great for JSON-shaped answers). **Full example:** ```java String answer = chatClient.prompt() .system("You are a helpful travel guide.") .user("Name three things to do in Lisbon.") .call() .content(); ``` **Gotchas / when to use:** - Use **ChatClient** for almost all app code — it is the recommended entry point. Drop to **ChatModel** only for low-level needs. - `call()` **blocks** the calling thread until the whole reply is generated; for token-by-token output use `stream()` instead. - Build the client **once** and reuse it; don't rebuild per request. - `content()` can be null if the model returned no text (e.g. a pure tool-call response) — prefer `chatResponse()` when you need to inspect that. - The same code works across providers; only the starter dependency and config change.

  • How do you get token usage or the finish reason instead of just the text?
    Call .chatResponse() instead of .content(); ChatResponse.getMetadata() exposes Usage (prompt/completion/total tokens) and each Generation carries metadata like the finish reason.
  • Should you create a ChatClient per request?
    No. It is thread-safe and expensive-ish to build; create it once (e.g. in the constructor or a @Bean) and reuse it. The per-request state lives in the prompt() spec, not the client.

saying these in an interview costs you the question

  • Thinking ChatClient is provider-specific (e.g. an 'OpenAI client')
  • Believing call() streams tokens incrementally
  • Rebuilding the client on every request
  • Confusing content() (String) with chatResponse() (full envelope)

context

open as a page

What is Spring AI's ChatMemory abstraction and why is it needed when talking to an LLM?

level: juniorimportance: must knowfreq 70%

basics

~20 s

LLM calls are stateless — the model forgets previous turns. ChatMemory is a Spring AI interface that stores a conversation's past messages and replays them into each new request, so the model appears to remember the dialogue.

open as a page

What is an EmbeddingModel in Spring AI and what does it produce?

level: juniorimportance: must knowfreq 60%

basics

~10 s

EmbeddingModel is a Spring AI interface that turns text into an embedding — a list of floating-point numbers (a vector) representing the text's meaning. Similar texts get numerically similar vectors, which powers semantic search.

open as a page

In Spring AI, what is a Prompt, and what roles do SystemMessage and UserMessage play inside it?

level: juniorimportance: must knowfreq 70%

basics

~10 s

A Prompt is the request sent to the model: a list of Messages plus optional options. A SystemMessage sets the model's behavior/persona; a UserMessage carries the user's actual input.

open as a page

What is RAG in Spring AI, and what does the QuestionAnswerAdvisor do?

level: juniorimportance: must knowfreq 70%

basics

~10 s

RAG (Retrieval-Augmented Generation) fetches relevant documents from a vector store and adds them to the prompt so the LLM answers from your data. Spring AI's QuestionAnswerAdvisor does this automatically on each ChatClient call.

open as a page

What is tool (function) calling in Spring AI, and why would you use it?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Tool calling lets the LLM ask your app to run a Java method (e.g. fetch weather, query a DB) and use the result in its answer. In Spring AI you annotate a method with @Tool and register it with the ChatClient.

open as a page

What is the difference between ChatModel and ChatClient in Spring AI, and how does ChatModel provide provider portability?

level: middleimportance: must knowfreq 65%

basics

~20 s

ChatModel is the low-level portable interface every provider (OpenAI, Anthropic, Ollama, Azure) implements. ChatClient is the fluent, higher-level API built on top of a ChatModel to make calls ergonomic. Your code targets both, so swapping providers just means swapping the starter dependency.

open as a page

How does conversationId work in Spring AI ChatMemory, and how do you keep two users' conversations separate?

level: middleimportance: must knowfreq 65%

basics

~20 s

conversationId is the key that partitions stored messages. Each user/session gets its own id; you pass it per request via the advisor param ChatMemory.CONVERSATION_ID. If you omit it, everything shares the default id "default" and users see each other's history.

open as a page

What is the VectorStore abstraction and the Document model, and how do you add data to a store?

level: middleimportance: must knowfreq 62%

basics

~10 s

VectorStore is Spring AI's interface for saving and searching embeddings. You wrap text plus metadata in Document objects and call vectorStore.add(documents); the store embeds and persists them. Later similaritySearch finds the closest documents.

open as a page

Describe Spring AI's ETL pipeline for ingesting documents into a vector store (Reader, Transformer, Writer).

level: middleimportance: must knowfreq 65%

basics

~10 s

You read source files into Documents with a DocumentReader, split/enrich them with a DocumentTransformer (e.g., TokenTextSplitter), then write them into the VectorStore with a DocumentWriter, which embeds and stores them for later retrieval.

open as a page

How do you define a tool with @Tool / @ToolParam and register it with ChatClient, and what does the execution loop look like?

level: middleimportance: must knowfreq 50%

basics

~20 s

Put @Tool(description=...) on a method and @ToolParam on its arguments. Register the containing object with ChatClient via .tools(new MyTools()) per request or .defaultTools(...) on the builder. Spring advertises the schema, runs the method when the model asks, and loops until a final answer.

open as a page

How does similaritySearch with SearchRequest work, and how do you apply metadata filters?

level: seniorimportance: must knowfreq 55%

basics

~20 s

You build a SearchRequest with the query text, a topK (how many results), a similarityThreshold, and a filter expression on metadata. The store embeds the query, finds the nearest vectors that also match the filter, and returns them as Documents with scores.

open as a page

How does BeanOutputConverter turn a free-text LLM reply into a typed Java POJO?

level: seniorimportance: must knowfreq 68%

basics

~10 s

BeanOutputConverter<T> generates a JSON-schema format instruction from the target type via getFormat(); you append that to the prompt so the model replies in JSON, then convert(reply) deserializes that JSON into your POJO with Jackson.

open as a page

What is the Model Context Protocol (MCP), and what do the spring-ai-starter-mcp-client and spring-ai-starter-mcp-server starters give you?

level: juniorimportance: should knowfreq 40%

basics

~20 s

MCP is an open protocol that standardizes how AI apps connect to external tools and data over a client-server link. The client starter lets your app call remote MCP servers; the server starter lets your app publish tools other MCP clients can use.

open as a page

What is the difference between call() and stream() on ChatClient, and when would you use each?

level: middleimportance: should knowfreq 55%

basics

~20 s

call() runs synchronously and blocks until the whole reply is ready, giving you a String or ChatResponse. stream() is reactive: it returns a Flux that emits the reply piece by piece as the model generates it, so you can show tokens live.

open as a page

On the MCP server side, how do you expose @Tool-annotated beans as MCP tools in a Spring AI application?

level: middleimportance: should knowfreq 38%

basics

~20 s

Add spring-ai-starter-mcp-server, annotate your methods with @Tool (params with @ToolParam), and register a ToolCallbackProvider bean built via MethodToolCallbackProvider from those objects. The MCP server auto-config discovers the provider and publishes each @Tool method as an MCP tool.

open as a page

When would you use ListOutputConverter versus BeanOutputConverter, and what are ListOutputConverter's limits?

level: middleimportance: should knowfreq 45%

basics

~20 s

Use ListOutputConverter for a simple List<String> — it tells the model to reply as a comma-separated list and splits it. Use BeanOutputConverter for structured objects (POJOs) or lists of objects, which come back as JSON.

open as a page

How does PromptTemplate (and SystemPromptTemplate) render variables into messages in Spring AI?

level: middleimportance: should knowfreq 60%

basics

~10 s

PromptTemplate holds a template string with {placeholder} slots. You call create(Map) with variable values; it substitutes them and returns a Prompt (or a Message). SystemPromptTemplate does the same but produces a SystemMessage.

open as a page

When and how would you build tools programmatically (FunctionToolCallback / ToolCallbacks.from) and pass out-of-band data with ToolContext?

level: middleimportance: should knowfreq 35%

basics

~20 s

Besides @Tool methods, you can build a ToolCallback programmatically — e.g. FunctionToolCallback wraps a Function/BiFunction with a name, description and input type. ToolCallbacks.from(obj) turns @Tool methods into callbacks. ToolContext lets you pass data (like a user ID) to the tool without exposing it to the model.

open as a page

What are Advisors in Spring AI's ChatClient, and what does the advisor chain let you do?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Advisors are interceptors in ChatClient that wrap each request and response, letting you inject behavior around the model call — like adding conversation memory, retrieving documents for RAG, logging, or guarding content — without changing your prompt code. They run as an ordered chain, similar to a filter chain.

open as a page

What is the difference between MessageChatMemoryAdvisor and PromptChatMemoryAdvisor?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Both replay conversation history, but differently. MessageChatMemoryAdvisor injects the history as a list of separate role-tagged messages (USER/ASSISTANT). PromptChatMemoryAdvisor flattens the history into text inside the system prompt. Message-based preserves structure; prompt-based is a single blended system message.

open as a page

Explain MessageWindowChatMemory's windowing behavior and how ChatMemory differs from ChatMemoryRepository.

level: seniorimportance: should knowfreq 45%

basics

~20 s

MessageWindowChatMemory keeps only the most recent N messages (default 20), evicting the oldest so the prompt stays bounded — but it retains system messages. ChatMemory is the policy layer (windowing); ChatMemoryRepository is the pure storage layer it delegates persistence to.

open as a page

How do you choose between VectorStore implementations like PgVector, Redis, and Chroma, and what configuration matters?

level: seniorimportance: should knowfreq 40%

basics

~20 s

All implement the same VectorStore interface, so code stays the same — you pick based on your existing infrastructure. Use PgVector if you already run Postgres, Redis if you need speed and already run Redis, Chroma for a lightweight AI-focused store. Key config: dimensions, distance metric, and index type.

open as a page

How do you consume an external MCP server's tools inside a Spring AI ChatClient, i.e. as ToolCallbacks?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Add spring-ai-starter-mcp-client and configure the server connection (stdio command or SSE url). The auto-config exposes a ToolCallbackProvider that adapts each remote tool into a Spring AI ToolCallback. Inject it and pass it to ChatClient via defaultToolCallbacks — the model can then call remote tools like local ones.

open as a page

How does QuestionAnswerAdvisor perform context injection, and how do you control what it retrieves?

level: seniorimportance: should knowfreq 55%

basics

~20 s

It runs a similarity search using a SearchRequest (topK, similarityThreshold, filter), formats the returned Documents' text, and substitutes them into a prompt template at the {question_answer_context} placeholder before the model call. You tune retrieval via the SearchRequest and override the template.

open as a page

How does RetrievalAugmentationAdvisor differ from QuestionAnswerAdvisor, and what modular stages does it support?

level: seniorimportance: should knowfreq 40%

basics

~10 s

RetrievalAugmentationAdvisor is Spring AI's modular RAG advisor. Unlike the simple QuestionAnswerAdvisor, it lets you compose pre-retrieval query transformation/expansion, a pluggable DocumentRetriever, post-retrieval processing, and a QueryAugmenter that injects context and can handle empty results.

open as a page

Explain internal vs user-controlled tool execution and the returnDirect option. When would you disable internal execution?

level: seniorimportance: should knowfreq 30%

basics

~20 s

By default Spring runs tools internally: it executes the requested tool and re-calls the model automatically until a final answer. You can disable this (internalToolExecutionEnabled=false) to get the raw tool-call request and run it yourself. returnDirect=true returns the tool's result straight to the caller instead of sending it back to the model.

open as a page

How does multimodality work in Spring AI — how do you send an image (or audio) to a chat model, and what are the constraints?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Multimodality means the model accepts more than text — e.g. images or audio — as input. In Spring AI you attach a Media object (MimeType + data) to the user message: ChatClient.prompt().user(u -> u.text("...").media(MimeTypeUtils.IMAGE_PNG, resource)). Only vision/audio-capable models (e.g. GPT-4o, Claude) support it.

open as a page

How do you make Spring AI chat memory durable in production using JDBC or Cassandra ChatMemoryRepository, and what operational concerns arise?

level: principalimportance: should knowfreq 40%

basics

~20 s

Replace the default in-memory store with a persistent ChatMemoryRepository — JdbcChatMemoryRepository (rows in a relational table) or CassandraChatMemoryRepository (wide-column, supports TTL). Add the matching Spring Boot starter; it auto-configures the bean, and you plug it into MessageWindowChatMemory.

open as a page

How do embeddings and a VectorStore fit into a RAG retrieval pipeline, and what are the main failure modes to design around?

level: principalimportance: should knowfreq 42%

basics

~20 s

In RAG you ingest documents (read, chunk, embed, store), then at query time embed the user's question, run similaritySearch to fetch the most relevant chunks, and stuff them into the prompt as context. Main risks: bad chunking, model mismatch, weak filtering, and stale data.

open as a page

showing 1–30 of 36