What is Spring AI's ChatClient, and how do you make a basic call to a model with it?
answer
- prompt().user().call().content()
- fluent client over ChatModel
- create() vs builder()
- content / chatResponse / entity
- call() blocks
basics
~10 sChatClient is Spring AI's fluent client for talking to an LLM. You write chatClient.prompt().user("...").call().content() to send a user message and get the reply back as a String.
solid answer
~40 sChatClient is the high-level, fluent API in Spring AI for calling a chat model (an LLM). You build it once from an auto-configured ChatModel bean via ChatClient.create(chatModel) or ChatClient.builder(chatModel).build(). To send a request you chain: prompt() starts a request spec, user("...") sets the user message (optionally system("...") for a system prompt), then call() executes synchronously. From the call spec you extract the result: content() gives the reply text as a String, chatResponse() gives the full ChatResponse (tokens, metadata, finish reason), and entity(MyType.class) maps the reply into a typed object (structured output). It mirrors the ergonomics of WebClient/RestClient, so the fluent chain reads left-to-right from building the prompt to extracting the answer.
code
java · 18 lines@Service
class AssistantService {
private final ChatClient chatClient;
AssistantService(ChatModel chatModel) { // auto-configured bean
this.chatClient = ChatClient.builder(chatModel)
.defaultSystem("You are a helpful travel guide.")
.build();
}
String tips(String city) {
return chatClient.prompt()
.user(u -> u.text("Name three things to do in {city}.")
.param("city", city))
.call()
.content(); // String reply
}
}go deeper
Know the happy-path chain prompt().user().call().content() and that it returns the reply text.
Distinguish content() vs chatResponse() vs entity(); know it wraps a ChatModel and is built via builder() with defaults.
Explain thread-safety/reuse, structured output via entity(), and that the fluent spec is the per-request state.
Frame ChatClient as the app-facing seam that keeps business code provider-agnostic and testable (mock the ChatModel or client).
**Spring AI** is Spring's integration library for building AI/LLM features. An **LLM (large language model)** is a text-generation model like OpenAI GPT, Anthropic Claude, or a local Ollama model. Spring AI gives you two layers to call one: - **ChatModel** — the low-level portable interface (one method conceptually: take a Prompt, return a ChatResponse). Spring Boot auto-configures a ChatModel bean for whichever provider starter is on the classpath. - **ChatClient** — a **fluent (builder-style) client** layered on top of ChatModel that removes boilerplate: assembling messages, setting options, attaching advisors, and extracting the result. **Creating a ChatClient.** Inject the auto-configured ChatModel and build once (typically in a @Configuration or @Service): ```java ChatClient chatClient = ChatClient.create(chatModel); // quick ChatClient chatClient = ChatClient.builder(chatModel) // configurable .defaultSystem("You are a terse assistant.") .build(); ``` ChatClient is thread-safe and meant to be reused as a bean. **Making a call — the fluent chain:** 1. `prompt()` — begins a *request spec*. Overloads: `prompt(String)` sets the user text directly, or `prompt(Prompt)` passes a pre-built Prompt object. 2. `.system("...")` / `.user("...")` — set the **system message** (instructions/persona) and **user message** (the actual question). You can also pass a lambda to add parameters for template substitution. 3. `.call()` — executes the request **synchronously (blocking)** and returns a *call response spec*. 4. Extract the result: - `.content()` → `String`, just the reply text. - `.chatResponse()` → `ChatResponse`, the full envelope: generations, `ChatResponseMetadata` (token usage, model name, finish reason). - `.entity(Class<T>)` / `.entity(ParameterizedTypeReference<T>)` → maps the reply to a typed object using structured-output converters (great for JSON-shaped answers). **Full example:** ```java String answer = chatClient.prompt() .system("You are a helpful travel guide.") .user("Name three things to do in Lisbon.") .call() .content(); ``` **Gotchas / when to use:** - Use **ChatClient** for almost all app code — it is the recommended entry point. Drop to **ChatModel** only for low-level needs. - `call()` **blocks** the calling thread until the whole reply is generated; for token-by-token output use `stream()` instead. - Build the client **once** and reuse it; don't rebuild per request. - `content()` can be null if the model returned no text (e.g. a pure tool-call response) — prefer `chatResponse()` when you need to inspect that. - The same code works across providers; only the starter dependency and config change.
- How do you get token usage or the finish reason instead of just the text?Call .chatResponse() instead of .content(); ChatResponse.getMetadata() exposes Usage (prompt/completion/total tokens) and each Generation carries metadata like the finish reason.
- Should you create a ChatClient per request?No. It is thread-safe and expensive-ish to build; create it once (e.g. in the constructor or a @Bean) and reuse it. The per-request state lives in the prompt() spec, not the client.
saying these in an interview costs you the question
- Thinking ChatClient is provider-specific (e.g. an 'OpenAI client')
- Believing call() streams tokens incrementally
- Rebuilding the client on every request
- Confusing content() (String) with chatResponse() (full envelope)