skip to content

ChatClient & ChatModel

ChatClient is a fluent API over a portable ChatModel, so the same code runs against different providers, with advisors wrapping each call. Interviewers ask what portability actually buys you, and swapping providers without rewriting is the case.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

5

What is Spring AI's ChatClient, and how do you make a basic call to a model with it?

level: juniorimportance: must knowfreq 70%

answer

  1. prompt().user().call().content()
  2. fluent client over ChatModel
  3. create() vs builder()
  4. content / chatResponse / entity
  5. call() blocks

basics

~10 s

ChatClient is Spring AI's fluent client for talking to an LLM. You write chatClient.prompt().user("...").call().content() to send a user message and get the reply back as a String.

solid answer

~40 s

ChatClient is the high-level, fluent API in Spring AI for calling a chat model (an LLM). You build it once from an auto-configured ChatModel bean via ChatClient.create(chatModel) or ChatClient.builder(chatModel).build(). To send a request you chain: prompt() starts a request spec, user("...") sets the user message (optionally system("...") for a system prompt), then call() executes synchronously. From the call spec you extract the result: content() gives the reply text as a String, chatResponse() gives the full ChatResponse (tokens, metadata, finish reason), and entity(MyType.class) maps the reply into a typed object (structured output). It mirrors the ergonomics of WebClient/RestClient, so the fluent chain reads left-to-right from building the prompt to extracting the answer.

code

java · 18 lines
java
@Service
class AssistantService {
    private final ChatClient chatClient;

    AssistantService(ChatModel chatModel) {           // auto-configured bean
        this.chatClient = ChatClient.builder(chatModel)
            .defaultSystem("You are a helpful travel guide.")
            .build();
    }

    String tips(String city) {
        return chatClient.prompt()
            .user(u -> u.text("Name three things to do in {city}.")
                        .param("city", city))
            .call()
            .content();                               // String reply
    }
}

go deeper

for a junior

Know the happy-path chain prompt().user().call().content() and that it returns the reply text.

for a middle

Distinguish content() vs chatResponse() vs entity(); know it wraps a ChatModel and is built via builder() with defaults.

for a senior

Explain thread-safety/reuse, structured output via entity(), and that the fluent spec is the per-request state.

for a principal

Frame ChatClient as the app-facing seam that keeps business code provider-agnostic and testable (mock the ChatModel or client).

**Spring AI** is Spring's integration library for building AI/LLM features. An **LLM (large language model)** is a text-generation model like OpenAI GPT, Anthropic Claude, or a local Ollama model. Spring AI gives you two layers to call one: - **ChatModel** — the low-level portable interface (one method conceptually: take a Prompt, return a ChatResponse). Spring Boot auto-configures a ChatModel bean for whichever provider starter is on the classpath. - **ChatClient** — a **fluent (builder-style) client** layered on top of ChatModel that removes boilerplate: assembling messages, setting options, attaching advisors, and extracting the result. **Creating a ChatClient.** Inject the auto-configured ChatModel and build once (typically in a @Configuration or @Service): ```java ChatClient chatClient = ChatClient.create(chatModel); // quick ChatClient chatClient = ChatClient.builder(chatModel) // configurable .defaultSystem("You are a terse assistant.") .build(); ``` ChatClient is thread-safe and meant to be reused as a bean. **Making a call — the fluent chain:** 1. `prompt()` — begins a *request spec*. Overloads: `prompt(String)` sets the user text directly, or `prompt(Prompt)` passes a pre-built Prompt object. 2. `.system("...")` / `.user("...")` — set the **system message** (instructions/persona) and **user message** (the actual question). You can also pass a lambda to add parameters for template substitution. 3. `.call()` — executes the request **synchronously (blocking)** and returns a *call response spec*. 4. Extract the result: - `.content()` → `String`, just the reply text. - `.chatResponse()` → `ChatResponse`, the full envelope: generations, `ChatResponseMetadata` (token usage, model name, finish reason). - `.entity(Class<T>)` / `.entity(ParameterizedTypeReference<T>)` → maps the reply to a typed object using structured-output converters (great for JSON-shaped answers). **Full example:** ```java String answer = chatClient.prompt() .system("You are a helpful travel guide.") .user("Name three things to do in Lisbon.") .call() .content(); ``` **Gotchas / when to use:** - Use **ChatClient** for almost all app code — it is the recommended entry point. Drop to **ChatModel** only for low-level needs. - `call()` **blocks** the calling thread until the whole reply is generated; for token-by-token output use `stream()` instead. - Build the client **once** and reuse it; don't rebuild per request. - `content()` can be null if the model returned no text (e.g. a pure tool-call response) — prefer `chatResponse()` when you need to inspect that. - The same code works across providers; only the starter dependency and config change.

  • How do you get token usage or the finish reason instead of just the text?
    Call .chatResponse() instead of .content(); ChatResponse.getMetadata() exposes Usage (prompt/completion/total tokens) and each Generation carries metadata like the finish reason.
  • Should you create a ChatClient per request?
    No. It is thread-safe and expensive-ish to build; create it once (e.g. in the constructor or a @Bean) and reuse it. The per-request state lives in the prompt() spec, not the client.

saying these in an interview costs you the question

  • Thinking ChatClient is provider-specific (e.g. an 'OpenAI client')
  • Believing call() streams tokens incrementally
  • Rebuilding the client on every request
  • Confusing content() (String) with chatResponse() (full envelope)

context

open as a page

What is the difference between ChatModel and ChatClient in Spring AI, and how does ChatModel provide provider portability?

level: middleimportance: must knowfreq 65%

basics

~20 s

ChatModel is the low-level portable interface every provider (OpenAI, Anthropic, Ollama, Azure) implements. ChatClient is the fluent, higher-level API built on top of a ChatModel to make calls ergonomic. Your code targets both, so swapping providers just means swapping the starter dependency.

open as a page

What is the difference between call() and stream() on ChatClient, and when would you use each?

level: middleimportance: should knowfreq 55%

basics

~20 s

call() runs synchronously and blocks until the whole reply is ready, giving you a String or ChatResponse. stream() is reactive: it returns a Flux that emits the reply piece by piece as the model generates it, so you can show tokens live.

open as a page

What are Advisors in Spring AI's ChatClient, and what does the advisor chain let you do?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Advisors are interceptors in ChatClient that wrap each request and response, letting you inject behavior around the model call — like adding conversation memory, retrieving documents for RAG, logging, or guarding content — without changing your prompt code. They run as an ordered chain, similar to a filter chain.

open as a page

How would you implement a custom Advisor, and how does its ordering in the chain affect behavior on the call() and stream() paths?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

You write a class implementing CallAdvisor (for blocking calls) and/or StreamAdvisor (for streaming), mutate the request, delegate to the next advisor in the chain, then optionally transform the response. getOrder() decides its position — a lower order runs earlier/further from the model. If you support both call() and stream(), you implement both interfaces so the two paths behave the same.

open as a page