What is the difference between call() and stream() on ChatClient, and when would you use each?
answer
- call() = blocking, whole answer
- stream() = Flux, token-by-token
- Flux<String> via .content()
- SSE + WebFlux endpoint
- entity() is call()-side
basics
~20 scall() runs synchronously and blocks until the whole reply is ready, giving you a String or ChatResponse. stream() is reactive: it returns a Flux that emits the reply piece by piece as the model generates it, so you can show tokens live.
solid answer
~40 sBoth start from prompt(), but they finish the request differently. call() executes synchronously and blocks the calling thread until the model finishes, then returns the complete result: .content() (String), .chatResponse() (full ChatResponse), or .entity(...) (typed). stream() executes reactively and returns a Reactor Flux — .content() gives Flux<String> of incremental text chunks and .chatResponse() gives Flux<ChatResponse> — emitting as the provider streams tokens (server-sent events under the hood). Use call() for simple request/response, batch jobs, or when you need the full object (token usage, structured output). Use stream() for chat UIs where you want a typewriter effect and lower time-to-first-token, typically wired to an SSE endpoint via a WebFlux controller returning Flux<String>. Note that per-chunk metadata like final token usage may only be complete on the terminal event.
code
java · 14 lines// Blocking: get the whole answer
String full = chatClient.prompt().user("Explain SSE").call().content();
// Reactive: stream chunks to an SSE endpoint
@RestController
class ChatController {
private final ChatClient chatClient;
ChatController(ChatModel m) { this.chatClient = ChatClient.create(m); }
@GetMapping(value = "/chat", produces = MediaType.TEXT_EVENT_STREAM_VALUE)
Flux<String> chat(@RequestParam String q) {
return chatClient.prompt().user(q).stream().content(); // emits as tokens arrive
}
}go deeper
Know call() waits for the full answer and stream() sends it in pieces for live display.
Explain Flux<String>, SSE, time-to-first-token, and when to pick each.
Discuss metadata-on-final-chunk, reactive error handling, and WebFlux wiring; know entity() is a call-side extractor.
Reason about UX latency budgets, back-pressure, and mixing streaming UI with structured/aggregated post-processing.
After you build a request with `prompt()....`, you terminate the chain with one of two operations that differ in **execution model**: **call() — synchronous/blocking.** - Returns a call-response spec whose extractors block the current thread until the model produces the *entire* answer: - `.content()` → `String` - `.chatResponse()` → `ChatResponse` (generations + metadata: token usage, finish reason) - `.entity(Class<T>)` → typed structured output - Simple mental model: one request, one complete response. Good for MVC controllers, background/batch processing, or anywhere you need the whole result before proceeding (e.g. to parse structured output or read token counts). **stream() — reactive/incremental.** - Returns a stream-response spec backed by **Project Reactor**: - `.content()` → `Flux<String>` — emits partial text as tokens arrive - `.chatResponse()` → `Flux<ChatResponse>` — emits partial ChatResponse objects - Under the hood the provider streams tokens (typically **Server-Sent Events**). This lowers **time-to-first-token** (the user sees output almost immediately) and enables a typewriter UX. - Typically exposed through a **Spring WebFlux** controller that returns the Flux directly (as `text/event-stream`) so the browser renders chunks live: ```java @GetMapping(value = "/chat", produces = MediaType.TEXT_EVENT_STREAM_VALUE) Flux<String> chat(@RequestParam String q) { return chatClient.prompt().user(q).stream().content(); } ``` **Choosing between them.** - **call()**: batch/back-end work, need the complete ChatResponse (usage, finish reason), structured output via entity(), simplest code. - **stream()**: interactive chat, long answers, perceived-latency-sensitive UIs. **Gotchas / edge cases.** - stream() requires **reactive plumbing**. In a blocking Spring MVC app you can still consume a Flux, but the natural fit is WebFlux; blocking on the Flux (e.g. `.block()`) throws away the streaming benefit. - **Metadata timing**: token usage and finish reason may only be fully populated on the **final** emission, not every chunk. Aggregate if you need the full text and metadata. - **Structured output (entity())** is a call()-side extractor — you generally can't map a half-formed streamed reply into an object mid-stream; collect the stream first or use call(). - **Error handling** differs: with a Flux, provider/network errors surface as `onError` signals you handle reactively (retry/onErrorResume), not as a thrown exception at the call site. - Advisors and default options apply identically to both; the streaming path uses the StreamAdvisor side of the advisor chain.
- Can you get total token usage from stream()?Yes, but usage/finish-reason are typically fully populated only on the final ChatResponse emission of the Flux (stream().chatResponse()), not on every chunk — so read it from the terminal element rather than assuming each chunk carries it.
- Why is stream() usually paired with WebFlux rather than MVC?stream() returns a Reactor Flux; WebFlux can return it directly as text/event-stream and back-pressure it. In blocking MVC you'd have to bridge/collect it, which loses the incremental delivery that makes streaming worthwhile.
saying these in an interview costs you the question
- Saying call() streams tokens incrementally
- Claiming stream() returns a String
- Assuming every streamed chunk carries final token usage
- Thinking you can reliably entity()-map a mid-stream partial reply