skip to content

What is the difference between call() and stream() on ChatClient, and when would you use each?

level: middleimportance: should knowfreq 55%

answer

  1. call() = blocking, whole answer
  2. stream() = Flux, token-by-token
  3. Flux<String> via .content()
  4. SSE + WebFlux endpoint
  5. entity() is call()-side

basics

~20 s

call() runs synchronously and blocks until the whole reply is ready, giving you a String or ChatResponse. stream() is reactive: it returns a Flux that emits the reply piece by piece as the model generates it, so you can show tokens live.

solid answer

~40 s

Both start from prompt(), but they finish the request differently. call() executes synchronously and blocks the calling thread until the model finishes, then returns the complete result: .content() (String), .chatResponse() (full ChatResponse), or .entity(...) (typed). stream() executes reactively and returns a Reactor Flux — .content() gives Flux<String> of incremental text chunks and .chatResponse() gives Flux<ChatResponse> — emitting as the provider streams tokens (server-sent events under the hood). Use call() for simple request/response, batch jobs, or when you need the full object (token usage, structured output). Use stream() for chat UIs where you want a typewriter effect and lower time-to-first-token, typically wired to an SSE endpoint via a WebFlux controller returning Flux<String>. Note that per-chunk metadata like final token usage may only be complete on the terminal event.

code

java · 14 lines
java
// Blocking: get the whole answer
String full = chatClient.prompt().user("Explain SSE").call().content();

// Reactive: stream chunks to an SSE endpoint
@RestController
class ChatController {
    private final ChatClient chatClient;
    ChatController(ChatModel m) { this.chatClient = ChatClient.create(m); }

    @GetMapping(value = "/chat", produces = MediaType.TEXT_EVENT_STREAM_VALUE)
    Flux<String> chat(@RequestParam String q) {
        return chatClient.prompt().user(q).stream().content(); // emits as tokens arrive
    }
}

go deeper

for a junior

Know call() waits for the full answer and stream() sends it in pieces for live display.

for a middle

Explain Flux<String>, SSE, time-to-first-token, and when to pick each.

for a senior

Discuss metadata-on-final-chunk, reactive error handling, and WebFlux wiring; know entity() is a call-side extractor.

for a principal

Reason about UX latency budgets, back-pressure, and mixing streaming UI with structured/aggregated post-processing.

After you build a request with `prompt()....`, you terminate the chain with one of two operations that differ in **execution model**: **call() — synchronous/blocking.** - Returns a call-response spec whose extractors block the current thread until the model produces the *entire* answer: - `.content()` → `String` - `.chatResponse()` → `ChatResponse` (generations + metadata: token usage, finish reason) - `.entity(Class<T>)` → typed structured output - Simple mental model: one request, one complete response. Good for MVC controllers, background/batch processing, or anywhere you need the whole result before proceeding (e.g. to parse structured output or read token counts). **stream() — reactive/incremental.** - Returns a stream-response spec backed by **Project Reactor**: - `.content()` → `Flux<String>` — emits partial text as tokens arrive - `.chatResponse()` → `Flux<ChatResponse>` — emits partial ChatResponse objects - Under the hood the provider streams tokens (typically **Server-Sent Events**). This lowers **time-to-first-token** (the user sees output almost immediately) and enables a typewriter UX. - Typically exposed through a **Spring WebFlux** controller that returns the Flux directly (as `text/event-stream`) so the browser renders chunks live: ```java @GetMapping(value = "/chat", produces = MediaType.TEXT_EVENT_STREAM_VALUE) Flux<String> chat(@RequestParam String q) { return chatClient.prompt().user(q).stream().content(); } ``` **Choosing between them.** - **call()**: batch/back-end work, need the complete ChatResponse (usage, finish reason), structured output via entity(), simplest code. - **stream()**: interactive chat, long answers, perceived-latency-sensitive UIs. **Gotchas / edge cases.** - stream() requires **reactive plumbing**. In a blocking Spring MVC app you can still consume a Flux, but the natural fit is WebFlux; blocking on the Flux (e.g. `.block()`) throws away the streaming benefit. - **Metadata timing**: token usage and finish reason may only be fully populated on the **final** emission, not every chunk. Aggregate if you need the full text and metadata. - **Structured output (entity())** is a call()-side extractor — you generally can't map a half-formed streamed reply into an object mid-stream; collect the stream first or use call(). - **Error handling** differs: with a Flux, provider/network errors surface as `onError` signals you handle reactively (retry/onErrorResume), not as a thrown exception at the call site. - Advisors and default options apply identically to both; the streaming path uses the StreamAdvisor side of the advisor chain.

  • Can you get total token usage from stream()?
    Yes, but usage/finish-reason are typically fully populated only on the final ChatResponse emission of the Flux (stream().chatResponse()), not on every chunk — so read it from the terminal element rather than assuming each chunk carries it.
  • Why is stream() usually paired with WebFlux rather than MVC?
    stream() returns a Reactor Flux; WebFlux can return it directly as text/event-stream and back-pressure it. In blocking MVC you'd have to bridge/collect it, which loses the incremental delivery that makes streaming worthwhile.

saying these in an interview costs you the question

  • Saying call() streams tokens incrementally
  • Claiming stream() returns a String
  • Assuming every streamed chunk carries final token usage
  • Thinking you can reliably entity()-map a mid-stream partial reply

context