skip to content

After moving a streaming app to claude-sonnet-5, thinking blocks arrive empty — why?

level: seniorimportance: should knowfreq 32%

answer

  1. nothing errored, the UI just went quiet
  2. a default flipped between generations
  3. visibility only — still thought, still billed
  4. display defaults to omitted here
  5. ask for summarized explicitly

basics

~20 s

On claude-sonnet-5 the thinking parameter's display field defaults to omitted, so thinking blocks stream with empty text. The previous generation defaulted to summarized. Set thinking to {"type": "adaptive", "display": "summarized"} to get readable reasoning back.

solid answer

~50 s

This is a silent default change rather than a bug in your stream handling. On `claude-sonnet-5` the `display` field inside the `thinking` parameter defaults to `"omitted"`, which means thinking blocks are still emitted but their text is empty; on the previous Sonnet generation the default was `"summarized"`, so the same code used to render readable reasoning. In a streaming UI the symptom is characteristic: a long pause with nothing on screen while the model reasons, then the final answer arrives in a burst. The fix is to ask for it explicitly — `thinking: {"type": "adaptive", "display": "summarized"}`. Note that `display` controls visibility only. Thinking still happens and is billed identically under every setting, so switching to `"omitted"` is not a cost optimisation, and the raw chain of thought is never returned on any current model — `"summarized"` gives you a readable summary, not the verbatim reasoning.

code

python · 13 lines
python
from anthropic import Anthropic

client = Anthropic()

with client.messages.stream(
    model="claude-sonnet-5",
    max_tokens=8192,
    thinking={"type": "adaptive", "display": "summarized"},
    messages=[{"role": "user", "content": "Plan a zero-downtime schema migration."}],
) as stream:
    final = stream.get_final_message()

print(final.content)

go deeper

for a junior

Know that thinking has a display setting, and that seeing empty thinking text usually means it is set to omitted rather than that something broke. Setting display to summarized brings the readable version back.

for a middle

Explain that display defaults differ by model generation, that it controls visibility only while the reasoning still happens and is still billed, and that effort — not display — is the lever that changes reasoning spend.

for a senior

Show the diagnosis path: no errors plus a UX regression right after a model bump points at a changed default, so compare your request against the new model's defaults before debugging your own stream assembly, and pin UX-relevant settings explicitly.

for a principal

Own the upgrade discipline: model bumps must go through a canary that asserts on stream shape and user-visible behaviour, not just final text, and UX-critical request settings should be explicit in config so a silent default change can never reach users unannounced.

## The symptom A streaming client that previously rendered a "thinking…" panel goes blank after the model string is changed to `claude-sonnet-5`. The stream is healthy: events arrive, thinking content blocks are present in the sequence, the final text is correct. What is missing is the text inside those thinking blocks — it is an empty string. Users experience it as a long dead pause followed by a sudden answer, which reads as a hang. ## The cause: a default that flipped The `thinking` parameter carries an optional `display` field. `"summarized"` returns a readable summary of the model's reasoning; `"omitted"` emits the thinking blocks with empty text. On `claude-sonnet-5` the default is `"omitted"`. On the previous generation it was `"summarized"`. Nothing in your code changed, and nothing errors — the request is perfectly valid either way — so this is a behaviour regression that only surfaces in the UI. The fix is a one-line explicit setting: `thinking: {"type": "adaptive", "display": "summarized"}`. Any application that streams reasoning to users should set `display` explicitly rather than relying on the default, precisely because the default is a per-model property that can differ across generations. ## What display does and does not control Three things are worth separating, because interviews probe exactly this confusion. First, `display` is about visibility, not computation. The model reasons the same amount and you are billed the same either way. Choosing `"omitted"` to save money does nothing; if you want to spend less on reasoning, lower `output_config.effort` instead, which genuinely changes how much thinking happens. Second, `"summarized"` is a summary, not a transcript. The raw, verbatim chain of thought is not exposed on current models under any setting. Product copy that promises users "see exactly how the model thought" is overclaiming, and any logic that tries to parse the summary as a structured plan is building on sand — treat it as human-facing text. Third, `display` is independent of whether thinking is on. Thinking is governed by `thinking.type`; on `claude-sonnet-5` `{"type": "adaptive"}` is the on-mode and omitting the parameter entirely also runs adaptive. `display` merely decides what you get to look at. ## The other half: echoing thinking blocks back A related and more consequential rule concerns multi-turn conversations. When you continue a conversation on the same model, append the assistant's content blocks back into `messages` unchanged, thinking blocks included. Applications that extract only the visible text and append that string lose the thinking blocks, which degrades continuation quality. This is easy to get wrong precisely when `display` is `"omitted"`, because the blocks look empty and therefore look safe to drop — the block still matters even when its rendered text does not. Switching models mid-conversation is the exception: other models silently ignore thinking blocks they did not produce, so you do not need to strip them by hand. ## Diagnosing this class of problem The general lesson is that a model upgrade can change defaults without changing the contract. Nothing throws, nothing 400s, and unit tests that assert on the final text keep passing. Guard against it in three ways. Set the settings that matter to your UX explicitly rather than inheriting defaults — `display`, `effort`, and `max_tokens` are the usual suspects. Keep at least one integration check that exercises a real streamed request and asserts on the *shape* of what arrives, not only on the final answer. And when a model bump produces a UX change with no errors in the logs, compare the request you are sending against the new model's documented defaults before you go hunting in your own stream-assembly code, which is where teams usually waste the afternoon. ## Quick checklist for a Sonnet generation bump Set `display: "summarized"` if you render reasoning. Confirm the thinking configuration uses `{"type": "adaptive"}` rather than an older fixed-budget form. Verify you are appending full content blocks, not extracted text, on multi-turn calls. Re-check any dashboard that infers "reasoning happened" from non-empty thinking text, because that inference is now wrong by default.

  • Does setting display to omitted reduce what you are billed for reasoning?
    No. Thinking happens and is billed identically under every display setting — the field controls visibility only. If you want to spend less on reasoning, lower `output_config.effort`, which actually changes how much thinking the model does. Treating display as a cost lever is a common misreading and produces no savings at all.
  • What must you do with thinking blocks when continuing the same conversation on the same model?
    Append the assistant's full content blocks back into `messages` unchanged, thinking blocks included, rather than extracting the visible text and appending a string. Dropping them degrades continuation quality, and it is easiest to do accidentally when display is omitted and the blocks look empty. Switching to a different model is the exception — other models silently ignore thinking blocks they did not produce.
  • Does summarized give you the model's verbatim reasoning?
    No. It returns a readable summary of the reasoning; the raw chain of thought is not exposed on current models under any setting. So do not promise users a literal transcript, and do not parse the summary as if it were a structured plan — treat it as human-facing prose whose wording and length can change between model generations.

saying these in an interview costs you the question

  • Blames the SSE parser when thinking text is empty
  • Thinks omitted display saves reasoning tokens
  • Believes summarized returns the verbatim chain of thought
  • Assumes empty thinking text means thinking was off
  • Drops thinking blocks from history because they look empty

context