skip to content

A reasoning model thinks 11 seconds before the first word — what should the UI show?

level: seniorimportance: should knowfreq 40%

answer

  1. blank screen reads as hang
  2. the first answer token cannot move
  3. move the first visible token instead
  4. show the model's work, not a spinner
  5. provisional region, never the answer

basics

~20 s

Fill the gap with real progress, not a spinner. Stream the model's thinking summary as visible status lines in a clearly separate, provisional region, keep a cancel control available, and treat a long silence as a stall worth surfacing.

solid answer

~50 s

Eleven seconds of blank screen reads as a hang, and users reload — which costs you the whole generation twice. The fix is to move the first *visible* token earlier even though the first *answer* token cannot move: providers typically expose a summarized stream of the model's reasoning, and rendering it as a status line ("checking the wire copy…") turns dead air into evidence of work. Present it as visibly distinct from the answer — a muted, collapsible region that the answer does not inherit — because it is exploratory work, not a committed claim, and it may contradict the final output. Keep an explicit cancel affordance, since a long wait is exactly when users want out. Add a stall watchdog: if nothing arrives for some seconds, say so rather than letting an unchanging screen imply failure. Fake progress bars are worse than honest silence once users notice they are fake.

go deeper

for a junior

Know that reasoning models can take many seconds before any answer text, and that showing nothing at all during that time makes users think the app has frozen.

for a middle

Explain that the first answer token cannot move, so you move the first visible token instead — typically by streaming the model's summarized thinking as status, kept separate from the answer.

for a senior

Show production judgment: provisional rendering rules, a stall watchdog distinguishing thinking from a dead stream, an always-available cancel, and client-measured time to first visible token as the metric you hold yourself to.

for a principal

Own the tradeoff between reasoning depth and perceived responsiveness across request classes, and set the product rule for what model-produced intermediate output may ever be shown to a user.

## The problem this creates Extended-thinking models spend real time reasoning before emitting an answer token. As of mid-2026 that phase routinely runs for seconds and, at high effort settings on the largest models, for tens of seconds. From the client's point of view the request has been accepted, the stream is open, and nothing readable is arriving. Users interpret that exactly as an application hang: they reload, they resubmit, or they leave. A reload is the worst outcome available — the original generation usually keeps running and gets billed, and the user now waits a second full cycle. So the design problem is not "make the model faster". It is: the first answer token cannot move, therefore move the first *visible* token. ## Fill the gap with real signal A spinner conveys only "something is happening", which is precisely the claim the user has stopped believing after five seconds. The strong option is to show the model's own progress. Providers typically expose the reasoning phase as a stream of summarized thinking, and rendering those summaries as short status lines gives the user genuine, changing evidence that work is progressing — and it usually tells them something useful about how the request was understood, which is a chance to cancel early if the model has misread the task. Where no thinking stream is available, the next-best signals are honest and structural: name the phase ("reading the attached documents"), show elapsed time, or show progress through pipeline stages your harness actually knows about. Any of these beats an animation with no information content. ## Present it as provisional, not as the answer The thinking stream must be visually and semantically separate from the answer. It is exploratory: the model may consider an approach and reject it, and a summary line can contradict what the final answer says. If a user reads it as a claim, you have shipped a correctness bug through the UI. Practical rules that hold up in review: put it in a distinct, muted region; collapse or replace it once the answer begins, so the finished output does not carry rejected reasoning next to it; do not let it be copied or exported as part of the answer; and label it clearly as a summary of the model's process rather than as verbatim internal state. ## Keep the exit available A long wait is the moment users most want a way out, so the cancel control has to be present and obvious throughout the thinking phase — not appearing only once tokens start. Cancelling must actually propagate so generation stops rather than continuing invisibly. This pairs with the thinking summary: because the user can see what the model has understood, they can make an informed decision to stop within the first few seconds instead of waiting out a wrong answer. ## Distinguish thinking from stalled Long thinking and a broken stream look identical from the outside. Run a watchdog on the client: if no event of any kind has arrived for some threshold, change the message — acknowledge the delay, keep the elapsed timer running, and offer retry alongside cancel. If your transport sends periodic keep-alive events, absence of those is a stronger stall signal than absence of content. The important part is that the interface never sits unchanged for a long period, because an unchanging interface is how users conclude a system is dead. ## What not to do Do not fabricate progress. A progress bar timed to a historical average is a lie that breaks the moment a request runs long, and users who catch it stop trusting every indicator you show. Do not silently buffer the whole answer so it appears "all at once, quickly" — you have then given up the streaming benefit entirely and made the blank period longer. Do not report your latency as time to first byte and declare the problem solved: the byte arrived, the user saw nothing, and only a client-measured time to first *visible* token reflects what actually happened. ## The judgment call There is a real product tradeoff underneath: heavier reasoning buys better answers and costs perceived responsiveness. The UI techniques above buy you room to spend that time, but they do not make it free — a user staring at a status line for forty seconds is still waiting. Deciding how much thinking a given request class should get is a separate lever, and the honest framing in an interview is that the interface can make a long wait tolerable and legible, not that it can make it invisible.

  • Why present the thinking stream as a summary in its own region rather than as the beginning of the answer?
    Because reasoning is exploratory. The model may propose and then discard an approach, so a status line can contradict the final output. Rendering it as answer text turns rejected reasoning into an apparent claim, and users copy it or quote it. A muted, collapsible region that clears when the answer starts keeps the provisional material visibly provisional.
  • How do you tell a long thinking phase apart from a dead stream on the client?
    Run a watchdog on time since the last event of any kind, not since the last visible token. If the transport emits periodic keep-alives, their absence is a strong stall signal. Past a threshold, change the message to acknowledge the delay, keep an elapsed timer running, and offer retry beside cancel — so the interface never sits unchanged long enough to look dead.
  • What is wrong with a progress bar based on the historical average duration?
    It is fabricated. It reads well until a request runs longer than average, at which point the bar sits at 95% or completes while nothing has happened, and users learn that the indicator carries no information. That damages trust in every other status signal in the product. Honest elapsed time or real phase labels are less pretty and far more durable.

saying these in an interview costs you the question

  • Shows a bare spinner through a ten-second thinking phase
  • Renders thinking summaries inline as part of the final answer
  • Fabricates a progress bar from average response times
  • Hides the cancel control until the first answer token arrives
  • Reports time to first byte and declares the wait solved

context