skip to content

Scratchpad and Thinking Tokens

Reasoning tokens that are generated but not necessarily shown — extended-thinking modes, hidden scratchpads, and the API controls over how much of it you buy. Expect questions on the latency and cost you trade for accuracy, and on how much of the trace you can actually see.

on this pageshow

questions

4

What are thinking tokens in an extended-thinking LLM, and are you billed for them?

level: juniorimportance: must knowfreq 68%

answer

  1. generated before the reply
  2. a separate output channel
  3. you pay whether or not you see them
  4. summary returned, not the raw pass
  5. counts against the output-token cap

basics

~20 s

Thinking tokens are reasoning a model generates before its final answer, on a separate channel. They are generated tokens like any others, so you pay for them and wait for them even when the API returns only a summary, or nothing at all.

solid answer

~50 s

Extended-thinking (or reasoning) modes let a model produce a scratchpad pass before it produces the reply. Those reasoning tokens are decoded one at a time exactly like answer tokens, so they cost output-token money and wall-clock latency; hiding them from the user is a presentation choice, not a discount. What actually comes back over the API varies: some providers return the full trace, some return a short model-written summary of it, and some return an opaque or encrypted block you can pass back but not read — usually alongside a token count so you can reconcile the bill. Reasoning tokens also typically count against the same output allowance as the answer, so a tight output cap can be consumed by thinking and leave a truncated or empty reply. Assume you are paying for every one of them.

go deeper

for a junior

Be ready to say plainly that thinking tokens are generated text before the answer, billed like output and adding wall-clock delay. Knowing that some providers return only a summary already puts you ahead.

for a middle

Explain the three return shapes — full trace, summary, opaque blob — and why reasoning tokens counting against the output cap produces empty or truncated replies when the cap is set too tight.

for a senior

Show that you instrument reasoning tokens separately in cost and latency dashboards, and that you size output caps as budget plus answer plus headroom. Expect questions about variance, not averages.

for a principal

Own the framing that a reasoning model shifts spend from predictable to input-dependent. Talk about capacity planning from measured token distributions and about which product surfaces can absorb a pre-answer silence at all.

## The idea A plain language model turns your prompt straight into a reply. An **extended-thinking** model (also called a reasoning model) first generates a private working pass — a scratchpad — and only then generates the reply the user sees. Mechanically nothing exotic is happening: the model is autoregressively decoding tokens one at a time, as always. What is new is that those tokens are tagged as a distinct channel and the product decides whether to render them. That framing answers most interview questions on this subject. **Thinking tokens are ordinary generated tokens wearing a label.** Everything true of generated tokens is true of them: they cost output-rate money, they take time to produce, they occupy the context window, and they can be truncated. ## Visible, summarised, hidden There are three shapes you will meet in practice, and you should know which one you are buying: - **Visible trace.** The API returns the reasoning text alongside the answer. You can log it, inspect it, and (if you choose) render it. - **Summarised trace.** The provider returns a shorter, model-written account of the reasoning rather than the raw tokens. This is a *description* of the pass, not the pass itself, so it is not a faithful transcript and should not be treated as one. - **Opaque or encrypted trace.** You get a token count and a blob you may be required to hand back on the next call, but cannot read. This exists partly for provider IP reasons and partly for safety: raw traces contain candidate content the model considered and rejected. In all three cases the tokens were generated. Billing is based on generation, not on delivery. ## Where the cost and latency land Two practical consequences fall out immediately. **Cost.** Reasoning tokens are billed at the output rate, which is the expensive rate. A task where the model thinks for 3,000 tokens and answers in 200 is billed like a 3,200-token answer. If you only instrument the visible reply length, your cost model will be wrong by an order of magnitude, and it will be wrong *variably*, because how much a model thinks depends on how hard it judges the input to be. **Latency.** The user sees nothing until the reasoning pass finishes and the answer channel starts. On an interactive surface this shows up as a long silence before the first visible character — the time-to-first-answer-token is pushed out by the entire thinking pass. This is why extended thinking is a poor fit for tightly turn-taken interfaces (a voice ordering flow, for example, where several seconds of silence breaks the conversational rhythm) and a fine fit for a batch document pipeline where nobody is waiting. ## The output-allowance trap A specific failure a junior engineer hits on day one: reasoning tokens usually count against the *same* maximum-output-token limit as the answer. If you cap output at 1,024 tokens and the model spends 1,000 of them thinking, you get a truncated or empty answer and a confusing bill. The fix is to size the output cap as *thinking budget plus expected answer length plus headroom*, and to treat a suspiciously empty response as a budget symptom rather than a model failure. ## Controlling it Providers expose the dial in one of two shapes: a **token budget** (an upper bound on how much reasoning to spend) or a coarse **effort level** such as low/medium/high, sometimes including an option to turn thinking off entirely. Both are ceilings, not quotas: a model given a large budget on an easy input will typically use a fraction of it. That variance is why per-request cost on a reasoning model is far less predictable than on a plain model, and why capacity planning has to work from measured distributions rather than a single average. ## What thinking tokens are not They are not a reliable explanation of the answer. A trace is text the model generated on the way to a reply; it is not an audited account of the computation that produced it, and studies of reasoning faithfulness show traces that omit the influence actually driving the answer. Treat a trace as a debugging aid and a cost line item, not as evidence you can show a regulator or paste into a user-facing justification. They are also not free context. Within a single request the reasoning occupies context; across turns, providers differ on whether prior thinking is carried forward at all, and many drop it by default so that a long conversation does not accumulate scratchpads. ## How to answer this in an interview Say what they are (a separate generated channel before the reply), then immediately say the two things that matter operationally: you pay output rate for them whether or not you see them, and they land on the latency path in front of the first visible token. Mentioning that the returned trace may be a summary rather than the raw pass signals that you have actually shipped against one of these APIs. Practice varies by provider and moves quickly, so the durable knowledge is the mechanism, not any one vendor's current field names.

  • If the provider returns only a summary of the reasoning, what does that cost you operationally?
    You lose the ability to debug the actual pass. The summary is itself model-generated, so it can be tidier and more coherent than the real reasoning, which makes it misleading evidence when you are diagnosing a wrong answer. It is still useful for spotting gross misreadings of the input, and you still pay for the underlying tokens, so budget from the reported reasoning-token count rather than from the summary's length.
  • Your average reply is 150 tokens but your bill looks like 2,000. How do you confirm reasoning is the cause?
    Read the per-request token accounting the API returns — reasoning tokens are reported as their own count, distinct from the visible completion. Aggregate that field across a day and compare it with visible output length. If the gap tracks input difficulty, that is thinking, not a leak. Then check whether you enabled a high effort level or a large budget by default on requests that do not need it.
  • Does turning extended thinking off make a model worse at every task?
    No. It removes a pass that helps on multi-step, constraint-heavy work and does little for lookup, formatting, extraction or classification tasks that a single forward pass already handles. On simple inputs the extra pass mostly buys latency and cost. The right move is to measure on your own task set rather than assume the reasoning mode dominates.

saying these in an interview costs you the question

  • Hidden reasoning is free because it is not returned
  • Thinking adds no latency since users do not see it
  • The returned summary is the verbatim reasoning trace
  • Reasoning tokens are billed at the input rate
  • A reasoning trace is an audit-grade explanation of the answer

context

open as a page

How do you choose an extended-thinking budget when doubling it doubles cost per document?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Treat the budget as an empirical dial, not a preference. Sweep several budgets over a labelled eval set, plot accuracy against spend, and pick the knee where extra reasoning stops buying correctness. Remember the budget is a ceiling the model often underspends.

open as a page

With extended thinking enabled, does adding "think step by step" to the prompt still help?

level: middleimportance: should knowfreq 44%

basics

~20 s

Usually not. A reasoning model already runs its own scratchpad pass, so generic step-by-step instructions mostly duplicate it, burn tokens, and can push reasoning-style prose into the visible reply. Task-specific guidance about what to check and how to format still helps.

open as a page

How should you handle hidden reasoning traces in logs, retention and multi-turn calls?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Treat a reasoning trace as sensitive conversation data: same access controls, retention window and redaction as the transcript, never surfaced to end users. Across turns, follow the provider's rule on returning thinking blocks unmodified, since editing or dropping them can break the call.

open as a page