With extended thinking enabled, does adding "think step by step" to the prompt still help?
answer
- the mode already does it
- boilerplate versus substance
- watch the visible channel stay clean
- specific checks beat generic nudges
- non-reasoning models are the exception
basics
~20 sUsually not. A reasoning model already runs its own scratchpad pass, so generic step-by-step instructions mostly duplicate it, burn tokens, and can push reasoning-style prose into the visible reply. Task-specific guidance about what to check and how to format still helps.
solid answer
~50 sExtended thinking gives the model a dedicated reasoning channel it was trained to use, so telling it to reason is redundant — the instruction is already satisfied by the mode. In practice, generic step-by-step boilerplate on a reasoning model buys little accuracy and has two side effects: it spends prompt and output tokens, and it can leak reasoning-shaped prose into the answer channel, breaking output contracts that expect a clean structured reply. What still pays is *specific* content: the checks that matter for this task, the domain constraints, edge cases to watch for, and the exact output shape. That is guidance the model cannot infer, as opposed to a nudge to do something it is already doing. On a plain non-reasoning model the calculus flips entirely — there prompt instructions are the only lever you have. So the answer is: measure on your task, and prefer substance over reasoning boilerplate.
go deeper
Know that reasoning models already produce a reasoning pass, so telling them to think step by step largely repeats what the mode does. Instructions about what to check and how to format the answer still matter.
Explain the two mechanisms separately — prompt-elicited reasoning in the visible channel versus a mode-level thinking channel — and why the redundant instruction leaks deliberation into a reply that was supposed to be structured.
Demonstrate that you settle this with an A/B on a real eval set measuring accuracy, output tokens and format-violation rate, and that you keep prompt-level reasoning for non-reasoning models and for cases needing a visible, structured trace.
Own the prompt-portfolio question: the same task may run against reasoning and non-reasoning models across a fleet, so decide whether prompts are written once for the weakest model or maintained per tier, and who owns that duplication.
## Two ways to get a model to reason There are two distinct mechanisms that produce intermediate reasoning, and confusing them is a common interview stumble. **Prompt-level elicitation** works on any model: you ask for reasoning in the instructions, and the model writes it into the ordinary output before the answer. The reasoning lands in the visible reply unless you also impose an output contract that separates it. **Extended thinking** is a mode of the model and the serving stack: reasoning is generated on a separate channel before the reply, and post-training has specifically shaped how the model uses that channel. You do not ask for it in the prompt; you enable it in the request. Once the second mechanism is on, the first is largely redundant. The model is already doing the thing the instruction asks for, on a channel built for it. ## What redundancy actually costs It is not free to leave the boilerplate in. - **Tokens both ways.** The instruction occupies prompt space, and inviting reasoning into the answer channel inflates output length on top of the thinking you already bought. - **Contract leakage.** This is the failure that bites in production. A service expecting a JSON object gets a paragraph of deliberation followed by the object, because the prompt asked for step-by-step working and the model complied — in the visible channel. Parsers break; strict-output modes fight the instruction. - **Style collision.** Reasoning models are tuned to a particular way of using their scratchpad. Heavy prescription of a different reasoning style in the prompt can work against that tuning rather than with it, and the effect is task-dependent enough that only your eval set can tell you. ## What still helps Drop the ritual, keep the substance. Instructions that continue to earn their place on a reasoning model are the ones carrying information the model cannot derive: - **Domain constraints** — which regulation applies, which fields are authoritative, what a valid value range is. - **Task-specific checks** — "verify the totals reconcile before answering", "confirm every cited clause exists in the provided text". This is not "think step by step"; it names *what* to think about. - **Output shape** — the schema, the field names, the requirement that the reply contain no deliberation. On a reasoning model this instruction becomes *more* important, not less, because it is the thing keeping the answer channel clean. - **Worked examples** where the value is demonstrating a format or an unusual convention, not demonstrating that reasoning exists. A useful test: would this sentence tell a competent human colleague something they did not know? If yes, keep it. If it just says "be thorough", cut it. ## When you still reach for prompt-level reasoning Three cases keep it alive: 1. **Non-reasoning models.** Cheaper or older models have no thinking channel; instructions are the whole toolkit. Many production systems route the easy majority to such a model, and there the prompt does the work. 2. **You need the reasoning visible and structured.** If a downstream reviewer must see the working in a specific shape, and the provider only returns a summary or an opaque blob, generating it in the answer channel under an explicit format is the only way to get it. 3. **Thinking is disabled for latency.** On a latency-bound surface you may turn extended thinking off and accept a shorter, prompt-elicited reasoning pass because you can bound and stream it. ## The honest state of it This is empirically contested ground. Reports differ on whether explicit reasoning instructions on reasoning models are neutral, mildly harmful, or occasionally helpful on specific task families, and the answer has moved with each model generation. What has held up is the direction: as reasoning modes matured, generic CoT boilerplate went from essential to roughly inert, while precise task instructions stayed valuable. In an interview, say that, then say you would A/B it on your own eval set — one arm with the boilerplate, one without, same budget — and read accuracy, output-token count, and format-violation rate. That answer survives the next model release; a confident universal claim will not.
- Your reasoning model started returning deliberation inside a JSON field. What is the first thing you check?The prompt's reasoning instructions. A leftover "show your working" or "think step by step" line tells the model to put deliberation in the visible channel, which is exactly what it did. Remove it, state explicitly that the reply must contain only the structured object, and use a strict structured-output mode if the provider offers one. Verify with a format-violation rate over the eval set rather than by eyeballing a few responses.
- If you deploy the same prompt across a reasoning model and a cheap non-reasoning one, how do you handle this?Either maintain two prompt variants, or write one prompt whose reasoning instruction is harmless on the stronger model and load-bearing on the weaker one — typically by naming specific checks rather than issuing a generic think-step-by-step. Evaluate both models against the same set; a prompt tuned only for the strong model tends to under-serve the weak one, which is where most of your volume usually goes.
- Does giving a reasoning model a worked example of the reasoning style still make sense?Rarely for the reasoning itself, often for the output. The model has been trained on how to use its own scratchpad, so demonstrating a reasoning style adds tokens and can conflict with that training. Demonstrating an unusual output convention, a domain-specific format, or a tricky edge case still pays, because that information is not derivable from the task description alone.
saying these in an interview costs you the question
- Adding step-by-step instructions always improves any model
- Reasoning models ignore the prompt's instructions entirely
- Prompt-elicited reasoning and extended thinking are the same mechanism
- Output-format instructions matter less once thinking is enabled
- A universal answer exists without evaluating your own task