skip to content

In prompting, how does zero-shot CoT differ from few-shot CoT?

level: juniorimportance: must knowfreq 72%

answer

  1. two ways to switch reasoning on
  2. instruction versus demonstration
  3. one line versus repeated tokens
  4. examples also fix the output format

basics

~20 s

Zero-shot CoT adds only an instruction to reason step by step, and the model invents its own reasoning shape. Few-shot CoT puts solved examples in the prompt that demonstrate both how to reason and how the answer should look.

solid answer

~50 s

Both aim at the same thing — intermediate reasoning before the answer — but they carry different information. **Zero-shot CoT** is a bare instruction, classically "Let's think step by step", appended to the question. It costs one line, works on any input, and leaves the model free to choose what a step is and how to end. **Few-shot CoT** puts two to a handful of worked problems in the prompt, each with its reasoning written out and its answer in a fixed layout. Those demonstrations transfer domain procedure and output format, not just the general idea of reasoning. The practical consequences follow from that: few-shot buys format control and house-specific method at the price of tokens re-sent on every call and a tendency to pull odd inputs toward the demonstrated shape; zero-shot buys flexibility and a tiny prompt at the price of inconsistent structure.

go deeper

for a junior

Be able to state the difference plainly: zero-shot CoT is an instruction only, few-shot CoT includes solved examples. Name the classic trigger phrase and say that examples show both the reasoning and the answer layout.

for a middle

Explain what each actually supplies to the model — a mode versus a demonstrated procedure plus a format — and give the token consequence: exemplars are part of the input on every call, a trigger phrase is a few tokens once per call.

for a senior

Show you would pick based on measurement, not folklore: run both variants on a real eval slice, look at format consistency and accuracy separately, and be explicit about the input distribution the examples represent.

for a principal

Own the standardization call across a fleet of prompts: where exemplars are worth their recurring token and data-exposure cost, where a format template suffices instead, and how you keep a shared exemplar library from silently drifting away from the traffic it was written for.

## The distinction Chain-of-thought prompting means getting a model to write intermediate reasoning before committing to a final answer. There are two standard ways to elicit it, and they differ in what the prompt actually carries. **Zero-shot CoT** carries an *instruction* and nothing else — a directive such as "Let's think step by step" or "Work through the relevant factors before answering." No solved problems appear. The model decides what counts as a step, how many steps to produce, in what order, and how to present the conclusion. **Few-shot CoT** carries *demonstrations* — a small number of solved problems, each showing an input, a written-out reasoning trace, and a final answer in a consistent layout. The model infers the pattern from those examples and continues it for the new input. ## What a trigger phrase supplies A trigger supplies a mode, not knowledge. It tells the model that a longer, decomposed response is wanted rather than an immediate guess. What it cannot supply is anything specific to your problem: which considerations matter in your domain, which checks come first, what your house thresholds are, or how the answer should be laid out for downstream parsing. Two runs of the same zero-shot prompt on similar inputs can produce reasoning of very different length and structure, and the final answer may be buried in prose rather than sitting in a predictable slot. ## What worked examples supply Exemplars supply *method* and *format* simultaneously, and it is easy to forget the second half. Method: the examples encode a procedure — the order in which factors are weighed, the intermediate quantities that get named, the rule that is applied at the end. Format: they pin the output — roughly how long the reasoning runs, what sections it has, and exactly how the final answer is written, which is what makes the result reliably parseable without a separate extraction step. That second effect is the one candidates miss. Examples are not neutral illustrations; they are a strong prior over the shape of the response. ## Cost A trigger phrase is a handful of tokens. Exemplars are typically hundreds to thousands, and they are part of the input on *every* request, not a one-time setup. On a high-volume path that is a real line item, and it also consumes context that could hold retrieved material or conversation history. Caching a shared prompt prefix reduces the price of those repeated tokens but does not make them disappear, and it does not change the fact that whatever is in the exemplars is transmitted and logged on each call. ## When each is preferred Reach for **few-shot CoT** when the reasoning shape is unusual or house-specific — an internal escalation ladder, a scoring rubric, a domain notation — or when downstream code needs a strict, predictable output layout, or when zero-shot output is measurably inconsistent across similar inputs. Reach for **zero-shot CoT** when no worked solutions exist yet — a brand-new internal DSL with no solved cases is the clean example, and you simply cannot write exemplars for reasoning nobody has done before. Also prefer it when inputs vary so widely that any small set of examples would misrepresent most of them, when the token budget per call is tight relative to the task's difficulty, or when honest exemplars would have to be written from sensitive records. The two are not exclusive. A common production shape is a short instruction plus a strict output template and no full worked traces: you get format control without paying for complete exemplar reasoning. ## What has changed by mid-2026 Models trained to reason produce chains by default, so the literal phrase "let's think step by step" adds far less than it did when the technique was first reported, and on such models the effort dial has largely moved out of the prompt text and into a provider-level control. The interesting question has shifted from *how do I switch reasoning on* to *do I hand the model exemplars at all*, which is a format-and-cost decision rather than a reasoning-elicitation one. State your assumed model class when answering, because the honest answer differs between a small instruction-tuned model and a reasoning-trained frontier model. ## Common mistakes Treating few-shot as strictly better: on a task whose inputs are heterogeneous, examples can narrow the model rather than help it. Treating the trigger phrase as a universal accuracy booster: it does little on trivially easy tasks and on models that already reason. And forgetting that exemplars are re-billed and re-transmitted on every single request.

  • If downstream code has to parse the answer, which of the two makes your life easier?
    Few-shot, usually. The exemplars end with the answer in a fixed layout, and the completion tends to follow it, so a simple parser works. Zero-shot ends in free prose and needs either an explicit output template in the instruction or a separate extraction step. If you only need the format and not the demonstrated method, a template instruction is the cheaper way to get the same benefit.
  • What do you do when the task is a brand-new internal DSL with no worked solutions in existence?
    Start zero-shot, because there is nothing to demonstrate — you cannot write exemplar reasoning for a procedure nobody has performed yet. Give the model the grammar and the constraints, let it reason freely, and evaluate. Once the format has stabilized and you have verified good runs, a few of them can be hand-checked and promoted into exemplars, at which point you can compare the two variants on the same eval set.
  • On a reasoning-trained model, does adding "let's think step by step" still help?
    Usually very little, as of mid-2026. Those models already produce internal reasoning without being asked, so the phrase is mostly redundant and can even make the visible output more verbose without improving the answer. The lever that still matters on them is how much effort or thinking budget you allow, which is a provider-level setting rather than prompt text. Measure rather than assume — on smaller instruction-tuned models the trigger still earns its keep.

A trigger phrase is telling a new analyst to show their working; exemplars are handing them three completed worksheets and saying "like these".

saying these in an interview costs you the question

  • Says few-shot CoT is always more accurate than zero-shot
  • Treats "let's think step by step" as a universal accuracy booster
  • Forgets exemplar tokens are re-sent and billed on every request
  • Thinks exemplars only convey method, not output format
  • Assumes zero-shot CoT output is as easy to parse as few-shot

context