skip to content

Why does zero-shot CoT often need a second call to extract the answer?

level: middleimportance: should knowfreq 46%

answer

  1. reasoning first, answer second
  2. free prose has no answer slot
  3. the "therefore the answer is" cue
  4. two calls means paying twice
  5. parsing is the whole reason

basics

~20 s

A bare step-by-step instruction returns free-form prose with no fixed answer slot. The original zero-shot CoT recipe therefore uses two stages: generate the reasoning, then re-prompt with that text plus an extraction cue that yields a short, parseable answer.

solid answer

~50 s

Zero-shot CoT gives the model no template, so it ends wherever it ends — sometimes with a clean "the answer is 42", sometimes with a hedged paragraph, sometimes with two candidate answers discussed in sequence. That is fine for a human reader and hostile to a parser. The classic recipe splits it in two: the **reasoning stage** appends the trigger to the question and lets the model think freely, then the **extraction stage** re-sends the question plus the generated reasoning with a cue such as "Therefore, the answer is", so the completion is a short span you can parse deterministically. Few-shot CoT rarely needs this, because the exemplars already end in a fixed layout that the completion imitates. The cost is a second round trip and a second billing of the reasoning tokens, which is why modern practice usually collapses it into one call by asking for the reasoning followed by a delimited final line, or by constraining the response to a schema.

code

python · 10 lines
python
import re

CUE = re.compile(r"therefore,? the answer is\s*(.+?)\s*\.?$", re.IGNORECASE)

def extract_answer(extraction_stage_output: str) -> str | None:
    matches = CUE.findall(extraction_stage_output.strip())
    return matches[-1] if matches else None

print(extract_answer("Therefore, the answer is 42."))
print(extract_answer("The steps were long and inconclusive."))

go deeper

for a junior

Know that zero-shot CoT returns free prose, so something has to turn that prose into a value your code can use — either a second extraction prompt or an explicit output format in the instruction.

for a middle

Be able to describe both stages concretely, including the answer cue, and explain why few-shot prompts usually avoid the extra call because their exemplars fix the ending layout.

for a senior

Talk about what you would run in production: single call with a delimited final line or a response schema, a parser anchored on the last delimiter, loud failures, and an extraction-failure rate you actually monitor.

for a principal

Frame it as an interface contract between the model and the rest of the system: where output shape is enforced versus requested, what the fallback behaviour is on malformed output, and who owns that contract when the model behind the endpoint is upgraded.

## The problem the second call solves Zero-shot CoT trades format control for flexibility. You tell the model to reason and it does, but nothing in the prompt says where the answer goes or what it looks like. Across a batch of similar inputs you will see all of these endings: a bare value, a value inside a sentence with qualifications, a restatement of the whole chain, a correct value followed by a caveat that gestures at a different one, or reasoning that simply stops without concluding. A regular expression written against yesterday's output shape breaks on today's. That is a parsing problem, not a reasoning problem — and it is worth naming it that way in an interview, because the fix is a formatting mechanism, not a change to how the model thinks. ## The two-stage pattern The original zero-shot CoT recipe runs two prompts against the model: 1. **Reasoning stage.** Send the question with the trigger appended — question plus "Let's think step by step". Take the whole free-form completion. 2. **Extraction stage.** Send the question, the reasoning text from stage one, and an answer cue — a phrase engineered so that the only natural continuation is the answer itself. "Therefore, the answer (a number) is" is the canonical shape; the parenthetical type hint matters, because it steers the completion toward a bare value rather than another paragraph. Stage two is doing something narrow and mechanical: it converts prose into a slot-filling task. The model is not re-deciding the answer, it is reading its own transcript and reporting the conclusion. ## What it costs Two round trips instead of one, so roughly double the latency floor. The reasoning tokens are sent again as input on the second call, so you pay for them twice — once as output, once as input. On a high-volume path both matter, which is the main argument for collapsing the pattern. ## Collapsing it into one call Modern instruction-following models follow output-format directions well enough that the second call is usually avoidable. The single-call version asks for the reasoning and then a delimited final line — for example, instructing that the response must end with a line beginning `FINAL:` and containing only the answer. You parse the last such line. Where the provider supports constrained or schema-shaped responses, a two-field object with a reasoning field and an answer field gives the same guarantee without any string matching. Be careful with one variant: instructing the model to output *only* the answer, with the reasoning suppressed entirely, removes the benefit you were paying for. The reasoning has to be generated somewhere for it to help. ## Failure modes to watch **Ambiguous cues.** "The answer is" appears inside reasoning too, so anchoring a parser on the first occurrence rather than the last picks up an intermediate value. **Silent extraction failures.** If the parser returns nothing and the code falls back to an empty string, you get quiet wrong answers rather than errors. Fail loudly, count the failures, and alert on the rate — a rising extraction-failure rate is an early signal that the prompt or the model behind it has changed. **Truncation.** If the response is cut off by an output limit mid-reasoning, there is no final line to parse at all. Handle that case distinctly from a malformed answer, because the remedy is different. **Disagreement between the two stages.** Occasionally the extraction reports something the reasoning did not conclude. Treat that as a defect to measure, not a rounding error; if it is frequent, the reasoning was probably ambiguous and the prompt needs the answer criteria stated up front. ## Why few-shot mostly sidesteps this In a few-shot prompt every example ends with the answer in the same layout, and next-token prediction being what it is, the completion for a new input tends to land in that same layout. You still validate, but a single call plus a stable parser is usually enough. This is one of the concrete reasons teams accept exemplar token cost: it buys deterministic output shape as a side effect.

  • Why anchor the parser on the last occurrence of the cue rather than the first?
    Because the phrase can appear inside the reasoning itself — a model often writes "so the answer is X" mid-chain and then revises. Taking the first match captures an intermediate value that the model later abandoned. The last occurrence is the conclusion, which is what the extraction cue was placed to produce. Better still, use a delimiter the model is told to emit exactly once.
  • How would you avoid the second call entirely without losing parseability?
    Keep the reasoning but constrain the ending: instruct that the response must finish with a single line starting with a fixed marker containing only the answer, then parse that line. Where the provider supports a response schema, a two-field object with separate reasoning and answer fields is stronger, since the structure is enforced rather than requested. Both keep the reasoning that made CoT worth using in the first place.
  • What should happen when extraction finds nothing?
    Fail loudly and count it. Silent fallbacks to an empty string or to a default answer turn a formatting failure into an undetectable wrong answer. Emit a distinct error, keep the raw text for inspection, and track the extraction-failure rate as an operational metric — a step change in it is a reliable early signal that the prompt, the input distribution, or the model behind the endpoint has shifted.

saying these in an interview costs you the question

  • Thinks the second call re-computes the answer rather than reporting it
  • Suppresses the reasoning entirely and keeps only the answer instruction
  • Anchors the parser on the first "the answer is" in the text
  • Falls back to an empty string when extraction fails
  • Assumes few-shot CoT needs the same extraction stage

context