Chain of Thought
Making the model show its work: explicit intermediate steps, the zero-shot and few-shot ways to elicit them, and sampling several chains to vote on an answer. Interviewers ask about it constantly, including the part people skip — when a reasoning trace is unfaithful or actively hurts.
on this pageshowhide
explore
- Intermediate Reasoning Steps5 questions
- Zero-Shot vs Few-Shot CoT4 questions
- Self-Consistency Decoding10 questions
- Sampling for Diverse Reasoning Paths4 questions
- Answer Aggregation & Verification6 questions
- Scratchpad and Thinking Tokens4 questions
- Limitations and Failure Modes5 questions
- CoT Evaluation Methods5 questions
questions
page 2 of 2Zero-shot or few-shot CoT for a bank's KYC risk memos — how do you choose?
basics
~20 sCost and data exposure decide it. Exemplars are re-sent and billed on every request, and exemplars written from real customer files put personal data into every call and every log. Prefer synthetic policy-derived examples, or zero-shot with a strict output template.
When does weighting self-consistency votes by model confidence beat a plain count?
basics
~20 sWeighted voting helps mainly when the sample count is small and the answer space is short and constrained, so a per-sample likelihood score is meaningful. It hurts when confidence is miscalibrated, because it lets a few confidently wrong samples outvote a correct plurality.
What does an LLM gain and lose by reasoning in words rather than latent state?
basics
~20 sVerbalizing forces each reasoning step through one discrete token, discarding most of the model's richer internal state - but it produces a trace that can be read, checked, cached, edited and monitored. Latent reasoning keeps the bandwidth and gives up the audit surface.
showing 31–33 of 33