What is least-to-most prompting, and how does it differ from a fixed pipeline?
answer
- plan first, then solve in order
- the model writes the subproblem list
- each solved answer joins the next context
- adaptivity bought with unpredictability
- a wrong plan sinks every later step
basics
~20 sLeast-to-most prompting asks the model to first list the subproblems a hard question contains, then solve them in dependency order, each answer feeding the next. The decomposition is produced at run time by the model, not authored in advance by an engineer.
solid answer
~50 sLeast-to-most prompting is a two-phase technique. The first phase asks the model to break the problem into simpler subproblems and order them; the second phase solves those subproblems one at a time, appending each solved answer to the context so later subproblems can use it. Take a multi-clause tax question: phase one enumerates the sub-questions, such as which income counts as taxable, which allowance applies to it, and how the residual is banded; phase two answers them in that order. The contrast with a fixed engineered pipeline is who authors the stages. In a fixed pipeline you wrote them, so cost, latency, and per-stage evals are predictable and each stage can have its own validator. With least-to-most the plan adapts to inputs of varying shape, but the plan itself is unverified, stage count varies per input, and a wrong or wrongly ordered decomposition sinks everything downstream.
go deeper
Recall the two phases in order: first ask for the list of simpler subproblems, then solve them one by one, carrying each answer forward into the next prompt.
Explain why the ordering helps, that each step is simpler because its prerequisites are already concrete values, and contrast model-generated plans with stages you author yourself.
Show how you would keep it safe in production: validate the plan structurally, cap the subproblem count with a rejection path, and attribute regressions to the decomposition or the solving phase.
Take a position on where adaptivity is worth its unpredictability. Argue for fixing the stages you can test and reserving model-generated decomposition for the input classes you cannot enumerate.
## The technique Least-to-most prompting attacks problems that are hard mainly because they are compositional: the final answer depends on several intermediate results that must be established in the right order. It runs in two phases. **Phase one, decomposition.** The model is prompted to restate the problem as an ordered list of simpler subproblems, easiest and most foundational first. Nothing is solved yet; the output is a plan. **Phase two, sequential solving.** Each subproblem is solved in turn, and each solved subproblem, with its answer, is carried into the context for the next. By the time the model reaches the final subproblem, the facts it needs have already been established in front of it rather than having to be derived in one leap. The name captures the ordering principle: start with the least complex piece and work up to the most complex, which is usually the original question restated. ## Why the ordering matters The reason this helps is not that the model gets more tokens to think with. It is that each solving step faces a problem strictly simpler than the original, with its prerequisites already resolved to concrete values. Consider a multi-clause tax question: gross income from two sources, one allowance that phases out above a threshold, and a band structure applied to what remains. Answered in one leap, the model frequently applies the allowance to the wrong base or bands the pre-allowance figure. Decomposed, the first subproblem yields a total, the second yields the allowance given that total, the third bands the residual. Each step is arithmetic a weaker model can do, and the dependency order stops the classic error of using a value before it exists. ## Against a fixed engineered pipeline A fixed pipeline hard-codes the same idea. You decided in advance that the stages are total, then allowance, then banding, and you wrote a prompt for each. The difference is authorship of the plan, and it drives every practical tradeoff. **Predictability.** A fixed pipeline has a known number of calls, so cost and latency are budgetable and a p95 is meaningful. Least-to-most may emit three subproblems for one input and eleven for another; the tail is unbounded unless you cap it. **Testability.** Fixed stages can each have a golden set and a validator, and a regression can be attributed to a stage. With a model-generated plan, the unit under test is the plan itself, which is harder to score than a field extraction. **Generality.** The fixed pipeline is brittle exactly where the input shape varies. If some questions involve a credit rather than an allowance, the fixed pipeline either has a stage that does nothing or needs a branch. Least-to-most handles heterogeneous inputs without you enumerating the cases, which is its real selling point. **Failure mode.** In a fixed pipeline, failure usually looks like a stage doing its job badly. In least-to-most, the dangerous failure is a confidently wrong plan: a missing subproblem, an inverted dependency, or a decomposition that quietly changes the question. Every downstream step then executes faithfully and the output is wrong in a way that looks reasoned. ## Making it survivable in production Treat the plan as an artifact, not as invisible scaffolding. Log it, and where a real pipeline exists, validate it: check that the number of subproblems is within a sane range, that required subproblems appear for known input classes, and that no step references a value not produced by an earlier step. A cheap deterministic check on the plan is worth more than extra prompting on the solving steps, because the plan is the single point of leverage. Sampling the decomposition more than once and comparing plans is a reasonable smoke test during development: if three samples produce structurally different plans for the same input, the problem statement is ambiguous or the task is at the edge of the model's competence, and no amount of downstream care will fix that. A pragmatic middle position is common in real systems. You fix the stages you understand and are willing to test, and you let the model decompose only inside the one stage that faces genuinely open-ended input. That keeps the budget and the evals of a pipeline while buying adaptivity exactly where it is needed. ## When it is not the answer When the model already solves the problem reliably in a single call, decomposition adds calls and latency for nothing, and modern models handle many compositional tasks that used to require it. When the task is not compositional at all, such as summarising, decomposition has nothing to order and tends to fragment the output. And when the problem needs several candidate decompositions to be explored and compared rather than one committed to, that is a different family of methods entirely; least-to-most commits to its first plan and executes it.
- How would you evaluate the decomposition phase separately from the solving phase?Score them on different data. For solving, feed a known-good plan and check the answers, which isolates arithmetic or lookup errors. For the plan, keep reference decompositions for a labelled set and check structural properties rather than exact wording: are the required subproblems present, is the dependency order valid, does any step consume a value no earlier step produced. That tells you which phase a regression belongs to.
- What do you do when the model's decomposition varies between runs on the same input?Treat the variance as a signal about the task, not just noise. Sample several plans and compare structure; if they disagree materially, the problem statement is ambiguous or the model is out of depth. Fixes in order: tighten the problem statement, provide worked example decompositions, or replace the open decomposition with fixed stages for that input class and keep the model's plan only where inputs are genuinely heterogeneous.
- Why not simply cap the number of subproblems the model may produce?A cap is a good budget control and a poor quality control. It bounds cost and latency and stops runaway plans, which is worth having. But truncating a genuinely eleven-step problem to five produces a plan that omits prerequisites, and the solving phase will happily execute the mutilated plan. Pair a cap with a rejection path: if the plan hits the cap, route the input to a fallback or to review rather than proceeding.
saying these in an interview costs you the question
- Says least-to-most just means telling the model to think step by step
- Assumes the generated plan is correct because the model produced it
- Ignores that stage count and cost vary per input
- Thinks it explores alternative decompositions rather than committing to one
- Uses it on non-compositional tasks like summarisation