skip to content

Task Decomposition & Prompt Chaining

Splitting a hard request into ordered sub-prompts instead of one mega-prompt — least-to-most prompting, prompt chaining, and skeleton-first drafting. Interviewers ask when to chain versus when to trust a single call, because chaining buys controllability at the price of latency, cost, and compounding errors.

part ofPrompt engineeringoverview, primer and where to startread it →
on this pageshow

questions

5

What is prompt chaining, and why split a task across several model calls?

level: juniorimportance: must knowfreq 62%

answer

  1. ordered calls, not one big instruction
  2. output of one stage becomes next prompt
  3. the seam is where you can assert
  4. controllability bought with latency
  5. errors from stage one travel forward

basics

~20 s

Prompt chaining runs one task as an ordered series of model calls, each stage's output feeding the next stage's prompt. Every stage gets a narrow instruction and a checkable result, so a single step can be validated, retried, or replaced.

solid answer

~50 s

Prompt chaining decomposes a request into ordered sub-prompts: stage one produces an intermediate artifact, that artifact is inserted into stage two's prompt, and so on until the final answer. The point is controllability. A single mega-prompt that asks the model to extract, judge, and write in one breath gives you exactly one output to inspect, and when it is wrong you cannot tell which instruction it dropped. Splitting the work gives each stage a narrow job and a machine-checkable output shape, so you can validate between stages, retry only the failing one, send mechanical stages to a cheaper model, and log per-stage inputs and outputs when debugging. The price is real: end-to-end latency becomes the sum of the stages, context is re-sent on every call, and an early stage's mistake propagates silently into later ones. Chain where you need a checkpoint, not everywhere.

go deeper

for a junior

Be able to say plainly that a chain is several calls where one call's output goes into the next call's prompt, and give one reason to do it, such as being able to check the middle result.

for a middle

Explain what a stage boundary buys concretely: validation, targeted retry, cheaper models on mechanical stages, per-stage logs. Then name the costs, including additive latency and repeated context.

for a senior

Show judgement about where to put the seam. Argue from a measurement of the single-call baseline rather than a preference, and talk about how you would debug a wrong final answer using per-stage traces.

for a principal

Own the standing question of whether the pipeline is still the right shape as models improve, and the maintenance cost of every extra stage: its tests, its failure handling, and the eval that justifies keeping it.

## The mechanic Prompt chaining means running one logical task as a sequence of separate model calls. Stage 1 receives the raw input and produces an intermediate artifact: a list, a JSON object, an outline, a draft. Your application code takes that artifact, optionally checks or transforms it, and embeds it in the prompt for stage 2. The chain continues until a stage produces the answer the user asked for. The key word is *separate*. Asking a model to think step by step inside one call is a different technique: there, the intermediate steps are tokens in a single response and your code never touches them. In a chain, the seam between steps is real program state. That is exactly what makes chaining an engineering tool rather than a prompting trick, because anything your code can hold, it can inspect. ## A worked shape Suppose you turn supplier-contract text into a risk summary for a procurement team. As one call, the prompt says: find the termination, liability, and indemnity clauses; classify each against our policy; and write a one-paragraph summary. As a chain it becomes three calls. Stage 1 extracts the relevant clauses verbatim with their section numbers. Stage 2 receives only those clauses plus the policy rules and emits a structured verdict per clause. Stage 3 receives the verdicts and writes the paragraph. Each stage's instruction now fits on a few lines, and each stage's output has a shape you can assert on. ## What a stage boundary actually buys **Validation.** Between calls you can parse JSON, check that a required field is present, or confirm an extracted quote appears in the source. A malformed stage fails loudly instead of being smoothed over by the next instruction. **Targeted retry.** If stage 2 returns an unparseable verdict, you re-run stage 2 with the error appended, not the whole pipeline. That is cheaper and it keeps the good work already done. **Model and effort routing.** Mechanical stages such as extraction or reformatting often run acceptably on a small, fast model, while the judgement stage needs the strong one. A single call forces one model for everything. **Observability.** Per-stage logs make failures locatable. When the summary is wrong you can ask whether extraction missed the clause, whether the verdict was wrong, or whether the writer distorted a correct verdict. With one call you get a single blob and a guess. **Focused attention.** Long multi-part instructions are a common source of dropped requirements. Each stage seeing only what it needs is usually the cheapest fix for output that quietly ignores half the brief. ## What it costs Latency is additive. Three calls that each take two seconds give a six-second answer, and there is no partial credit if the user is watching a blank screen. Token cost usually rises too, because the shared framing and any source document get re-sent per stage, though a stable prompt prefix can often be cached by the provider so the repeat is much cheaper than the first send. Errors compound. If each stage is independently right ninety-five percent of the time, four chained stages are right about eighty-one percent of the time, and the failures are the nastiest kind: an early stage returns something plausible but wrong, later stages accept it without complaint, and the final answer is confidently incorrect with no error anywhere in the logs. There is also information loss at the seam. Whatever stage 1 chose not to emit is gone. If stage 3 would have benefited from a nuance in the original document, and stage 1 summarised it away, no amount of prompt tuning downstream recovers it. A common mitigation is to pass the source alongside the intermediate artifact rather than replacing it. Finally, a chain is code. It needs error handling, timeouts, tests, and a deployment story per stage. That maintenance cost is invisible in a demo and very visible in year two. ## When one call is enough Use one call when the task genuinely is one instruction, when the intermediate result has no independent use, when latency is the dominant requirement, or when the model already performs the whole task reliably on your eval set. Capable models absorb more per call every year, so a pipeline written when models were weaker is a standing candidate for consolidation. The honest test is not aesthetic preference but measurement: run both variants against a labelled set and compare quality, p95 latency, and cost. ## Choosing the seam Good seams are places where something checkable exists: a structured artifact, a deterministic tool call, a point where a cheaper model suffices, or a decision that must be recorded. Bad seams split what is really one judgement into halves that each see too little context. If you cannot say what you would assert at a boundary, that boundary is probably paying latency for nothing.

  • You added a chain and quality improved, but p95 latency tripled. What do you try before reverting?
    First check whether any stages are independent and can run concurrently instead of in series. Then look at whether a small fast model is adequate for the mechanical stages, whether a stable prompt prefix can be cached so repeated context is cheap, and whether an early stage's output can be streamed or shown as progress. Only after that consider merging two adjacent stages whose boundary carries no validation.
  • How do you decide what an intermediate stage should output?
    Pick the smallest artifact the next stage genuinely needs, in a shape your code can validate: structured fields rather than prose, with identifiers that point back at the source. If you cannot write an assertion over it, the boundary is not earning its latency. Where later stages may need nuance the artifact drops, pass the original source alongside it rather than replacing it.

saying these in an interview costs you the question

  • Claims chaining removes hallucination because each stage is smaller
  • Thinks chaining and step-by-step reasoning in one call are the same thing
  • Ignores that latency and token cost add up across stages
  • Assumes a later stage will notice an earlier stage's mistake
  • Chains every task by default without measuring the single-call baseline

context

open as a page

A four-stage screening chain extracts fields then judges inclusion; how do you stop stage-2 errors poisoning stage 3?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Put a check at the seam. Validate the extracted fields against a schema, require each value to appear verbatim in the source, allow the stage to abstain rather than guess, and pass the source forward so the judging stage can disagree with the extraction.

open as a page

What is least-to-most prompting, and how does it differ from a fixed pipeline?

level: middleimportance: should knowfreq 45%

basics

~20 s

Least-to-most prompting asks the model to first list the subproblems a hard question contains, then solve them in dependency order, each answer feeding the next. The decomposition is produced at run time by the model, not authored in advance by an engineer.

open as a page

How would you decide whether to collapse a five-call prompt chain into one call?

level: principalimportance: should knowfreq 38%

basics

~20 s

Treat it as an experiment, not a preference. Run both variants over a labelled set, compare quality, p95 latency, and cost per request including retries, then price the observability and routing you give up. Collapse adjacent stages whose boundary carries no check.

open as a page

What is skeleton-of-thought prompting, and when does parallel expansion hurt quality?

level: middleimportance: nice to knowfreq 26%

basics

~20 s

Skeleton-of-thought first asks for a terse outline, then expands each outline point in its own call, running the expansions concurrently and stitching the results. It cuts wall-clock time for long outputs but breaks down when the sections depend on each other.

open as a page