skip to content

What is skeleton-of-thought prompting, and when does parallel expansion hurt quality?

level: middleimportance: nice to knowfreq 26%

answer

  1. outline first, sections after
  2. branches run at the same time
  3. wall clock beats token count
  4. independence is the load-bearing assumption
  5. stapled memos, not one argument

basics

~20 s

Skeleton-of-thought first asks for a terse outline, then expands each outline point in its own call, running the expansions concurrently and stitching the results. It cuts wall-clock time for long outputs but breaks down when the sections depend on each other.

solid answer

~50 s

Skeleton-of-thought is a two-phase pattern for long-form generation. One call produces a skeleton, a short numbered outline of the points the answer will cover. A second wave of calls expands each point independently, and because those calls do not depend on each other they run concurrently, so end-to-end time approaches the slowest section rather than the sum of all of them. For a twelve-section market report that is a large wall-clock win, and the outline also gives you structural control and a cheap review point before any expensive expansion happens. It costs more tokens, since every branch re-sends the shared framing, and it fails where sections are genuinely interdependent: a stepwise proof, a narrative, or code whose later parts use earlier definitions. Passing the full skeleton and a scope budget into each branch, then adding a cheap final consistency pass, limits the damage.

go deeper

for a junior

Recall the two phases: a first call writes a short outline, then separate calls expand each outline point, and the pieces are joined back together in order.

for a middle

Explain that the win is wall-clock latency because the expansions run concurrently, that token cost goes up rather than down, and that the pattern assumes sections are independent.

for a senior

Diagnose the assembled-output failures you would actually see: repeated background, divergent assumptions, missing transitions, runaway length, and say which mitigation addresses each.

for a principal

Decide when the complexity is justified at all, weighing it against simply streaming one response, and set the policy for bounded concurrency and cost when many requests fan out at once.

## The pattern Skeleton-of-thought splits long-form generation into an outline phase and an expansion phase. The first call sees the request and returns a skeleton: a short numbered list of section headings or one-line point summaries, deliberately terse. The application then issues one call per skeleton point, each asking for that point expanded to full prose, and concatenates the results in skeleton order. The distinguishing feature is that the expansion calls are independent of one another, so they are issued concurrently rather than in series. That is the opposite of a normal prompt chain, where stage two waits on stage one. ## Why it helps **Wall-clock latency.** Generation is sequential per response: a long answer takes long mainly because tokens come out one after another. Splitting a twelve-section report into twelve concurrent expansions means total time is roughly the outline call plus the slowest section, not the sum of twelve sections. On a long report this is the difference between a minute and a handful of seconds, and it is the pattern's original motivation. **Structural control.** You get to see and, if you like, edit or validate the outline before paying for any expansion. Wrong shape, missing section, or a heading outside the brief can be caught for the cost of one small call. That review point is worth having independently of the latency win. **Length control per section.** A single long generation drifts in section length, front-loading detail and trailing off. Per-section calls with an explicit word budget produce more even coverage. ## What it costs Token cost rises. Each branch re-sends the shared framing, the skeleton, and any source material, so a twelve-way expansion pays that overhead twelve times. A stable prefix can often be cached by the provider, which softens the bill considerably, but the request count and any per-request overhead remain. Concurrency also pushes you against provider rate limits far faster than a single call does, so a fan-out needs bounded parallelism. ## When parallel expansion degrades the output The pattern assumes the sections are independent. Where that assumption fails, the damage shows up in predictable ways. **Interdependent content.** A stepwise mathematical proof, a tutorial where step four uses the file created in step two, or code where later functions call earlier ones cannot be expanded in isolation. Each branch has to guess what its predecessor established, and the guesses do not agree. **Repetition.** Independent branches all reach for the same background. Every section opens by re-explaining the same premise, because no branch knows another already covered it. The concatenated result reads padded even though each section alone reads fine. **Inconsistent commitments.** Branches silently pick different assumptions, terminology, units, or examples. One section says the sample is regional, another treats it as national. Nothing is internally contradictory within a section; the contradiction only exists in the assembled document. **No connective tissue.** Sections cannot reference each other, so transitions are absent or generic, and the whole reads like a stapled set of memos rather than one argument. **Uncontrolled total length.** Each branch decides its own depth, so the assembled output can be several times the length anyone wanted. ## Mitigations Pass the whole skeleton into every expansion call, not just that branch's point. A branch that knows what sections three and five will cover stops re-explaining their material and can defer to them. Add an explicit scope and word budget per section, and state which section owns the shared background so exactly one branch writes it. Where a genuine dependency exists between two points, do not fight it: chain those two sequentially and parallelise only the rest, which is usually most of them. Finally, add a cheap assembly pass over the concatenated draft that removes duplicated background, harmonises terminology, and writes transitions. It is one extra call over material that already exists, it runs on a small model, and it recovers most of what the fan-out lost. That pass is what turns the pattern from a demo into something you would ship. ## When not to reach for it If the answer is short, the outline overhead dominates and there is nothing to parallelise. If the content is inherently sequential, expand in order and accept the latency, or stream the single response so the user sees progress instead. And if perceived latency is the real problem rather than total latency, streaming a single generation is simpler, cheaper, and usually enough.

  • How do you keep concurrent expansions from repeating the same background material?
    Give every branch the full skeleton, not only its own point, and name in the prompt which section owns the shared background so the others can assume it and move on. Set a scope sentence and a word budget per section. Then run a cheap assembly pass over the concatenated draft to strip surviving duplication and add transitions, which is far more reliable than hoping each branch stays in its lane.
  • When is streaming a single response the better answer than this pattern?
    When the complaint is that the user waits with nothing on screen rather than that the total generation is slow. Streaming shows tokens immediately, keeps the output internally consistent because it is one generation, costs no extra tokens, and adds no orchestration code. Fan-out earns its complexity when total wall-clock matters for a downstream consumer, or when you want to inspect and validate the outline before paying for expansion.

saying these in an interview costs you the question

  • Says it reduces token cost when it usually increases it
  • Applies it to content whose sections depend on each other
  • Gives each branch only its own outline point and expects coherence
  • Confuses it with step-by-step reasoning inside a single response
  • Ignores rate limits when fanning out many concurrent calls

context