skip to content

What are induction heads, and how do they relate to in-context learning?

level: seniorimportance: nice to knowfreq 20%

answer

  1. a circuit, not a single head
  2. two heads working together across layers
  3. match the earlier occurrence, copy its successor
  4. forms abruptly, visible in the loss curve
  5. explains copying better than abstraction

basics

~20 s

An induction head is an attention circuit that spots an earlier occurrence of the current token and copies whatever followed it. Such circuits form during pretraining and their appearance coincides with a jump in the model's ability to exploit patterns in its prompt.

solid answer

~50 s

Induction heads come from mechanistic interpretability. The circuit spans two layers: a previous-token head writes, at each position, information about the token before it; a second head then attends from the current token back to positions that matched it earlier and copies the successor. The behavioural signature is completing `... A X ... B Y ... A ->` with `X`. What makes them interesting is timing: in the models where this has been studied closely, these circuits form abruptly early in training, and their formation lines up with a bump in the loss curve and with a sudden improvement in how much the model benefits from tokens already in its context. That is the strongest mechanistic story available for why prompt patterns get continued at all. Be careful with the claim, though — the evidence is largely correlational and concentrated on small models, and it explains copying and pattern continuation far better than it explains few-shot performance on abstract tasks.

code

python · 8 lines
python
prompt = "\n".join([
    "Request log:",
    "req-7f2a  ->  ACCEPTED",
    "req-91bd  ->  REJECTED",
    "req-7f2a  ->",
])
# The completion 'ACCEPTED' is produced by matching the earlier
# occurrence of req-7f2a and copying the token that followed it.

go deeper

for a junior

Recognise the term and the behaviour it names: the model completes a pattern by finding where the current token appeared earlier and repeating what came after it.

for a middle

Describe the two-layer composition — a previous-token head feeding a head that attends back to matching positions and copies the successor — and give the AX / BY / AX example.

for a senior

Connect the mechanism to observed training dynamics and to prompting practice: circuits forming abruptly alongside a context-use jump, and the prediction that any repeated structure in your examples, including its defects, gets continued.

for a principal

Be candid about epistemic status — largely correlational evidence, mostly small models, and no settled account of abstract in-context learning — so the team treats interpretability results as useful intuition rather than as guarantees to design against.

## The question this tries to answer In-context learning is easy to describe behaviourally — put a pattern in the prompt and the model continues it — and hard to explain mechanically. Induction heads are the best-known partial mechanical explanation, and they are the reason the phrase appears in senior interviews about LLM capabilities. ## The circuit An induction head is not a single component but a composition of two attention heads in different layers: 1. **A previous-token head** in an earlier layer attends from each position to the one immediately before it, writing information about the preceding token into the current position's representation. 2. **The induction head** in a later layer attends *from* the current token *to* positions whose representation says "the token before me was this same token", and copies that position's token into the output prediction. The composite behaviour is: find where this token appeared before, look at what came next, predict that. Given `... A X ... B Y ... A`, the model predicts `X`. The two heads communicate through the residual stream — the first writes a signal the second reads — which is why it is described as a circuit rather than a head. ## Why it is linked to in-context learning Two observations connect the circuit to the capability. **Timing.** In the models where this has been examined closely, induction heads do not appear gradually. They form in a narrow window early in training, and that window coincides with a visible bump or kink in the training loss curve — a moment where the model gets abruptly better at something. **Correlation with a context-use score.** In-context learning ability can be operationalised without any benchmark: measure how much lower the model's loss is on a token late in the context than on a token early in the context. If the model exploits what it has already seen, later tokens are easier. That score jumps at the same point in training that the induction heads form. The natural reading is that the circuit is a substantial part of what makes context useful at all. ## What the story does and does not explain It explains a great deal of everyday prompting behaviour: - **Format continuation.** If your prompt shows three blocks of `Claim: ... / Label: ...`, the model continues the fourth block in the same shape. Literal structural copying is exactly what this circuit does. - **Quirk propagation.** Copying is indiscriminate. A stray trailing space, an inconsistent quote style or a typo repeated in two examples gets continued into the output. Practitioners who have wondered why an accidental artifact in an example block keeps reappearing are watching pattern continuation at work. - **Repeated-identifier completion.** Prompts full of repeated keys, ids or tags get their associations completed with striking reliability. It explains much less about the harder case: performing a genuinely abstract task from a handful of examples, where the correct answer is nowhere in the prompt to be copied. Copying cannot produce a label the model has not been shown in that exact context. Later interpretability work explores more abstract mechanisms — for instance the idea that processing the demonstrations produces a compact internal representation of the task that is then applied to the query — but this is an active area, not a settled account. ## How to state the claim honestly Three caveats belong in any senior answer: - **The evidence is correlational.** Circuit formation coinciding with a capability jump is strong evidence of involvement, not proof that the circuit is the cause of all in-context learning. - **The models studied are small.** Much of the clearest work is on small attention-only transformers and small production-class models. Extrapolating a mechanism to a frontier-scale model with a different architecture is an assumption, not a finding. - **It is a component, not a knob.** Nobody configures an induction head. It is a description of structure discovered in trained weights, useful for intuition and for interpretability research, not something exposed to an application developer. ## What to do with it as a practitioner The practical payoff is a mental model for prompt structure. If part of what makes examples work is literal pattern continuation, then keeping the demonstration format byte-identical to the output you want is not a stylistic nicety — it is directly exploiting the mechanism. It also predicts a failure mode: whatever is repeated in your examples will be continued, including things you did not intend to teach. Reviewing example blocks for accidental consistency is a cheap and surprisingly effective debugging step.

  • Does this account explain few-shot learning on an abstract classification task?
    Only partly. Copying the successor of an earlier match explains structural continuation and repeated-association completion, but not producing a label that never appears in a copyable position. Interpretability research explores more abstract mechanisms — such as a compact internal task representation formed while reading the demonstrations and then applied to the query — but no complete account of abstract in-context learning is settled.
  • What concrete prompting behaviour does the induction-head view predict?
    That anything structurally repeated across your examples will be continued, wanted or not. Keeping the demonstration format byte-identical to the output you intend to parse exploits the mechanism deliberately. The flip side is that a stray trailing space, an inconsistent quoting style or a typo repeated in two examples will reliably reappear in the model's output.
  • How is in-context learning ability measured without a benchmark in this line of work?
    By comparing the model's loss on a token late in the context with its loss on a token early in the context. If the model exploits what it has already read, later tokens are cheaper to predict, and the gap is a continuous score. This score rises sharply at the same point in training where induction circuits form, which is what links the two.

saying these in an interview costs you the question

  • Describes induction heads as something a developer configures
  • Claims they fully explain few-shot learning on abstract tasks
  • Confuses the circuit with a separate memory module or cache
  • Says every attention head performs copying
  • Treats findings from small models as proven for frontier models

context