skip to content

What does in-context learning change in an LLM if no weights are updated?

level: juniorimportance: must knowfreq 76%

answer

  1. ask where the state actually lives
  2. no training run happens here
  3. parameters stay frozen
  4. conditioning the next-token distribution
  5. examples re-sent and re-billed each call

basics

~20 s

In-context learning changes only the next-token probabilities for that one request. Examples and instructions in the prompt condition the output distribution during the forward pass; the parameters stay frozen, so nothing carries over to the next call.

solid answer

~40 s

A language model maps a token sequence to a distribution over the next token. Everything you send — instructions, worked examples, the real input — is one sequence, and attention lets the position being generated read all of it, so the demonstrations shift the probabilities the model emits. That is the whole mechanism: no optimizer runs, no gradient is computed, and the parameters are identical before and after the call. Two consequences follow. First, the effect is scoped to the request: a correction the model "accepted" in one conversation is gone in the next unless your system re-sends it. Second, the examples are re-tokenized and re-processed on every call, so they are a permanent per-request cost in tokens, latency and context budget rather than a one-off training cost.

go deeper

for a junior

Be ready to say plainly that prompt examples never change the model's weights. They shape the answer to that one request only, and your application has to resend them every time.

for a middle

Explain the mechanism: demonstrations are part of the input the forward pass conditions on, shifting next-token probabilities. Then price it — those tokens are re-processed on every call and compete for the context window.

for a senior

Show that you design around statelessness: the prompt is versioned and logged like code, a behaviour fix has to be rolled out to every call site, and example blocks are budgeted against the real input at production volume.

for a principal

Own the framing that in-context learning converts a one-off training budget into a permanent per-request context budget, and make that recurring cost visible in capacity, latency and pricing plans rather than treating prompts as free.

## The claim in one line In-context learning (ICL) is the ability of a large language model to perform a task it was never specifically trained for, using only the instructions and worked examples placed inside its prompt. Nothing about the model changes. No optimizer step runs, no gradient is computed, and the parameters are identical before and after the call. The word "learning" is historical and slightly misleading; **conditioning** is the accurate description. ## What the model actually does with your examples A language model is a function from a sequence of tokens to a probability distribution over the next token. Everything in a request — the system instructions, the four labelled example claims you pasted, the new claim you want triaged — is one flat token sequence. During the forward pass, attention lets the position where the answer will be generated read every earlier position, so the content of the demonstrations shows up in the hidden state at that position, and therefore in the probabilities that come out. A token is sampled from that distribution, appended, and the process repeats. When the response ends, that working state is discarded. Providers often cache the internal state of a repeated prompt prefix so the next call with the same prefix is cheaper or faster. That is an infrastructure optimisation of the same computation — the cached state is derived entirely from tokens you sent — not a form of learning. The weights are still untouched. ## Why being pedantic about this pays off Three practical facts follow directly: 1. **Nothing carries over.** If an insurance claim-triage prompt produces the right label after you add a worked example, that improvement exists only in requests that carry the example. It is a property of your template, not of the model. 2. **Nothing leaks sideways.** Your examples do not change behaviour for other users or other endpoints of the same model. Statelessness is a safety property as much as a limitation. 3. **The prompt is the entire configuration.** To explain why a production answer came out the way it did, you need the exact text that was sent. Log it, version it like code, and treat a prompt edit as a deploy. ## What conditioning buys you Speed of iteration is the whole point. A claim-triage prompt that shows four adjudicated claims labelled `total_loss`, `repairable` and `fraud_review` fixes both the label vocabulary and the output shape immediately: the model now knows which strings are legal outputs and what a well-formed answer looks like. Changing that behaviour is a text edit, testable in seconds, revertible instantly, and variable per request — you can send a different example set to a different segment on the same model, which no weight-level change can do cheaply. ## What it costs The cost is recurring and structural: - **Tokens and latency.** Sixteen demonstrations are re-processed on every single call. At a million calls a day that is a permanent line item, whereas a training cost is paid once. - **Context budget.** Examples compete for the same window as the actual input, retrieved documents and conversation history. Long example blocks crowd out the material the answer actually depends on. - **Fragility.** Because the effect is purely conditioning, it is sensitive to how the examples are written, formatted and ordered. There is no artifact you can inspect to know what the model "knows"; there is only the prompt you sent. - **No sideways generalisation.** Fixing a failure in one prompt does not fix the same failure in another call site that happens to use the same model. ## The contrast to name in an interview Methods that update parameters produce a persistent change: the behaviour is baked in, per-request prompts get shorter, and the cost moves to a one-off training run plus the operational weight of owning a model version. In-context learning produces no persistent change and moves the cost into every request forever. Which one a given system should use is a separate decision with its own tradeoffs; the point here is that they are different *kinds* of change, and confusing them is the classic junior error. ## Common confusions worth pre-empting - **"The model learned my writing style."** No — the template that carries your style samples did. Delete the samples and the style disappears. - **"The examples fine-tune the model a little."** There is no gradient anywhere in a normal inference request. - **"Same prompt, same answer."** Conditioning reshapes a distribution; sampling still draws from it, so identical requests can differ unless randomness is constrained. - **"The model remembers what I told it earlier."** Only because your client re-sends the conversation. Statelessness is the default; any memory is something your system built around the model.

  • If the examples work so well, why does per-request cost grow with them?
    Every demonstration is part of the input on every call, so it is re-tokenized and re-processed each time and billed as input. A sixteen-example prompt pays that tax on request number one and request number ten million alike. Prefix caching can reduce the price of a repeated block, but the tokens are still part of each request and still occupy the context window.
  • Two identical few-shot requests returned different answers. Does that contradict the idea of conditioning?
    No. Conditioning reshapes the probability distribution over next tokens; a sampler then draws from it. Unless randomness is constrained, two draws from the same distribution can differ. Some deployments expose sampling controls and some deliberately lock them, but in every case conditioning narrows the distribution rather than pinning a single output.
  • How do you make a behaviour that in-context learning produces survive across sessions?
    Store it outside the model and re-insert it. The instruction or example block lives in your prompt template, a config store or a retrieved record, and your application puts it back into every relevant request. The alternative is changing the weights, which is a different kind of investment. There is no third option where the model quietly keeps it for you.

It is closer to showing a temp worker three completed forms before they fill in the fourth than to sending them on a training course: the moment the desk is handed back, nothing has been retained.

saying these in an interview costs you the question

  • Says the prompt examples retrain or fine-tune the model
  • Claims few-shot examples change answers for other users too
  • Describes in-context learning as gradient descent inside the API
  • Assumes examples are free because no training run happens
  • Thinks the model remembers a correction after the conversation ends

context