skip to content

Ordering Effects & Label Bias

The uncomfortable finding that permuting the same examples can swing accuracy by double digits: recency bias toward the last demonstration, majority-label bias, and calibration fixes. Asked to see whether you have ever actually measured prompt sensitivity instead of assuming it away.

part ofPrompt engineeringoverview, primer and where to startread it →
on this pageshow

questions

4

Why can permuting the same four few-shot examples swing accuracy by double digits?

level: middleimportance: must knowfreq 62%

answer

  1. same examples, different score
  2. position carries information about labels
  3. the last shot pulls hardest
  4. report spread across permutations, not one run

basics

~20 s

Order is not cosmetic: the model reads the demonstration sequence as evidence about which label is likely, and examples nearest the end pull hardest. The same four exemplars in different orders can move a sentiment task from the fifties into the high eighties.

solid answer

~50 s

A demonstration block does two things at once: it shows the task, and it implicitly sets a prior over the answer space. Position changes that prior. Predictions skew toward the labels of the demonstrations closest to the query (recency bias), toward whichever label appears most often, and toward label words that are simply frequent in ordinary text. With four exemplars there are twenty-four orders, and on a binary sentiment task the spread across them can run from roughly chance into the high eighties with the *identical* examples and wording. The practical consequence is about measurement, not just tuning: a single-order accuracy number is one draw from a distribution. Report the mean and the spread over several permutations, otherwise any other prompt change you evaluate is measured against noise larger than the effect you are chasing.

code

python · 14 lines
python
from itertools import permutations
import statistics

shots = [("great film", "positive"), ("waste of time", "negative"),
         ("loved it", "positive"), ("boring", "negative")]

def evaluate(order):
    # stand-in for a real evaluation run over a fixed test set
    tail_bonus = 0.12 if order[-1][1] == "negative" else -0.05
    return 0.70 + tail_bonus + 0.02 * sum(i for i, s in enumerate(order) if s[1] == "positive")

scores = [evaluate(p) for p in permutations(shots)]
print(len(scores), round(min(scores), 3), round(max(scores), 3),
      round(statistics.pstdev(scores), 3))

go deeper

for a junior

Know that few-shot examples are part of the prompt the model conditions on, and that changing their order can change the answer. Say plainly that you would test more than one order before trusting a score.

for a middle

Be ready to name the mechanisms — recency bias, majority-label bias, common-token bias — and explain that the demonstration block sets an implicit prior over labels, not just an illustration of the mapping.

for a senior

Show the measurement discipline: fix the data, sweep permutations, report mean and spread, and refuse to compare prompt variants whose gap is smaller than the permutation standard deviation.

for a principal

Own the tradeoff between tuning around the sensitivity and engineering it away — more shots, calibration, order ensembling or fine-tuning — and set the policy for re-measuring variance after every model upgrade.

## The phenomenon Few-shot prompting means putting solved examples (demonstrations, exemplars, or "shots") in the prompt before the real input. Intuitively the examples are just an illustration, so their sequence should not matter much. Empirically it matters a great deal. Take four labelled sentiment examples, hold the wording, the template and the test set fixed, and evaluate all twenty-four orderings: the accuracy distribution is wide, with the worst orders near chance and the best ones far above. Double-digit swings on classification tasks are routine, and this was one of the first uncomfortable findings of the prompting literature — the prompt you happened to type is a sample from a distribution of prompts, not a fixed artifact. ## Why position carries information A language model predicts the next token conditioned on everything before it. The demonstration block therefore acts as evidence about the *joint* distribution of inputs and labels, not only about the mapping. Three distinct biases interact: - **Recency bias.** Tokens closer to the point of prediction exert more influence. If the last demonstration is labelled `negative`, the model's probability mass shifts toward `negative` for the query. Primacy effects exist too, but the tail of the prompt usually dominates in short shot blocks. - **Majority-label bias.** If three of four demonstrations carry the same label, the model infers that this label is common and predicts it more often than the true base rate warrants. - **Common-token bias.** Label words that are frequent in ordinary text are easier for the model to emit than rare ones, so the answer space is unevenly reachable before the task even starts. Ordering interacts with the other two: a majority label placed last compounds the skew, while placing the minority label last partly cancels it. That interaction is why order effects look erratic rather than monotone. ## Does a bigger or better model fix it? Not reliably. Sensitivity does not fall smoothly with scale, and instruction-tuned chat models are calmer but not immune — the effect is strongest when the shot count is small, the labels are imbalanced, the task is ambiguous, or the query sits near a decision boundary. Treat "we use a frontier model, so ordering does not matter" as a claim to be measured rather than assumed, and re-measure after any model upgrade. ## Measuring it instead of assuming it away The discipline is simple and rarely practised: 1. Fix the exemplar set and the template. 2. Sample several permutations — all of them if the shot count is small, a random sample if not. 3. Evaluate each on the same dataset. 4. Report mean, min, max and standard deviation, not a single number. The spread is the quantity of interest. If permutation standard deviation is four points and your new prompt idea gains two points, you have measured nothing. Teams that skip this step routinely ship "improvements" that are re-rolls of the same dice, and equally often reject good ideas that happened to land on a bad order. ## Reducing the sensitivity Several levers shrink the variance rather than gaming it: - **More shots.** Variance typically falls as the demonstration count rises, because no single position dominates. This trades tokens and latency for stability. - **Calibration.** Estimate the model's label prior and divide it out of the output probabilities, which absorbs a large part of the majority-label and common-token component. - **Ensembling over orders.** Run k permutations and take a majority vote. Costs k× tokens, but converts a variance problem into a slightly slower, much steadier system. - **Choosing an order deliberately.** Search a set of candidate orders on held-out data instead of accepting whichever one you typed first, and validate that the winner is real rather than selection noise. - **Fine-tuning.** When the task is high-volume and stable, training the behaviour into weights removes the prompt-permutation surface entirely — at the cost of a training and refresh pipeline. ## The engineering takeaway Order sensitivity is a property of the *system*, not a bug in the model. The mature response is to (a) quantify it once so you know the noise floor of your evaluation, (b) drive it down structurally where the task allows, and (c) re-check it whenever the model, the shot count or the label mix changes. The failure mode to avoid is a prompt that scores beautifully in a notebook because it was, silently, the lucky permutation.

  • If you can only afford one extra evaluation run, what would you spend it on?
    A second, randomly chosen permutation of the same exemplars. One extra order gives you the cheapest possible signal about whether your accuracy number is stable or a lucky draw. If the two runs differ by more than the effect size you care about, every subsequent prompt comparison you make is untrustworthy until you widen the sample.
  • Does raising the shot count from four to sixteen eliminate order effects?
    It usually shrinks them but does not eliminate them. With more demonstrations, no single position dominates the prior, so permutation variance falls. But recency still favours the tail, imbalanced labels still skew the prior, and you now pay more tokens and latency per call. Measure the variance at your chosen shot count rather than assuming a threshold makes it vanish.
  • How would you tell an ordering effect apart from plain evaluation noise?
    Hold the evaluation set fixed and vary only the order — then any difference cannot come from sampling different test items. Run each order more than once at your production temperature to separate decoding noise from ordering effect. If the ranking of orders is stable across repeats on the same data, the effect is real.

saying these in an interview costs you the question

  • Claiming example order is purely cosmetic presentation
  • Assuming frontier models are order-invariant without measuring
  • Reporting one accuracy number from one arbitrary permutation
  • Confusing ordering effects with sampling temperature noise
  • Believing more examples always removes the effect entirely

context

open as a page

How does contextual calibration use a content-free input like "N/A" to correct label bias?

level: middleimportance: should knowfreq 30%

basics

~20 s

Feed the prompt an input carrying no task evidence, such as "N/A" or an empty string, and read the label probabilities it returns. Those probabilities are the prompt's built-in prior; dividing real predictions by them and renormalising removes most of the bias.

open as a page

Why does a three-approve, one-deny shot set push a few-shot classifier toward approve?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Majority-label bias: the model reads the label mix in the demonstrations as evidence about how often each outcome occurs, so a 3:1 approve-heavy shot set inflates the predicted approve rate on borderline building-permit applications, independent of what each application actually says.

open as a page