How can nearest-neighbour exemplar retrieval hurt accuracy on a borderline query?
answer
- nearest is not the same as representative
- similarity and label are correlated
- borderline queries get a one-sided vote
- duplicates give k shots, one shot of signal
- floor plus fallback, cap per label
basics
~20 sNearest neighbours are similar to the query, not representative of the task. For a borderline case they often all carry one label, or duplicate each other, so the prompt quietly argues for that one answer instead of showing the model where the decision boundary sits.
solid answer
~50 sSimilarity search optimises for closeness, and closeness correlates with the label. So the retrieved set for a genuinely ambiguous input tends to collapse into one class — the model sees four demonstrations all answering "reject" and follows the majority rather than reasoning about the boundary. Two related failures come from the same mechanism: a pool with near-duplicate entries returns k demonstrations that carry one example's worth of information, and a genuinely novel input has no close neighbour at all, so retrieval hands back distant, off-task examples that are worse than a curated fixed block. The guards are all at retrieval time: enforce a similarity floor and fall back to a static default block below it, cap how many retrieved exemplars may share one label so both sides of the boundary appear, and de-duplicate near-identical entries. Log the retrieved ids, scores and label mix so you can see the collapse when it happens.
code
python · 20 linesdef select_exemplars(query_vec, pool, k=4, floor=0.55, per_label_cap=2, fallback=()):
def cosine(a, b):
dot = sum(x * y for x, y in zip(a, b))
na = sum(x * x for x in a) ** 0.5
nb = sum(x * x for x in b) ** 0.5
return dot / (na * nb) if na and nb else 0.0
ranked = sorted(pool, key=lambda e: -cosine(query_vec, e["vec"]))
kept, per_label = [], {}
for entry in ranked:
if cosine(query_vec, entry["vec"]) < floor:
break
label = entry["label"]
if per_label.get(label, 0) >= per_label_cap:
continue
kept.append(entry)
per_label[label] = per_label.get(label, 0) + 1
if len(kept) == k:
break
return kept if len(kept) >= 2 else list(fallback)go deeper
Know that retrieval returns the closest examples it has, not the most useful ones, and that it never says "I found nothing good" unless you make it. Be able to name single-label neighbourhoods as a real risk.
Explain why similarity correlates with the label and how that makes borderline queries receive one-sided demonstrations. Name concrete retrieval-time guards — a similarity floor with a static fallback, a per-label cap, de-duplication — and what each one costs.
Show you would make the failure observable before it becomes an incident: log retrieved ids, scores and label mix, alert on single-label rate and sub-floor rate, and slice evaluation by those signals rather than trusting an aggregate score.
Own the trade between similarity and boundary coverage as a tunable product decision, and argue when the honest answer is to stop retrieving demonstrations for a segment and route it to a different treatment entirely.
## The mechanism behind the failure A nearest-neighbour retriever answers one question: which pool entries are most similar to this input? On most tasks similarity correlates strongly with the label — cases that look alike were usually decided alike. That correlation is exactly why retrieval works, and it is also exactly why it fails on the inputs that matter most. The hard queries are the ones near a decision boundary. For those, the nearest neighbourhood is drawn from whichever side of the boundary happens to be denser in the pool. The prompt then shows the model several demonstrations that all answer the same way. Few-shot prompting is imitation, so the model does what the demonstrations do. You have not shown it the boundary; you have shown it a lopsided vote. This is worth separating from ordering and majority-label effects in a *fixed* prompt. There, the imbalance is a property of a set you chose once and can inspect. Here it is generated fresh per request by the retriever, is different for every query, and never appears in code review — which is why it is diagnosed from logs rather than from reading the template. ## The three neighbourhood shapes that go wrong **Single-class collapse.** All k retrieved exemplars share one label. The model reads a unanimous block and answers with the majority. This hits borderline and minority-class queries hardest, so the damage concentrates precisely where accuracy already suffers. **Near-duplicate collapse.** Pools grown by write-back accumulate variants of the same case — the same incident reported five times, the same template ticket. Similarity search happily returns all five. Nominally the model gets 5-shot; informationally it gets 1-shot, with four tokens' worth of redundancy occupying the context and reinforcing one pattern. **Empty or distant neighbourhood.** A genuinely new kind of input has nothing close in the pool. The retriever does not refuse — it returns the least-far entries it has, which may be barely related. Distant demonstrations are not neutral: they actively pull the model toward the wrong output shape or the wrong label vocabulary. A curated static block would have done better, because at least its examples were chosen to be on-task. A fourth case is worth naming because it is silent in the other direction: **the query itself is in the pool**. If production inputs are written back and the same input arrives again, retrieval hands the model the answer verbatim. Live, that is often harmless or even desirable. In offline evaluation it is contamination — your measured accuracy is a lookup, and it will not survive contact with unseen inputs. ## Guards that live at retrieval time **Similarity floor with a static fallback.** Set a threshold below which retrieved entries are discarded, and fall back to a hand-curated default block when too few survive. This turns "no good neighbour" from a silent quality drop into an explicit, testable path. **Per-label cap.** Allow at most n retrieved exemplars per label, which forces the block to show more than one side of a boundary. It costs some similarity — the second-class examples are further away — and that trade is worth measuring rather than assuming. **De-duplication.** Collapse near-identical entries so k demonstrations carry k examples' worth of signal. This can be done at write time on the pool or as a filter over the retrieved candidates; doing it on retrieval also protects you from duplicates that arrived after the last curation pass. **Over-retrieve then filter.** Fetch more candidates than you need and apply the caps and de-duplication to that larger set, so the guards do not simply shrink the block. ## Making it observable None of these failures show up in an aggregate accuracy number until they are severe, because they concentrate on a minority of queries. Log, per request: the retrieved exemplar ids, their similarity scores, and the label distribution of the retrieved set. Then you can alert on two cheap signals — the share of requests whose retrieved set is single-label, and the share whose top similarity sits below the floor. Both are leading indicators: the first tells you the boundary is being hidden from the model, the second tells you the live input distribution has drifted away from the pool. When you evaluate, slice by those same signals. A dynamic-exemplar system almost always looks fine on the head of the distribution and does its damage in the tail, so an average over a balanced test set will happily hide the problem you were hired to find.
- Your similarity floor now rejects the neighbourhood on 30% of live requests. What does that tell you?That the live input distribution has moved away from the pool — new products, new phrasing, a new customer segment — or that the pool never covered that segment. It is a drift alarm, not a threshold-tuning problem. The fix is to widen the pool with labelled cases from the segment that is missing; lowering the floor only restores the silent failure you were catching.
- Why is a single-label retrieved set worse than an unbalanced fixed few-shot block?A fixed block's imbalance is visible in the template, reviewable, and identical for every request, so you can calibrate against it. A retrieved set's imbalance is generated per query, differs every call, and concentrates on exactly the ambiguous inputs where the model most needs to see both sides. It never appears in a diff, so it is found only by logging the label mix.
- How would you keep evaluation honest when production inputs are written back into the pool?Pin the evaluation to a pool snapshot taken before the eval inputs existed, and exclude any pool entry whose input matches an eval item. Otherwise retrieval can hand the model the exact answer and you measure a lookup, not generalisation — a score that collapses the moment genuinely unseen traffic arrives.
Asking the four colleagues sitting nearest your desk how to decide a contentious case: they are the easiest to reach and they mostly agree with each other, which feels like consensus and is really just proximity.
saying these in an interview costs you the question
- Assuming the most similar examples are automatically the most useful
- Ignoring that a novel input still gets k demonstrations returned
- Treating near-duplicate pool entries as k independent shots
- Judging the system on aggregate accuracy instead of the tail
- Fixing a low-similarity alarm by lowering the threshold