skip to content

Your needle-in-a-haystack eval passes but long-document QA fails in production — why?

level: seniorimportance: should knowfreq 48%

answer

  1. your eval is easier than production
  2. verbatim lookup is the easy case
  3. one needle versus several confusable ones
  4. hops across distant pages
  5. aggregation has no needle to find

basics

~20 s

Single-needle retrieval of a verbatim string is the easiest long-context task and is largely saturated on current models. Production work needs multi-hop reasoning, aggregation across the document, and resistance to distractors — none of which that eval measures.

solid answer

~50 s

A classic needle-in-a-haystack test plants one distinctive sentence in bland filler and asks the model to quote it back. That is near-exact string matching over a low-noise background, and modern models pass it comfortably — so passing tells you almost nothing about real work. Production long-document QA fails on three harder axes the eval never touches. Multi-hop: the answer requires joining facts at distant positions, for example a dose stated on page 40 and an amendment on page 260 that supersedes it, where the model happily returns the stale dose. Aggregation: the question needs the whole document, such as counting or reconciling every occurrence, so there is no single needle to find. Distraction: the real corpus is full of similar-but-wrong passages, unlike inert filler. Modern suites test these directly — RULER adds multi-key, multi-hop and aggregation variants; MRCR-style tests plant multiple confusable needles. Build your eval from real documents and real question shapes.

go deeper

for a junior

Know that finding one planted sentence in filler is a much easier task than answering real questions over a long document, so passing that test does not mean the system works.

for a middle

Explain the three axes a single-needle test misses — multi-hop, aggregation and distractors — and describe how you would extend the eval with several confusable needles and questions requiring two distant facts.

for a senior

Show you build evals from real documents and production question shapes, sweep length and position per category, repeat cells because failures are stochastic, and read the resulting staircase as a per-category break point.

for a principal

Own the decision the staircase drives: which question categories are served by one long prompt and which must be decomposed or routed. Treat the eval as the regression gate for every model version change.

## Why single-needle passes are cheap The original needle-in-a-haystack construction is deliberately simple: take filler text with no relationship to the question, insert one memorable and syntactically distinctive sentence, ask a question whose answer is that sentence, and check whether it comes back. Two properties make it easy. The needle is **lexically unique** — it shares almost no vocabulary with its surroundings, so even shallow matching finds it. And the answer is a **verbatim copy**, so no reasoning stands between finding it and emitting it. As of mid-2026 essentially every serious model passes this at most depths within its comfortable range. A green board on single-needle is therefore not evidence of long-context competence; it is evidence that you measured the one long-context task that got solved first. Teams that ship on the strength of it are surprised in production, and the surprise has a predictable shape. ## The three axes production adds **Multi-hop.** Real questions require joining facts that sit far apart. Take a clinical-trial protocol assistant: the starting dose is specified on page 40 of the protocol, and an amendment on page 260 supersedes it with a lower dose. Answering "what dose does this patient start on?" needs both spans, plus the reasoning that the later amendment wins. The characteristic failure is not "I don't know" — it is confidently returning the stale page-40 dose, because that span alone is a perfectly good answer to a single-needle-shaped reading of the question. Single-needle evals cannot detect this failure mode at all, because they never plant two facts that interact. **Aggregation.** Some questions are about the document as a whole: how many sites enrolled patients, which sections were modified by any amendment, is any statement in section 9 contradicted elsewhere. There is no needle; the model must use everything. Performance here degrades much earlier with length than retrieval does, and it degrades differently — you get plausible, well-formatted, partially-correct answers rather than misses, which makes it far more dangerous. **Distraction.** Filler in a synthetic eval is inert. A real corpus is adversarial by accident: superseded versions, near-duplicate boilerplate, sibling documents that answer a *similar* question. Confusable neighbours compete with the correct span in a way that bland filler never does, so accuracy measured over filler systematically overestimates accuracy over your corpus. ## What a serious long-context eval looks like The published suites moved in exactly this direction, and their designs are worth copying. RULER extends the haystack idea with configurable multi-key retrieval (several needles, only one relevant), multi-hop variable tracking, aggregation and question-answering categories, and reports a per-model effective length rather than a single score. MRCR-style tests plant several deliberately confusable needles and ask for a specific one, which measures exactly the discrimination that single-needle skips. Multi-hop long-document suites push the number of required hops up until models break. For your own system, four design rules carry most of the value: 1. **Use your real documents.** Filler flatters; your corpus does not. If you cannot use production data, synthesize a haystack with the same redundancy and near-duplicate structure. 2. **Match the question shapes to production.** Mine actual user questions and classify them — single-fact lookup, multi-hop, aggregation, comparison — then keep the mix in your eval proportional to the mix you serve. 3. **Sweep length and position anyway.** Keep the two axes from the original construction; you want to know at what length each *category* breaks, and they break at very different lengths. 4. **Repeat every cell and report a rate.** Long-context failures are stochastic. A category that passes 70% of the time looks green at n=1 and is unshippable in production. ## Reading the results Expect a staircase, not one number: single-fact retrieval holds far out, multi-hop breaks earlier, aggregation earliest. That staircase is the actionable output — it tells you which question categories may be answered by one long prompt and which must be restructured, decomposed, or routed elsewhere. It also gives you a regression instrument: when a model version changes, the staircase moves, and you will see it before your users do. The common failure of this whole exercise is building an eval that is easier than production and trusting it. If your eval never surfaces a wrong-but-confident answer, it almost certainly does not resemble the work.

  • How do you build a multi-hop long-context eval item without hand-writing hundreds of them?
    Template the hop structure over real documents. Pick a fact type that recurs — a dosage, a threshold, an effective date — locate its occurrences programmatically, and generate items that require joining an original statement with the later one that supersedes it. The document supplies the content, the template supplies the reasoning shape, and the ground truth is derivable from the extraction rather than written by hand. Spot-check a sample by hand to confirm the templates are producing genuinely two-hop questions.
  • What signal tells you a long-context failure is a reasoning failure rather than a retrieval failure?
    Check whether the required spans were present and whether the model referenced them. If you shrink the prompt to just the two required passages and the model still returns the stale answer, the joining logic is the problem. If it answers correctly on the short prompt and fails on the long one with identical evidence, the failure is long-context degradation. Asking the model to quote both spans before answering makes this distinction observable in the trace.
  • Why report a per-category break point instead of one long-context score?
    Because the categories break at very different lengths, and a single averaged score hides which ones. A model may retrieve reliably at 200K while its aggregation accuracy collapsed at 40K; one number smears both into a mid-range figure that describes neither. Per-category break points map directly onto engineering decisions — this question shape can go in one prompt, that one needs decomposition — which an average cannot support.

saying these in an interview costs you the question

  • Treating a passing needle-in-a-haystack run as long-context validation
  • Assuming retrieval accuracy predicts multi-hop accuracy
  • Using bland filler that is nothing like the real corpus
  • Running each eval cell once and reading it as pass or fail
  • Reporting one long-context score across all question shapes

context