When does a million-token context window replace RAG, and when not?
answer
- ask for volume, size, churn, audit needs
- input tokens are charged every request
- caching needs a stable prefix
- no window holds an unbounded corpus
- hybrid: retrieve sections, not fragments
basics
~20 sLong context wins for one-off or low-volume work over a bounded document set where any part may matter. Retrieval wins on economics, freshness and scale: resending the same corpus every request costs input tokens per call, and corpora larger than any window, or changing hourly, cannot be pasted at all.
solid answer
~60 sFrame it as economics and freshness rather than accuracy, because on a small enough document set a long-context prompt often reads *better* than eight retrieved chunks. **Cost.** Pasting a 300-page handbook — roughly 200k tokens — into every request means paying for 200k input tokens per question. At thousands of questions a day that dwarfs the cost of retrieving eight chunks. Prompt caching cuts the repeated-prefix cost substantially where the corpus is static and the traffic is dense enough to keep the cache warm, which narrows but does not close the gap. **Freshness and scale.** A corpus that changes hourly invalidates a cached prefix constantly, and a corpus of millions of documents does not fit in any window. Retrieval indexes update incrementally and are unbounded in corpus size. **Attribution.** Retrieval hands you the passage that supported each claim, which matters when the answer must be auditable. The common production answer is hybrid: retrieve coarsely, then give the model whole sections rather than small fragments, spending the large window on precision of *evidence* rather than on the whole corpus.
go deeper
Know that you can either paste documents into a large context window or retrieve the relevant parts, and that pasting everything costs tokens on every single request.
Explain the cost arithmetic with real numbers, and name freshness, corpus size and attribution as the reasons retrieval survives large windows.
Show you would measure before deciding: cost per answered question, update frequency against cache viability, and where a hybrid that retrieves whole sections beats both extremes.
Own the decision rule and the migration path. Define the thresholds on volume, corpus growth and churn that trigger a move, decide what auditability the domain requires, and avoid committing the organisation to an index it does not yet need or to a prompt that will not scale.
## Why this is now a real decision When context windows held a few thousand tokens, RAG was not a choice, it was the only way to use a large corpus at all. With frontier windows around a million tokens as of mid-2026, whole document sets fit, and the question becomes a genuine architectural tradeoff that interviewers ask precisely because the naive answer ("big windows killed RAG") is wrong in a specific and instructive way. ## Start with the economics Take a concrete comparison: a 300-page employee handbook, roughly 200,000 tokens. Option A pastes it into the prompt every turn. Option B indexes it and retrieves eight relevant chunks, maybe 4,000 tokens. Option A pays for 200,000 input tokens on every single question. Option B pays for 4,000 plus a cheap vector search. That is a fifty-fold difference in input volume, and it is charged per request, so it scales with traffic. At ten questions a day nobody cares. At fifty thousand questions a day it is the dominant line item in the system's cost, and it buys you very little on a corpus this small, because a well-tuned retriever finds the leave-policy paragraph reliably. Providers offer prompt caching for exactly this shape — a large, stable prefix reused across requests — and it materially reduces the repeated cost of the pasted handbook. Two conditions have to hold: the prefix must be genuinely stable, since any edit invalidates it, and traffic must be dense enough to keep entries warm within their lifetime. Caching narrows the gap; it rarely reverses the conclusion at high volume. Latency follows the same shape. Processing a very large prompt takes time before the first token appears, on every request, whereas retrieval front-loads the work into an offline indexing job. ## Then freshness This is the argument people skip, and it is often the decisive one. A retrieval index is updated incrementally: a document changes, you re-embed that document, and the next query sees it. A pasted corpus is assembled at request time from whatever your loader read, which is fine, but if you are relying on caching to make the economics work, every update invalidates the cache and the economics revert. So the pattern is: **static corpus plus high volume equals caching works; volatile corpus plus high volume equals retrieval**, because the volatility that breaks the cache does not bother an index at all. A price list, an on-call rotation or a ticket queue changes far too often to be a cacheable prefix. ## Then scale The window is a hard ceiling. A handbook fits; a 50-gigabyte document archive, a decade of support tickets, or an entire codebase across many repositories does not, and no near-term window growth changes that by the orders of magnitude required. Once the corpus exceeds the window, you are doing retrieval whether or not you call it that — something must choose what goes in. ## Then attribution When a compliance officer asks which clause supports an answer, retrieval hands over the passage, its document, its version and its retrieval score. A long-context system can be asked to cite, and usually does so competently, but the citation is generated rather than recorded, so it is a claim about provenance rather than evidence of it. In regulated domains that difference decides the architecture on its own. ## Where long context genuinely wins Be fair to the other side, because a candidate who only argues one way sounds rehearsed. Long context wins when the corpus is bounded and the question is *global*. "Summarise every obligation this 200-page contract places on us" or "where do these two policies contradict each other?" cannot be served by retrieving eight similar chunks, because the answer depends on material that is not similar to the query. Chunking also destroys long-range structure — a defined term on page 3 governing a clause on page 180 — that a single pass preserves. It also wins on volume: one-off analyses, low-traffic internal tools, and exploratory work where building and maintaining an index is more engineering than the use case justifies. And it wins as a prototype: paste the corpus first, establish the quality ceiling, and only build retrieval when volume or corpus size forces it. ## The hybrid that most production systems land on The useful synthesis is that a large window changes *how much* you retrieve, not whether. Instead of eight 400-token fragments, retrieve whole sections or whole documents and pass fifty thousand tokens of directly relevant material. You keep incremental freshness, unbounded corpus size and recorded provenance, while spending the window on evidence density rather than on resending everything. Some systems add a coarse routing step — identify the three relevant documents, then read them whole — which is retrieval at document granularity rather than chunk granularity. ## How to answer this in an interview Refuse the framing that one killed the other. Ask for the numbers: corpus size, update frequency, query volume, and whether answers must be auditable. Those four determine the answer, and quoting them back is what distinguishes a judgement answer from an opinion.
- How much does prompt caching change the calculation, and what has to be true for it to help?It substantially reduces the cost of a large, unchanging prefix reused across requests, so it can make long context viable at moderate volume. Two conditions must hold: the prefix must be stable, because any edit to the pasted corpus invalidates the entry, and traffic must be frequent enough that entries stay warm within their lifetime. Volatile corpora or bursty low traffic get little benefit, and the retrieval argument stands.
- Give a question over a bounded corpus that retrieval handles badly and long context handles well.Any global question whose answer is not similar to the query. "List every deadline this contract imposes on us" requires every obligation clause, and top-k similarity retrieval returns the k that look most like the question rather than all of them. Cross-document contradiction hunting is the same shape. These need full coverage, so either a long-context pass or a map-reduce over the whole corpus, not selective retrieval.
- Your handbook assistant is cheap today at low volume using a pasted corpus. What triggers the migration to retrieval?Set thresholds in advance rather than migrating in a panic. Trigger on sustained query volume where per-request input cost exceeds the amortised cost of running an index, on corpus growth approaching a meaningful fraction of the window, on update frequency rising to the point that caching stops paying, or on a compliance requirement for recorded provenance. Instrument cost per answered question so the crossing point is visible before it is expensive.
- Is a hybrid meaningfully different from just retrieving more chunks?Yes, in granularity and in what it preserves. Retrieving whole sections or documents keeps internal structure and long-range references that chunking severs, and it lets a coarse router work at document level where precision is easier. You still get incremental index updates, unbounded corpus size and recorded provenance, while spending the large window on dense relevant evidence rather than on the entire corpus.
Pasting the whole corpus is rereading the entire manual before answering each question; retrieval is a well-made index. For one question the reread is fine, and it catches things the index misses — but not fifty thousand times a day, and not when the manual is rewritten weekly.
saying these in an interview costs you the question
- Claiming large context windows made RAG obsolete
- Comparing the two on accuracy alone and ignoring cost per request
- Assuming prompt caching is free and always applicable
- Forgetting that corpora can exceed any window by orders of magnitude
- Treating a generated citation as recorded provenance