What failures push a RAG system from naive to advanced to modular?
answer
- three rungs, three failure classes
- single pass versus data-dependent control flow
- pre-retrieval and post-retrieval stages
- loops and routing arrive with modular
- measure before you climb
basics
~20 sNaive RAG retrieves top-k once and generates. Advanced RAG adds pre- and post-retrieval stages when recall or precision fails. Modular RAG turns those stages into swappable components with routing, conditionals and loops, so control flow depends on the query.
solid answer
~60 sThe three names form a maturity ladder, and each rung is forced by a different class of failure. **Naive RAG** chunks a corpus, embeds it, pulls top-k by vector similarity and stuffs the chunks into one prompt. It assumes retrieval is always needed, always fires once, and always returns something usable. **Advanced RAG** keeps the single pass but wraps it: better chunking and metadata, hybrid lexical plus dense search, query transformation before retrieval, reranking and compression after it. You climb here when measurement shows the generator is fine and the retriever is the problem — poor recall, or ten chunks of which two are relevant. **Modular RAG** is the architectural step: retrieval becomes one interchangeable module among routing, fusion, memory and generation, with conditionals and loops between them. You climb here when no single pass can succeed — the answer spans several hops, the corpus does not contain it, or some queries need no retrieval at all. Adaptive, self-reflective, corrective and agentic RAG are patterns that only become expressible once control flow is data-dependent.
go deeper
Be able to draw the naive pipeline end to end — chunk, embed, store, retrieve top-k, generate — and name one thing that goes wrong with it, such as the retriever missing a passage because the question uses different words.
Explain what advanced RAG adds on each side of retrieval and why reranking and hybrid search fix different problems. Say plainly that modular RAG is the point where control flow becomes data-dependent.
Show the diagnostic habit: attribute observed failures to retrieval or generation with per-stage evaluation, and justify each added component by the failure it removes and the latency and cost it adds.
Own the tradeoff between capability and operability. Every branch and loop multiplies the paths you must evaluate, monitor and explain to an auditor, so argue for the simplest architecture that meets the quality bar and set the evidence standard the team must meet before climbing a rung.
## What the three names actually describe Naive, advanced and modular RAG are not three products or three libraries. They describe what a retrieval-augmented system looks like after each class of failure has been engineered away. Reading them as a maturity ladder is the framing an interviewer is usually after, because it forces you to say *which measured failure* justified each added moving part rather than reciting a diagram. ## Rung one: naive RAG The baseline pipeline is: split documents into chunks, embed each chunk into a vector, store the vectors, embed the user question, take the k nearest chunks by similarity, paste them into a prompt template, and generate one answer. It is a straight line with no branches. That line carries four silent assumptions. First, that the question is worded like the answer, so a single similarity comparison finds it. Second, that the answer fits inside the chunks the splitter happened to cut. Third, that the top k chunks are worth reading — precision is never checked. Fourth, that retrieval should always happen, even for "hi" or for a question the model already knows. The symptoms are recognisable. Vocabulary mismatch: the user asks about "time off", the policy says "paid leave of absence", and cosine similarity misses it. Boundary damage: a table's header lands in one chunk and its rows in another. Dilution: eight of ten retrieved chunks are noise and the generator averages over them. Confident nonsense: retrieval returned nothing relevant, but nothing in the pipeline can say so, so the model answers anyway. ## Rung two: advanced RAG Advanced RAG keeps the single pass and attacks those symptoms with stages on either side of it. Pre-retrieval work improves what goes in: semantically aware or structure-aware chunking, chunk overlap, metadata and filters, richer indexing, and transformations applied to the query before it is embedded. Retrieval itself often becomes hybrid — a lexical index catches exact identifiers and rare terms that embeddings blur, a dense index catches paraphrase, and the two result lists are combined. Post-retrieval work improves what reaches the model: retrieve a wide candidate set, then rerank it with a cross-encoder or a model that scores query and passage jointly, and keep only the few best; compress or summarise passages so the prompt carries signal rather than volume; order passages deliberately rather than by raw score. The engineering discipline here is measurement. You climb to this rung when your evaluation separates retrieval quality from generation quality and shows the retriever is at fault: the answer was present in the corpus and the retriever did not surface it (a recall problem) or surfaced it buried in noise (a precision problem). ## Rung three: modular RAG Advanced RAG still executes one fixed sequence for every query. Modular RAG breaks the pipeline into interchangeable modules — routing, retrieval, fusion, memory, grading, generation — and, crucially, adds control flow between them: conditionals, branches and loops whose direction depends on the query and on intermediate results. You climb here when the remaining failures are ones no single pass can fix. The question needs two hops, where the second query can only be written after seeing the first result. The corpus genuinely does not contain the answer, and the correct behaviour is to say so or to consult another source. Query classes differ so much that one pipeline cannot serve them — a greeting, a lookup and a comparative analysis want different treatment. Retrieval quality varies per query, so the system must inspect what it got back before trusting it. Every named variant of the last two years lives on this rung. Adaptive RAG decides whether to retrieve at all. Self-reflective designs have the generator critique its own retrieval and grounding. Corrective designs grade retrieved passages and act on the grade. Agentic designs expose retrieval as a tool an agent may call repeatedly with reformulated queries. None of these is expressible without data-dependent control flow, which is precisely why modular is a genuine architectural step rather than more stages bolted on. ## What each rung costs The ladder is not a scoreboard to climb for its own sake. Advanced RAG adds an extra model call for reranking and often one for query transformation, which shows up as latency on every request. Modular RAG adds variable numbers of LLM calls, non-deterministic paths, and far more surface to evaluate and debug — a bad answer might now come from the router, the grader, the retriever or the generator. Cost per question stops being a constant. So the honest answer in an interview is diagnostic, not aspirational: instrument the naive pipeline first, attribute failures to retrieval or generation, and add exactly the mechanism that the measured failure demands. A team that jumps to an agentic loop because the demo felt weak usually has a chunking problem and now has a chunking problem wrapped in three extra LLM calls.
- How would you decide whether a bad answer is a retrieval failure or a generation failure?Evaluate the stages separately. Check whether the supporting passage exists in the corpus and whether the retriever returned it in the top k — that is recall. If it did and the answer is still wrong, the failure is generation or grounding, and you inspect whether the answer's claims are actually supported by the retrieved text. Only retrieval failures justify reranking, hybrid search or better chunking; generation failures need prompt, grounding or model work.
- Does modular RAG always beat advanced RAG on quality?No. Loops and routing help only when a single pass genuinely cannot succeed. On corpora where the answer usually sits in one well-chunked passage, an adaptive system mostly pays extra LLM calls and latency to reconfirm what one retrieval already found, and adds new failure modes — a mis-grading, a bad route, a loop that never converges. Quality gains have to be measured against a tuned single-pass baseline, not against a naive one.
- Where does reranking sit in this ladder, and why is it usually the highest-value addition?Reranking is post-retrieval advanced RAG: retrieve a wide candidate set cheaply by vector or hybrid search, then score each candidate against the query with a model that reads both together, and keep the few best. It pays well because it fixes precision without changing the architecture — one extra call, no loops, no branching — and precision is what the generator is most sensitive to.
saying these in an interview costs you the question
- Treating modular RAG as strictly better rather than as answering a different failure
- Adding an agentic loop before measuring whether chunking is the real problem
- Claiming advanced RAG means bigger k or a bigger model
- Describing the ladder as a version history of frameworks rather than of failures
- Assuming every question needs retrieval at all