skip to content

As an architect, how do you design and evaluate a production RAG pipeline in Spring AI end to end?

level: principalimportance: nice to knowfreq 30%

answer

  1. Two pipelines: offline ETL + online serving
  2. Chunking + topK + threshold + filter = the tuning knobs
  3. Embedding-model version must match index and query (re-embed on change)
  4. allowEmptyContext + tenant filter = anti-hallucination + anti-leakage
  5. Evaluate: recall@k, RelevancyEvaluator/FactCheckingEvaluator

basics

~20 s

Design two paths: an offline ETL (read, chunk, enrich, embed, write to VectorStore) and an online query path (retrieve with SearchRequest, augment, generate). Tune chunking, topK, threshold, and metadata filters; measure retrieval quality and guard against empty-context hallucination.

solid answer

~40 s

I separate **ingestion** from **serving**. Ingestion is an offline pipeline: `DocumentReader` -> `TokenTextSplitter` (chunking is the highest-leverage knob) -> optional enrichers -> `vectorStore.add()`, run as a batch/job with idempotent IDs and metadata (tenant, source, version) for later filtering. Serving uses a `ChatClient` advisor: `QuestionAnswerAdvisor` for simple cases, `RetrievalAugmentationAdvisor` when I need query rewriting, multilingual translation, multi-query expansion, or explicit empty-context guarding via `ContextualQueryAugmenter`. I tune `SearchRequest` (topK vs token budget, similarityThreshold, metadata filterExpression for multi-tenancy). Cross-cutting: pick an `EmbeddingModel` and version it (re-embed on change), choose a scalable `VectorStore` (PgVector/Redis), manage context-window budget, add prompt instructions to stay within context, and evaluate retrieval with labeled Q/A sets. I also handle cost/latency of extra LLM calls from transformers and observability on which chunks were retrieved.

code

java · 16 lines
java
// Serving path: modular RAG with tenant scoping + empty-context guard
Advisor rag = RetrievalAugmentationAdvisor.builder()
    .documentRetriever(VectorStoreDocumentRetriever.builder()
        .vectorStore(vectorStore)
        .topK(6)
        .similarityThreshold(0.7)
        // security: never retrieve another tenant's chunks
        .filterExpression(() -> new FilterExpressionBuilder()
            .eq("tenantId", currentTenant()).build())
        .build())
    .queryAugmenter(ContextualQueryAugmenter.builder()
        .allowEmptyContext(false) // refuse safely instead of hallucinating
        .build())
    .build();

String answer = chatClient.prompt().advisors(rag).user(userQuestion).call().content();

go deeper

for a junior

Not expected; awareness that RAG has an ingest side and a query side is enough.

for a middle

Should separate ingestion from serving and name the main tuning knobs.

for a senior

Discuss embedding-model versioning, VectorStore choice, empty-context guarding, and basic evaluation.

for a principal

Own the whole system: eval-driven tuning, cost/latency budgets, multi-tenant security, prompt-injection from retrieved content, and RAG-vs-alternatives decisions.

A production RAG system in Spring AI is two decoupled pipelines plus a set of cross-cutting concerns. **1. Offline ingestion (ETL).** - `DocumentReader` (Tika/PDF/JSON/Text) extracts `Document`s; attach rich metadata early: `tenantId`, `source`, `docVersion`, `lang`, `category`. - `TokenTextSplitter` chunks — the single most impactful parameter. Tune chunk size/overlap per corpus; evaluate empirically. - Optional enrichers (`KeywordMetadataEnricher`, `SummaryMetadataEnricher`) improve retrieval/filtering at the cost of ingest-time LLM calls. - `vectorStore.add()` embeds (via `EmbeddingModel`) and stores. Use **stable Document IDs** so re-ingestion upserts instead of duplicating. Run as a batch job or on-upload, never in the request hot path for large corpora. **2. Online serving.** - Choose the advisor: `QuestionAnswerAdvisor` (simple, cheap, one search) vs `RetrievalAugmentationAdvisor` (modular: query rewrite/translate/compress, multi-query expansion, post-retrieval rerank, `ContextualQueryAugmenter` with `allowEmptyContext` to prevent hallucination on no-hit). - Configure `SearchRequest`: `topK` balances recall against token budget and noise; `similarityThreshold` trims weak matches (too high => empty context); `filterExpression` enforces tenant/category scoping — **security-relevant**: filtering by tenant metadata is how you prevent cross-tenant leakage. **3. Cross-cutting concerns.** - **Embedding model versioning:** the query and the stored chunks must be embedded by the *same* model/version; changing models requires re-embedding the whole corpus. Treat it as a migration. - **VectorStore choice:** `SimpleVectorStore` (in-memory, dev only) vs `PgVectorStore`/`RedisVectorStore`/dedicated vector DBs for scale, persistence, filtering, and ANN indexing. - **Context-window budget:** topK × chunk size + system prompt + history must fit; degrade gracefully (compress or reduce topK). - **Grounding/anti-hallucination:** instruct the model to answer only from context; use `allowEmptyContext=false`; optionally return citations by surfacing the retrieved Documents. - **Cost & latency:** every query transformer/expander is an extra LLM round-trip; multi-query multiplies retrievals. Cache where possible; measure P95. - **Evaluation:** build a labeled question/answer/ground-truth set; measure retrieval metrics (recall@k, precision) and answer quality; Spring AI offers `Evaluator` abstractions (e.g., `RelevancyEvaluator`, `FactCheckingEvaluator`) to automate grading. Feed results back into chunking/topK/threshold tuning. - **Observability:** log/trace which chunks were retrieved and their scores (Micrometer/OTel via Spring AI observability) for debugging bad answers. - **Security & privacy:** tenant-scoped filters, PII handling at ingest, and prompt-injection awareness (retrieved content can contain adversarial instructions — sanitize/limit its authority). **When RAG is the wrong tool:** if the task needs behavioral/style change rather than knowledge, fine-tuning may fit better; if data is tiny and static, stuffing it directly into the prompt beats a vector store; if answers require multi-hop reasoning over structured data, an agent with tools/SQL may outperform pure similarity retrieval. **Key gotchas:** duplicate ingestion, mismatched embedding models between index and query, empty-context hallucination, cross-tenant leakage from missing filters, context overflow from an over-large topK, and unbounded cost from stacking query transformers.

  • What breaks if you swap the embedding model without re-ingesting?
    Query embeddings and stored chunk embeddings come from different vector spaces, so similarity scores are meaningless and retrieval degrades to noise. Changing embedding models requires re-embedding (re-ingesting) the entire corpus — treat it as a data migration.
  • How do you prevent one tenant from retrieving another tenant's documents?
    Tag every Document with a tenantId at ingest, then always apply a filterExpression scoping the SearchRequest/DocumentRetriever to the current tenant. Missing or bypassable filters are a cross-tenant data-leak vulnerability, so enforce it server-side, not from client input.
  • How would you measure whether the RAG system is actually good?
    Build a labeled eval set (questions + ground-truth answers/relevant chunks) and measure retrieval recall@k/precision plus answer quality, automating grading with Spring AI's Evaluator abstractions like RelevancyEvaluator and FactCheckingEvaluator, then tune chunking/topK/threshold from the results.

saying these in an interview costs you the question

  • Ignoring that the embedding model must be identical between indexing and querying
  • Running heavy ingestion synchronously in the request path
  • Omitting tenant/metadata filters, enabling cross-tenant leakage
  • Treating retrieved document text as trusted (prompt-injection risk)
  • Setting topK so high the context window overflows
  • Assuming RAG replaces fine-tuning for behavioral/style changes

context