How do embeddings and a VectorStore fit into a RAG retrieval pipeline, and what are the main failure modes to design around?
answer
- ingest: read -> TokenTextSplitter -> add()
- query: embed -> similaritySearch -> stuff prompt
- QuestionAnswerAdvisor automates retrieve+augment
- gotchas: chunking, model mismatch, filters, staleness, cost
- retrieval quality caps answer quality
basics
~20 sIn RAG you ingest documents (read, chunk, embed, store), then at query time embed the user's question, run similaritySearch to fetch the most relevant chunks, and stuff them into the prompt as context. Main risks: bad chunking, model mismatch, weak filtering, and stale data.
solid answer
~50 sRAG (Retrieval-Augmented Generation) has two phases. Ingestion: a DocumentReader loads sources, a TokenTextSplitter chunks them into passage-sized Documents, and VectorStore.add embeds and persists them. Query time: embed the question, similaritySearch(SearchRequest) with topK, threshold, and metadata filters returns the nearest chunks, which you inject into the ChatClient prompt as grounding context — Spring AI's QuestionAnswerAdvisor automates exactly this retrieve-then-augment step. The failure modes a principal designs around: (1) chunking — too big blurs meaning and blows token limits, too small loses context; (2) model consistency — query and corpus must use the same EmbeddingModel or scores are noise, and switching models means a full re-index; (3) retrieval quality — tune topK/threshold and use metadata filters for tenant/freshness scoping to avoid noise and leakage; (4) staleness/idempotency — stable ids and re-ingestion strategy; (5) cost/latency — batch embeddings, cache, and cap topK. Retrieval feeds generation, so retrieval quality caps answer quality.
code
java · 34 lines// Ingestion (ETL): read -> split -> store
@Service
class RagIngestion {
private final VectorStore vectorStore;
RagIngestion(VectorStore vectorStore) { this.vectorStore = vectorStore; }
void ingest(Resource pdf) {
var reader = new PagePdfDocumentReader(pdf);
var splitter = new TokenTextSplitter(); // chunk into passages
List<Document> chunks = splitter.apply(reader.get());
vectorStore.add(chunks); // embeds + persists
}
}
// Query time: let QuestionAnswerAdvisor retrieve + augment automatically
@Service
class RagQuery {
private final ChatClient chatClient;
RagQuery(ChatClient.Builder builder, VectorStore vectorStore) {
this.chatClient = builder
.defaultAdvisors(new QuestionAnswerAdvisor(
vectorStore,
SearchRequest.builder()
.topK(4)
.similarityThreshold(0.7)
.filterExpression("tenant == 'acme'") // isolation
.build()))
.build();
}
String ask(String question) {
return chatClient.prompt().user(question).call().content();
}
}go deeper
Know RAG = retrieve relevant text with similaritySearch, then add it to the prompt.
Describe the ingest (read/split/store) and query (embed/search/augment) phases and name TokenTextSplitter and QuestionAnswerAdvisor.
Discuss tuning topK/threshold/filters, chunking tradeoffs, and same-model consistency plus re-index on model change.
Own the full design: chunking strategy, model/dimension governance, multi-tenant isolation via filters, idempotent re-ingestion, cost/latency budget, and monitoring retrieval hit-rate as the cap on answer quality.
## What RAG is and why embeddings power it **RAG (Retrieval-Augmented Generation)** grounds an LLM's answer in your own data by retrieving relevant text and adding it to the prompt, instead of relying on the model's parametric memory. Embeddings + a `VectorStore` are the *retrieval* engine. ## Phase 1 — Ingestion (ETL) Spring AI models this as Extract-Transform-Load: - **Extract**: a `DocumentReader` (`TikaDocumentReader`, `PagePdfDocumentReader`, `JsonReader`, `TextReader`) turns raw sources into `Document`s. - **Transform**: a `DocumentTransformer` — most importantly `TokenTextSplitter` — chunks large docs into retrieval-sized passages; enrichers like `KeywordMetadataEnricher`/`SummaryMetadataEnricher` can add metadata. - **Load**: `VectorStore` (a `DocumentWriter`) embeds each chunk via the `EmbeddingModel` and persists content+vector+metadata. ## Phase 2 — Retrieval + generation 1. Embed the user query (same `EmbeddingModel`). 2. `similaritySearch(SearchRequest.builder().query(q).topK(k).similarityThreshold(t).filterExpression(f).build())` returns the nearest chunks. 3. Inject those chunks into the prompt as context, then call the `ChatClient`. Spring AI provides `QuestionAnswerAdvisor` (and the newer modular `RetrievalAugmentationAdvisor`) that wraps steps 1-3 automatically: attach it to a `ChatClient` and it retrieves and augments per request. You can pass a `SearchRequest` (with filters/topK/threshold) into the advisor. ## Failure modes a principal designs around **1. Chunking strategy.** Chunks that are too large produce one blurry vector spanning many topics (poor precision) and risk exceeding the embedding model's token limit; too small lose surrounding context. Tune chunk size and overlap. This is the single biggest lever on retrieval quality. **2. Model consistency.** Query and corpus must be embedded by the **same** model — mixing models yields meaningless scores. Switching models later forces a **full re-embed and re-index**. Governance: pin the model per index and treat model change as a data migration. **3. Retrieval tuning.** `topK` too small misses relevant context; too large floods the prompt with noise and burns tokens/cost. `similarityThreshold` filters weak matches but risks empty results. Use **metadata filters** for multi-tenant isolation, freshness windows, and access control — a missing filter can leak another tenant's data into an answer. **4. Data freshness & idempotency.** Use stable `Document` ids so re-ingestion updates rather than duplicates; define a re-index cadence; delete-by-filter to purge stale docs. Duplicates inflate topK with near-identical chunks. **5. Cost & latency.** Every embed call is billed and adds latency; batch ingestion embeddings, cache frequent query embeddings, cap topK, and consider approximate indexes (HNSW) for QPS — accepting slightly lower recall. **6. Grounding/hallucination limits.** RAG reduces but doesn't eliminate hallucination; if retrieval returns nothing relevant, the model may still answer from memory. Design prompts to say "I don't know" when context is empty, and monitor retrieval hit-rate. ## Distance metric & normalization Text embeddings are typically compared by **cosine similarity**; Spring AI normalizes `getScore()` so higher = more similar regardless of the store's raw metric. Keep the metric consistent with how the model was trained. ## When to use RAG (vs alternatives) Use RAG for large, changing, or private corpora where fine-tuning is too slow/expensive and you need citable, up-to-date grounding. For tiny static context, just put it in the prompt. For behavior/style change rather than knowledge, prefer fine-tuning. ## Summary Retrieval quality caps generation quality: no matter how good the LLM, if `similaritySearch` returns the wrong chunks, the answer is wrong. So the embedding model, chunking, filters, and store tuning are the real engineering surface — the `ChatClient` call is the easy part.
- Why is chunk size the biggest lever on RAG quality?The chunk is the unit that gets one embedding vector and one retrieval decision. Too large blurs multiple topics into one vector (low precision) and can exceed token limits; too small strips context so retrieved passages are unusable. Right-sizing (with overlap) directly controls what similaritySearch can find.
- How does Spring AI let you wire RAG without hand-coding the retrieve-then-augment loop?Attach a QuestionAnswerAdvisor (or RetrievalAugmentationAdvisor) to the ChatClient with a VectorStore and optional SearchRequest. On each prompt it embeds the question, runs similaritySearch with your topK/threshold/filters, injects the results as context, and calls the model — no manual retrieval code.
- How do metadata filters prevent data leakage in multi-tenant RAG?Store a tenant id in each Document's metadata at ingest, then include a filterExpression like tenant == 'acme' in the SearchRequest. The store restricts nearest-neighbor search to that tenant's rows server-side, so one tenant's chunks can never surface in another tenant's answer.
saying these in an interview costs you the question
- Treating the LLM call as the hard part while ignoring retrieval quality
- Embedding the query with a different model than the corpus
- Ingesting whole documents without chunking
- Omitting tenant/access metadata filters and leaking data across tenants
- Believing RAG fully eliminates hallucination