Describe Spring AI's ETL pipeline for ingesting documents into a vector store (Reader, Transformer, Writer).
answer
- Extract=Reader (Supplier), Transform=Transformer (Function), Load=Writer (Consumer)
- TokenTextSplitter chunks by tokens
- VectorStore IS a DocumentWriter -> add() embeds + stores
- Tika/PDF/JSON/Text readers
- Enrichers add metadata via LLM
basics
~10 sYou read source files into Documents with a DocumentReader, split/enrich them with a DocumentTransformer (e.g., TokenTextSplitter), then write them into the VectorStore with a DocumentWriter, which embeds and stores them for later retrieval.
solid answer
~40 sSpring AI's ingestion is a three-stage ETL over `Document` objects. **Extract:** a `DocumentReader` (a `Supplier<List<Document>>`) reads a source — `TextReader`, `JsonReader`, `PagePdfDocumentReader`, or `TikaDocumentReader` for PDF/HTML/Office. **Transform:** a `DocumentTransformer` (a `Function<List<Document>, List<Document>>`) reshapes them — most importantly `TokenTextSplitter`, which chunks large docs so each fits the embedding/context budget and retrieval is granular; enrichers like `KeywordMetadataEnricher` and `SummaryMetadataEnricher` add metadata via the LLM. **Load:** a `DocumentWriter` (a `Consumer<List<Document>>`) persists them; `VectorStore` implements `DocumentWriter`, so `vectorStore.add(docs)` embeds each chunk with the `EmbeddingModel` and stores the vectors. You compose these with plain functional calls. Good chunking is the highest-leverage step — chunks too large blow the context and dilute relevance; too small lose meaning.
code
java · 20 lines@Component
class IngestionService {
private final VectorStore vectorStore;
IngestionService(VectorStore vectorStore) {
this.vectorStore = vectorStore;
}
void ingest(Resource pdf) {
// Extract: read the source into Documents
DocumentReader reader = new TikaDocumentReader(pdf);
List<Document> docs = reader.get();
// Transform: chunk into token-sized pieces for good retrieval
List<Document> chunks = new TokenTextSplitter().apply(docs);
// Load: VectorStore is a DocumentWriter -> embeds + stores each chunk
vectorStore.add(chunks);
}
}go deeper
Name the three stages and that VectorStore.add stores documents for retrieval.
Map each stage to its functional interface, name concrete readers/transformers, and explain why chunking matters.
Discuss chunk-size tuning, metadata for later filtering, enricher cost/latency, and idempotent re-ingestion by id.
Treat ingestion as an offline data pipeline: batching, incremental updates, embedding-model versioning, and retrieval-quality evaluation feeding chunking choices.
Before RAG can retrieve anything, you must **ingest** your corpus. Spring AI models this as a classic **ETL (Extract, Transform, Load)** pipeline whose currency is the `Document` (text + metadata map). All three stages are just functional interfaces, so you can chain or wrap them freely. **1. Extract — `DocumentReader`** (`extends Supplier<List<Document>>`). It produces Documents from a source. Built-ins: - `TextReader` — plain text/`Resource`. - `JsonReader` — JSON, with configurable fields to extract. - `PagePdfDocumentReader` / `ParagraphPdfDocumentReader` — PDF (from `spring-ai-pdf-document-reader`). - `TikaDocumentReader` — Apache Tika, handles PDF, HTML, Word, PowerPoint, etc. (from `spring-ai-tika-document-reader`). Markdown reader also exists. Metadata like source filename/page is attached here. **2. Transform — `DocumentTransformer`** (`extends Function<List<Document>, List<Document>>`). Reshapes the list: - `TokenTextSplitter` — **the key one**: splits long Documents into chunks sized by token count (default ~800 tokens, with min-size and overlap-ish settings). Chunking matters because embeddings and the context window are bounded, and retrieval returns whole chunks — you want each chunk to be one coherent, retrievable idea. - `ContentFormatTransformer` — applies a `ContentFormatter` controlling how text+metadata render. - `KeywordMetadataEnricher` — uses the LLM to add keyword metadata. - `SummaryMetadataEnricher` — uses the LLM to add per-chunk summaries (optionally of neighboring chunks). **3. Load — `DocumentWriter`** (`extends Consumer<List<Document>>`). Persists the result: - `VectorStore` **implements `DocumentWriter`**, so calling `vectorStore.add(documents)` (or `accept(documents)`) triggers the `EmbeddingModel` to compute a vector per Document and stores vector+text+metadata. This is what makes them searchable later. - `FileDocumentWriter` — writes to a file (e.g., for debugging/inspection). **Composition:** there is no heavyweight framework — you call `writer.accept(splitter.apply(reader.get()))`. You can run it at startup, on upload, or as a batch job. **Gotchas & when-to-use:** - **Chunk size is the highest-leverage knob.** Too large: fewer, noisier retrievals and context bloat; too small: fragments lose meaning. Tune `TokenTextSplitter` per corpus. - **Enrichers cost LLM calls** (money + latency) at ingest time — worth it when metadata materially improves retrieval/filtering. - **Idempotency:** re-running ingestion can create duplicates unless you manage IDs; Documents carry an id and stores upsert by id. - **Metadata filtering:** metadata added here (e.g., `category`, `source`) can later drive `SearchRequest` filter expressions at retrieval time. - Ingestion is offline/asynchronous relative to querying; keep it out of the request path for large corpora.
- Why is TokenTextSplitter usually necessary?Embeddings and the model's context window are size-bounded, and similarity search returns whole chunks. Splitting long documents into coherent, token-sized chunks keeps retrieval granular and relevant and avoids context overflow.
- How does calling vectorStore.add(docs) turn text into something searchable?VectorStore implements DocumentWriter; add() invokes the configured EmbeddingModel to compute a vector for each Document, then stores the vector alongside the text and metadata so later similarity searches can find it.
saying these in an interview costs you the question
- Claiming the ETL pipeline is a separate heavyweight framework rather than composed functional interfaces (Supplier/Function/Consumer)
- Thinking the writer stores raw text only, not embeddings (VectorStore.add embeds via EmbeddingModel)
- Ignoring chunking and assuming whole documents are embedded as-is
- Confusing readers (extract) with the VectorStore (load) responsibilities