skip to content

Describe Spring AI's ETL pipeline for ingesting documents into a vector store (Reader, Transformer, Writer).

level: middleimportance: must knowfreq 65%

answer

  1. Extract=Reader (Supplier), Transform=Transformer (Function), Load=Writer (Consumer)
  2. TokenTextSplitter chunks by tokens
  3. VectorStore IS a DocumentWriter -> add() embeds + stores
  4. Tika/PDF/JSON/Text readers
  5. Enrichers add metadata via LLM

basics

~10 s

You read source files into Documents with a DocumentReader, split/enrich them with a DocumentTransformer (e.g., TokenTextSplitter), then write them into the VectorStore with a DocumentWriter, which embeds and stores them for later retrieval.

solid answer

~40 s

Spring AI's ingestion is a three-stage ETL over `Document` objects. **Extract:** a `DocumentReader` (a `Supplier<List<Document>>`) reads a source — `TextReader`, `JsonReader`, `PagePdfDocumentReader`, or `TikaDocumentReader` for PDF/HTML/Office. **Transform:** a `DocumentTransformer` (a `Function<List<Document>, List<Document>>`) reshapes them — most importantly `TokenTextSplitter`, which chunks large docs so each fits the embedding/context budget and retrieval is granular; enrichers like `KeywordMetadataEnricher` and `SummaryMetadataEnricher` add metadata via the LLM. **Load:** a `DocumentWriter` (a `Consumer<List<Document>>`) persists them; `VectorStore` implements `DocumentWriter`, so `vectorStore.add(docs)` embeds each chunk with the `EmbeddingModel` and stores the vectors. You compose these with plain functional calls. Good chunking is the highest-leverage step — chunks too large blow the context and dilute relevance; too small lose meaning.

code

java · 20 lines
java
@Component
class IngestionService {
    private final VectorStore vectorStore;

    IngestionService(VectorStore vectorStore) {
        this.vectorStore = vectorStore;
    }

    void ingest(Resource pdf) {
        // Extract: read the source into Documents
        DocumentReader reader = new TikaDocumentReader(pdf);
        List<Document> docs = reader.get();

        // Transform: chunk into token-sized pieces for good retrieval
        List<Document> chunks = new TokenTextSplitter().apply(docs);

        // Load: VectorStore is a DocumentWriter -> embeds + stores each chunk
        vectorStore.add(chunks);
    }
}

go deeper

for a junior

Name the three stages and that VectorStore.add stores documents for retrieval.

for a middle

Map each stage to its functional interface, name concrete readers/transformers, and explain why chunking matters.

for a senior

Discuss chunk-size tuning, metadata for later filtering, enricher cost/latency, and idempotent re-ingestion by id.

for a principal

Treat ingestion as an offline data pipeline: batching, incremental updates, embedding-model versioning, and retrieval-quality evaluation feeding chunking choices.

Before RAG can retrieve anything, you must **ingest** your corpus. Spring AI models this as a classic **ETL (Extract, Transform, Load)** pipeline whose currency is the `Document` (text + metadata map). All three stages are just functional interfaces, so you can chain or wrap them freely. **1. Extract — `DocumentReader`** (`extends Supplier<List<Document>>`). It produces Documents from a source. Built-ins: - `TextReader` — plain text/`Resource`. - `JsonReader` — JSON, with configurable fields to extract. - `PagePdfDocumentReader` / `ParagraphPdfDocumentReader` — PDF (from `spring-ai-pdf-document-reader`). - `TikaDocumentReader` — Apache Tika, handles PDF, HTML, Word, PowerPoint, etc. (from `spring-ai-tika-document-reader`). Markdown reader also exists. Metadata like source filename/page is attached here. **2. Transform — `DocumentTransformer`** (`extends Function<List<Document>, List<Document>>`). Reshapes the list: - `TokenTextSplitter` — **the key one**: splits long Documents into chunks sized by token count (default ~800 tokens, with min-size and overlap-ish settings). Chunking matters because embeddings and the context window are bounded, and retrieval returns whole chunks — you want each chunk to be one coherent, retrievable idea. - `ContentFormatTransformer` — applies a `ContentFormatter` controlling how text+metadata render. - `KeywordMetadataEnricher` — uses the LLM to add keyword metadata. - `SummaryMetadataEnricher` — uses the LLM to add per-chunk summaries (optionally of neighboring chunks). **3. Load — `DocumentWriter`** (`extends Consumer<List<Document>>`). Persists the result: - `VectorStore` **implements `DocumentWriter`**, so calling `vectorStore.add(documents)` (or `accept(documents)`) triggers the `EmbeddingModel` to compute a vector per Document and stores vector+text+metadata. This is what makes them searchable later. - `FileDocumentWriter` — writes to a file (e.g., for debugging/inspection). **Composition:** there is no heavyweight framework — you call `writer.accept(splitter.apply(reader.get()))`. You can run it at startup, on upload, or as a batch job. **Gotchas & when-to-use:** - **Chunk size is the highest-leverage knob.** Too large: fewer, noisier retrievals and context bloat; too small: fragments lose meaning. Tune `TokenTextSplitter` per corpus. - **Enrichers cost LLM calls** (money + latency) at ingest time — worth it when metadata materially improves retrieval/filtering. - **Idempotency:** re-running ingestion can create duplicates unless you manage IDs; Documents carry an id and stores upsert by id. - **Metadata filtering:** metadata added here (e.g., `category`, `source`) can later drive `SearchRequest` filter expressions at retrieval time. - Ingestion is offline/asynchronous relative to querying; keep it out of the request path for large corpora.

  • Why is TokenTextSplitter usually necessary?
    Embeddings and the model's context window are size-bounded, and similarity search returns whole chunks. Splitting long documents into coherent, token-sized chunks keeps retrieval granular and relevant and avoids context overflow.
  • How does calling vectorStore.add(docs) turn text into something searchable?
    VectorStore implements DocumentWriter; add() invokes the configured EmbeddingModel to compute a vector for each Document, then stores the vector alongside the text and metadata so later similarity searches can find it.

saying these in an interview costs you the question

  • Claiming the ETL pipeline is a separate heavyweight framework rather than composed functional interfaces (Supplier/Function/Consumer)
  • Thinking the writer stores raw text only, not embeddings (VectorStore.add embeds via EmbeddingModel)
  • Ignoring chunking and assuming whole documents are embedded as-is
  • Confusing readers (extract) with the VectorStore (load) responsibilities

context