skip to content

What is the VectorStore abstraction and the Document model, and how do you add data to a store?

level: middleimportance: must knowfreq 62%

answer

  1. add / delete / similaritySearch
  2. Document = id + text + metadata + vector
  3. store embeds on add (composes EmbeddingModel)
  4. metadata = filter surface
  5. swap Pg/Redis/Chroma via config

basics

~10 s

VectorStore is Spring AI's interface for saving and searching embeddings. You wrap text plus metadata in Document objects and call vectorStore.add(documents); the store embeds and persists them. Later similaritySearch finds the closest documents.

solid answer

~40 s

VectorStore is a provider-agnostic interface (implemented by PgVectorStore, RedisVectorStore, ChromaVectorStore, etc.) with three core operations: add(List<Document>), delete(...), and similaritySearch(...). The unit of storage is Document — it carries an id, the text content, and a Map<String,Object> metadata (used later for filtering), and after storage an embedding vector. When you call add, the store internally uses the configured EmbeddingModel to turn each Document's content into a vector and persists content + vector + metadata together. You never embed manually for storage — the store owns that. Spring Boot auto-configures a VectorStore bean when a store starter (e.g. spring-ai-starter-vector-store-pgvector) plus an EmbeddingModel are on the classpath. Because everything is behind one interface, you can swap Postgres for Redis or Chroma with only configuration changes, not code changes.

code

java · 29 lines
java
@Service
class IngestionService {

    private final VectorStore vectorStore;

    IngestionService(VectorStore vectorStore) {
        this.vectorStore = vectorStore; // EmbeddingModel is composed inside
    }

    void ingest() {
        List<Document> docs = List.of(
            new Document(
                "Spring Modulith enforces module boundaries.",
                Map.of("source", "docs", "topic", "modulith", "year", 2024)),
            new Document(
                "pgvector adds a vector column type to Postgres.",
                Map.of("source", "blog", "topic", "pgvector", "year", 2023))
        );

        // add() embeds each Document's text and persists content+vector+metadata
        vectorStore.add(docs);

        // Later: delete by id
        // vectorStore.delete(List.of(docs.get(0).getId()));

        List<Document> hits = vectorStore.similaritySearch("module isolation");
        hits.forEach(d -> System.out.println(d.getText() + " score=" + d.getScore()));
    }
}

go deeper

for a junior

Know Document holds text + metadata and add() saves it into the store.

for a middle

Explain the VectorStore interface (add/delete/search), that it embeds on add via a composed EmbeddingModel, and that swapping stores is config-only.

for a senior

Discuss the ETL pipeline (DocumentReader/TokenTextSplitter/VectorStore as writer), metadata as the filter surface, and id/idempotency behavior.

for a principal

Reason about ingestion architecture, chunking strategy, re-ingestion idempotency, and dimension/schema governance across store implementations.

## The `VectorStore` interface `org.springframework.ai.vectorstore.VectorStore` is Spring AI's uniform API over vector databases. It extends `DocumentWriter` (`Consumer<List<Document>>`). The operations that matter: - `void add(List<Document> documents)` — embed and persist documents. - `void delete(List<String> idList)` — delete by id. There is also `delete(Filter.Expression)` to delete by metadata predicate, and `delete(String filterExpression)` for the textual form. - `List<Document> similaritySearch(String query)` — convenience search. - `List<Document> similaritySearch(SearchRequest request)` — full-control search (topK, threshold, filters). Implementations: `PgVectorStore` (Postgres + pgvector), `RedisVectorStore`, `ChromaVectorStore`, plus Milvus, Qdrant, Weaviate, Elasticsearch, Azure, Neo4j, and an in-memory `SimpleVectorStore` for tests/prototyping. ## The `Document` model `org.springframework.ai.document.Document` is the unit of storage and retrieval: - **id**: a `String` (auto-generated if you don't set one). - **content / text**: the actual text (`getText()`). - **metadata**: `Map<String,Object>` — arbitrary key/value pairs like `source`, `category`, `year`. This is what powers metadata filtering at search time. - **embedding / score**: the vector is attached after storage; on search results, `getScore()` gives the similarity score. You build them directly (`new Document(text, metadata)`) or generate them from files using the ETL pipeline: a `DocumentReader` (e.g. `TikaDocumentReader`, `PagePdfDocumentReader`, `TextReader`) reads raw sources, a `DocumentTransformer` such as `TokenTextSplitter` chunks large text into retrieval-sized pieces, and the `VectorStore` is the `DocumentWriter` sink. ## Who does the embedding? On `add`, the `VectorStore` calls the injected `EmbeddingModel` for you — you do **not** pre-compute vectors for storage. This is why a `VectorStore` bean requires an `EmbeddingModel` bean; the store composes it. ## Auto-configuration With a starter like `spring-ai-starter-vector-store-pgvector` and store properties (`spring.ai.vectorstore.pgvector.*`), Boot wires the `VectorStore` bean. `initialize-schema` can auto-create the table/index for stores like pgvector. ## Gotchas - **Metadata is your filter surface** — anything you might filter or delete by later must be in `metadata` at insert time; you can't filter on content you didn't store as metadata. - **Chunk before adding**: embedding a whole 50-page PDF as one Document gives poor retrieval; split with `TokenTextSplitter`. - **ids and idempotency**: re-adding the same id may duplicate or upsert depending on the store — set stable ids if you re-ingest. - **Dimension match**: the store's vector column/index dimension must equal the EmbeddingModel's `dimensions()`. ## When to use Use a `VectorStore` whenever you need durable semantic retrieval — the retrieval half of RAG, semantic search, or recommendation. Prototype with `SimpleVectorStore`, then swap to Postgres/Redis/Chroma via config.

  • Do you call EmbeddingModel yourself before vectorStore.add()?
    No. The VectorStore composes the EmbeddingModel and embeds each Document's content internally on add(). You only pre-embed manually if you're doing something custom outside the store.
  • Why chunk documents with TokenTextSplitter before adding them?
    One giant Document produces one vector that blurs many topics, hurting retrieval precision and possibly exceeding the model's token limit. Splitting into passage-sized chunks yields focused vectors that match specific queries.

saying these in an interview costs you the question

  • Thinking you must embed text yourself before calling add()
  • Storing filterable attributes only inside the text instead of in metadata
  • Adding whole large files as a single Document with no chunking
  • Assuming re-adding the same content is always deduplicated

context