What is the VectorStore abstraction and the Document model, and how do you add data to a store?
answer
- add / delete / similaritySearch
- Document = id + text + metadata + vector
- store embeds on add (composes EmbeddingModel)
- metadata = filter surface
- swap Pg/Redis/Chroma via config
basics
~10 sVectorStore is Spring AI's interface for saving and searching embeddings. You wrap text plus metadata in Document objects and call vectorStore.add(documents); the store embeds and persists them. Later similaritySearch finds the closest documents.
solid answer
~40 sVectorStore is a provider-agnostic interface (implemented by PgVectorStore, RedisVectorStore, ChromaVectorStore, etc.) with three core operations: add(List<Document>), delete(...), and similaritySearch(...). The unit of storage is Document — it carries an id, the text content, and a Map<String,Object> metadata (used later for filtering), and after storage an embedding vector. When you call add, the store internally uses the configured EmbeddingModel to turn each Document's content into a vector and persists content + vector + metadata together. You never embed manually for storage — the store owns that. Spring Boot auto-configures a VectorStore bean when a store starter (e.g. spring-ai-starter-vector-store-pgvector) plus an EmbeddingModel are on the classpath. Because everything is behind one interface, you can swap Postgres for Redis or Chroma with only configuration changes, not code changes.
code
java · 29 lines@Service
class IngestionService {
private final VectorStore vectorStore;
IngestionService(VectorStore vectorStore) {
this.vectorStore = vectorStore; // EmbeddingModel is composed inside
}
void ingest() {
List<Document> docs = List.of(
new Document(
"Spring Modulith enforces module boundaries.",
Map.of("source", "docs", "topic", "modulith", "year", 2024)),
new Document(
"pgvector adds a vector column type to Postgres.",
Map.of("source", "blog", "topic", "pgvector", "year", 2023))
);
// add() embeds each Document's text and persists content+vector+metadata
vectorStore.add(docs);
// Later: delete by id
// vectorStore.delete(List.of(docs.get(0).getId()));
List<Document> hits = vectorStore.similaritySearch("module isolation");
hits.forEach(d -> System.out.println(d.getText() + " score=" + d.getScore()));
}
}go deeper
Know Document holds text + metadata and add() saves it into the store.
Explain the VectorStore interface (add/delete/search), that it embeds on add via a composed EmbeddingModel, and that swapping stores is config-only.
Discuss the ETL pipeline (DocumentReader/TokenTextSplitter/VectorStore as writer), metadata as the filter surface, and id/idempotency behavior.
Reason about ingestion architecture, chunking strategy, re-ingestion idempotency, and dimension/schema governance across store implementations.
## The `VectorStore` interface `org.springframework.ai.vectorstore.VectorStore` is Spring AI's uniform API over vector databases. It extends `DocumentWriter` (`Consumer<List<Document>>`). The operations that matter: - `void add(List<Document> documents)` — embed and persist documents. - `void delete(List<String> idList)` — delete by id. There is also `delete(Filter.Expression)` to delete by metadata predicate, and `delete(String filterExpression)` for the textual form. - `List<Document> similaritySearch(String query)` — convenience search. - `List<Document> similaritySearch(SearchRequest request)` — full-control search (topK, threshold, filters). Implementations: `PgVectorStore` (Postgres + pgvector), `RedisVectorStore`, `ChromaVectorStore`, plus Milvus, Qdrant, Weaviate, Elasticsearch, Azure, Neo4j, and an in-memory `SimpleVectorStore` for tests/prototyping. ## The `Document` model `org.springframework.ai.document.Document` is the unit of storage and retrieval: - **id**: a `String` (auto-generated if you don't set one). - **content / text**: the actual text (`getText()`). - **metadata**: `Map<String,Object>` — arbitrary key/value pairs like `source`, `category`, `year`. This is what powers metadata filtering at search time. - **embedding / score**: the vector is attached after storage; on search results, `getScore()` gives the similarity score. You build them directly (`new Document(text, metadata)`) or generate them from files using the ETL pipeline: a `DocumentReader` (e.g. `TikaDocumentReader`, `PagePdfDocumentReader`, `TextReader`) reads raw sources, a `DocumentTransformer` such as `TokenTextSplitter` chunks large text into retrieval-sized pieces, and the `VectorStore` is the `DocumentWriter` sink. ## Who does the embedding? On `add`, the `VectorStore` calls the injected `EmbeddingModel` for you — you do **not** pre-compute vectors for storage. This is why a `VectorStore` bean requires an `EmbeddingModel` bean; the store composes it. ## Auto-configuration With a starter like `spring-ai-starter-vector-store-pgvector` and store properties (`spring.ai.vectorstore.pgvector.*`), Boot wires the `VectorStore` bean. `initialize-schema` can auto-create the table/index for stores like pgvector. ## Gotchas - **Metadata is your filter surface** — anything you might filter or delete by later must be in `metadata` at insert time; you can't filter on content you didn't store as metadata. - **Chunk before adding**: embedding a whole 50-page PDF as one Document gives poor retrieval; split with `TokenTextSplitter`. - **ids and idempotency**: re-adding the same id may duplicate or upsert depending on the store — set stable ids if you re-ingest. - **Dimension match**: the store's vector column/index dimension must equal the EmbeddingModel's `dimensions()`. ## When to use Use a `VectorStore` whenever you need durable semantic retrieval — the retrieval half of RAG, semantic search, or recommendation. Prototype with `SimpleVectorStore`, then swap to Postgres/Redis/Chroma via config.
- Do you call EmbeddingModel yourself before vectorStore.add()?No. The VectorStore composes the EmbeddingModel and embeds each Document's content internally on add(). You only pre-embed manually if you're doing something custom outside the store.
- Why chunk documents with TokenTextSplitter before adding them?One giant Document produces one vector that blurs many topics, hurting retrieval precision and possibly exceeding the model's token limit. Splitting into passage-sized chunks yields focused vectors that match specific queries.
saying these in an interview costs you the question
- Thinking you must embed text yourself before calling add()
- Storing filterable attributes only inside the text instead of in metadata
- Adding whole large files as a single Document with no chunking
- Assuming re-adding the same content is always deduplicated