skip to content

Embeddings & Vector Stores

An embedding model turns text into vectors, and the VectorStore abstraction stores and searches them by similarity with metadata filters. Interviewers ask why you would filter on metadata as well as similarity, since relevance alone is rarely enough.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

5

What is an EmbeddingModel in Spring AI and what does it produce?

level: juniorimportance: must knowfreq 60%

answer

  1. text -> float[] vector
  2. similar meaning = close vectors
  3. embed(String) / embed(List) batch
  4. dimensions() fixed per model
  5. auto-configured from starter

basics

~10 s

EmbeddingModel is a Spring AI interface that turns text into an embedding — a list of floating-point numbers (a vector) representing the text's meaning. Similar texts get numerically similar vectors, which powers semantic search.

solid answer

~40 s

EmbeddingModel is Spring AI's abstraction over an embedding provider (OpenAI, Ollama, etc.). You call embed(String) to get a float[] vector, or embed(List<String>) to batch. Under the hood it wraps EmbeddingRequest/EmbeddingResponse and delegates to the provider's REST API via an auto-configured bean. The vector's length is fixed per model (e.g. text-embedding-3-small = 1536 dimensions), exposed via dimensions(). The core property is semantic: texts with similar meaning map to vectors that are close in space (usually by cosine similarity). You rarely use the raw numbers directly — instead you store them in a VectorStore and later run similaritySearch to find the nearest documents to a query. Spring Boot auto-configures the EmbeddingModel bean from the starter on the classpath plus API-key config; you just inject it.

code

java · 24 lines
java
@Service
class SemanticService {

    private final EmbeddingModel embeddingModel;

    SemanticService(EmbeddingModel embeddingModel) {
        this.embeddingModel = embeddingModel;
    }

    void demo() {
        // Single text -> one vector
        float[] vector = embeddingModel.embed("Spring makes Java productive");
        System.out.println("dims = " + vector.length); // e.g. 1536

        // Batch: one API call for many inputs
        List<float[]> vectors = embeddingModel.embed(
                List.of("cats", "dogs", "kotlin coroutines"));

        // Rich response with token-usage metadata
        EmbeddingResponse response =
                embeddingModel.embedForResponse(List.of("hello world"));
        var usage = response.getMetadata().getUsage();
    }
}

go deeper

for a junior

Know that EmbeddingModel turns text into a numeric vector and similar text gives similar vectors.

for a middle

Know the embed overloads, batching, dimensions(), and that Spring Boot auto-configures the bean from a starter + API key.

for a senior

Discuss EmbeddingRequest/Response, per-call options, token-usage metadata, and the re-embed-on-model-change constraint.

for a principal

Reason about cost/latency of batching at scale, dimension/model governance across environments, and coupling between model choice and store schema.

## What an embedding is An **embedding** is a fixed-length array of floating-point numbers (a **vector**) that numerically represents the *meaning* of a piece of text (or image/audio). The key property: semantically similar inputs produce vectors that are geometrically close, so you can measure relatedness with math (typically **cosine similarity**) instead of keyword matching. A model like OpenAI's `text-embedding-3-small` outputs 1536 numbers per input; that count is the **dimensionality** and is fixed for a given model. ## The `EmbeddingModel` interface `org.springframework.ai.embedding.EmbeddingModel` is Spring AI's provider-agnostic abstraction. Key methods: - `float[] embed(String text)` — one string → one vector. - `float[] embed(Document document)` — embeds a `Document`'s content. - `List<float[]> embed(List<String> texts)` — batch (one API call, far cheaper/faster than looping). - `EmbeddingResponse embedForResponse(List<String> texts)` — returns richer result incl. token-usage `Metadata`. - `EmbeddingResponse call(EmbeddingRequest request)` — the low-level call taking options (e.g. model name). - `int dimensions()` — the vector length (may trigger one probe call to discover it). ## Auto-configuration Add a starter such as `spring-ai-starter-model-openai` (or `-ollama`, `-mistral-ai`, etc.). Spring Boot then creates an `EmbeddingModel` bean from properties like `spring.ai.openai.api-key` and `spring.ai.openai.embedding.options.model`. You inject it by type. If multiple embedding starters are present you get multiple beans and must qualify. ## `EmbeddingRequest` / `EmbeddingResponse` `EmbeddingRequest(List<String> instructions, EmbeddingOptions options)` lets you override the model or other options per call. `EmbeddingResponse` holds a `List<Embedding>` (each wraps a `float[]` and its index) plus usage metadata. ## Gotchas - **Never mix models**: vectors from different models (or different dimensions) are not comparable — you must re-embed everything if you switch models. - **Batch to save money/latency**: `embed(List)` is one HTTP call; looping `embed(String)` is N calls. - **Cost & rate limits**: embedding calls hit the provider API and are billed per token. - **Dimensions must match your VectorStore column/index** — a 1536-dim model can't write into a store configured for 768. ## When to use Whenever you need semantic similarity: RAG retrieval, semantic search, clustering, deduplication, recommendation. You usually don't consume raw vectors — you hand them to a `VectorStore`.

  • Why should you batch with embed(List<String>) instead of looping embed(String)?
    Each embed call is a billed HTTP round-trip to the provider. The batch overload sends all inputs in one request, cutting latency and cost, and respecting rate limits far better than N separate calls.
  • What breaks if you change the embedding model after data is already stored?
    Vectors from the new model live in a different space (and possibly different dimensionality), so similarity scores against old vectors are meaningless. You must re-embed and re-index the entire corpus.

saying these in an interview costs you the question

  • Thinking embeddings store the literal text or are reversible back to the original words
  • Believing vectors from different models are comparable
  • Not knowing dimensionality is fixed per model
  • Calling embed in a loop instead of batching

context

open as a page

What is the VectorStore abstraction and the Document model, and how do you add data to a store?

level: middleimportance: must knowfreq 62%

basics

~10 s

VectorStore is Spring AI's interface for saving and searching embeddings. You wrap text plus metadata in Document objects and call vectorStore.add(documents); the store embeds and persists them. Later similaritySearch finds the closest documents.

open as a page

How does similaritySearch with SearchRequest work, and how do you apply metadata filters?

level: seniorimportance: must knowfreq 55%

basics

~20 s

You build a SearchRequest with the query text, a topK (how many results), a similarityThreshold, and a filter expression on metadata. The store embeds the query, finds the nearest vectors that also match the filter, and returns them as Documents with scores.

open as a page

How do you choose between VectorStore implementations like PgVector, Redis, and Chroma, and what configuration matters?

level: seniorimportance: should knowfreq 40%

basics

~20 s

All implement the same VectorStore interface, so code stays the same — you pick based on your existing infrastructure. Use PgVector if you already run Postgres, Redis if you need speed and already run Redis, Chroma for a lightweight AI-focused store. Key config: dimensions, distance metric, and index type.

open as a page

How do embeddings and a VectorStore fit into a RAG retrieval pipeline, and what are the main failure modes to design around?

level: principalimportance: should knowfreq 42%

basics

~20 s

In RAG you ingest documents (read, chunk, embed, store), then at query time embed the user's question, run similaritySearch to fetch the most relevant chunks, and stuff them into the prompt as context. Main risks: bad chunking, model mismatch, weak filtering, and stale data.

open as a page