skip to content

What is an EmbeddingModel in Spring AI and what does it produce?

level: juniorimportance: must knowfreq 60%

answer

  1. text -> float[] vector
  2. similar meaning = close vectors
  3. embed(String) / embed(List) batch
  4. dimensions() fixed per model
  5. auto-configured from starter

basics

~10 s

EmbeddingModel is a Spring AI interface that turns text into an embedding — a list of floating-point numbers (a vector) representing the text's meaning. Similar texts get numerically similar vectors, which powers semantic search.

solid answer

~40 s

EmbeddingModel is Spring AI's abstraction over an embedding provider (OpenAI, Ollama, etc.). You call embed(String) to get a float[] vector, or embed(List<String>) to batch. Under the hood it wraps EmbeddingRequest/EmbeddingResponse and delegates to the provider's REST API via an auto-configured bean. The vector's length is fixed per model (e.g. text-embedding-3-small = 1536 dimensions), exposed via dimensions(). The core property is semantic: texts with similar meaning map to vectors that are close in space (usually by cosine similarity). You rarely use the raw numbers directly — instead you store them in a VectorStore and later run similaritySearch to find the nearest documents to a query. Spring Boot auto-configures the EmbeddingModel bean from the starter on the classpath plus API-key config; you just inject it.

code

java · 24 lines
java
@Service
class SemanticService {

    private final EmbeddingModel embeddingModel;

    SemanticService(EmbeddingModel embeddingModel) {
        this.embeddingModel = embeddingModel;
    }

    void demo() {
        // Single text -> one vector
        float[] vector = embeddingModel.embed("Spring makes Java productive");
        System.out.println("dims = " + vector.length); // e.g. 1536

        // Batch: one API call for many inputs
        List<float[]> vectors = embeddingModel.embed(
                List.of("cats", "dogs", "kotlin coroutines"));

        // Rich response with token-usage metadata
        EmbeddingResponse response =
                embeddingModel.embedForResponse(List.of("hello world"));
        var usage = response.getMetadata().getUsage();
    }
}

go deeper

for a junior

Know that EmbeddingModel turns text into a numeric vector and similar text gives similar vectors.

for a middle

Know the embed overloads, batching, dimensions(), and that Spring Boot auto-configures the bean from a starter + API key.

for a senior

Discuss EmbeddingRequest/Response, per-call options, token-usage metadata, and the re-embed-on-model-change constraint.

for a principal

Reason about cost/latency of batching at scale, dimension/model governance across environments, and coupling between model choice and store schema.

## What an embedding is An **embedding** is a fixed-length array of floating-point numbers (a **vector**) that numerically represents the *meaning* of a piece of text (or image/audio). The key property: semantically similar inputs produce vectors that are geometrically close, so you can measure relatedness with math (typically **cosine similarity**) instead of keyword matching. A model like OpenAI's `text-embedding-3-small` outputs 1536 numbers per input; that count is the **dimensionality** and is fixed for a given model. ## The `EmbeddingModel` interface `org.springframework.ai.embedding.EmbeddingModel` is Spring AI's provider-agnostic abstraction. Key methods: - `float[] embed(String text)` — one string → one vector. - `float[] embed(Document document)` — embeds a `Document`'s content. - `List<float[]> embed(List<String> texts)` — batch (one API call, far cheaper/faster than looping). - `EmbeddingResponse embedForResponse(List<String> texts)` — returns richer result incl. token-usage `Metadata`. - `EmbeddingResponse call(EmbeddingRequest request)` — the low-level call taking options (e.g. model name). - `int dimensions()` — the vector length (may trigger one probe call to discover it). ## Auto-configuration Add a starter such as `spring-ai-starter-model-openai` (or `-ollama`, `-mistral-ai`, etc.). Spring Boot then creates an `EmbeddingModel` bean from properties like `spring.ai.openai.api-key` and `spring.ai.openai.embedding.options.model`. You inject it by type. If multiple embedding starters are present you get multiple beans and must qualify. ## `EmbeddingRequest` / `EmbeddingResponse` `EmbeddingRequest(List<String> instructions, EmbeddingOptions options)` lets you override the model or other options per call. `EmbeddingResponse` holds a `List<Embedding>` (each wraps a `float[]` and its index) plus usage metadata. ## Gotchas - **Never mix models**: vectors from different models (or different dimensions) are not comparable — you must re-embed everything if you switch models. - **Batch to save money/latency**: `embed(List)` is one HTTP call; looping `embed(String)` is N calls. - **Cost & rate limits**: embedding calls hit the provider API and are billed per token. - **Dimensions must match your VectorStore column/index** — a 1536-dim model can't write into a store configured for 768. ## When to use Whenever you need semantic similarity: RAG retrieval, semantic search, clustering, deduplication, recommendation. You usually don't consume raw vectors — you hand them to a `VectorStore`.

  • Why should you batch with embed(List<String>) instead of looping embed(String)?
    Each embed call is a billed HTTP round-trip to the provider. The batch overload sends all inputs in one request, cutting latency and cost, and respecting rate limits far better than N separate calls.
  • What breaks if you change the embedding model after data is already stored?
    Vectors from the new model live in a different space (and possibly different dimensionality), so similarity scores against old vectors are meaningless. You must re-embed and re-index the entire corpus.

saying these in an interview costs you the question

  • Thinking embeddings store the literal text or are reversible back to the original words
  • Believing vectors from different models are comparable
  • Not knowing dimensionality is fixed per model
  • Calling embed in a loop instead of batching

context