skip to content

How does similaritySearch with SearchRequest work, and how do you apply metadata filters?

level: seniorimportance: must knowfreq 55%

answer

  1. SearchRequest.builder(): query/topK/threshold/filter
  2. filter DSL: == != > >= in nin && || NOT
  3. FilterExpressionBuilder -> Filter.Expression
  4. filter pushed down to store (WHERE), not post-filter
  5. same EmbeddingModel for query and docs

basics

~20 s

You build a SearchRequest with the query text, a topK (how many results), a similarityThreshold, and a filter expression on metadata. The store embeds the query, finds the nearest vectors that also match the filter, and returns them as Documents with scores.

solid answer

~40 s

similaritySearch(SearchRequest) is the full-control retrieval call. SearchRequest.builder() lets you set query(text), topK(n), similarityThreshold(0..1, drop weaker matches), and filterExpression(...). At runtime the store embeds the query with the EmbeddingModel, then does an approximate-nearest-neighbor search combined with a metadata predicate — so only documents whose metadata satisfies the filter AND are close in vector space are returned, each with a getScore(). Filters can be expressed two ways: a textual DSL string like "genre == 'drama' && year >= 2020" (parsed by FilterExpressionTextParser), or programmatically with FilterExpressionBuilder producing a Filter.Expression. The DSL supports ==, !=, >, >=, <, <=, IN, NIN, AND/OR/NOT, and grouping. Crucially the filter is pushed down to the store (a WHERE clause in pgvector), so it's efficient — not a post-filter in Java. This lets one store serve multi-tenant or category-scoped retrieval.

code

java · 25 lines
java
// Textual DSL form
List<Document> results = vectorStore.similaritySearch(
    SearchRequest.builder()
        .query("how do module boundaries work?")
        .topK(5)
        .similarityThreshold(0.7)
        .filterExpression("topic == 'modulith' && year >= 2023")
        .build());

// Programmatic builder form (equivalent, type-safe)
var b = new FilterExpressionBuilder();
Filter.Expression exp = b.and(
        b.eq("topic", "modulith"),
        b.gte("year", 2023)).build();

List<Document> results2 = vectorStore.similaritySearch(
    SearchRequest.builder()
        .query("how do module boundaries work?")
        .topK(5)
        .similarityThreshold(0.7)
        .filterExpression(exp)
        .build());

results.forEach(d ->
    System.out.println(d.getScore() + " :: " + d.getText()));

go deeper

for a junior

Know similaritySearch returns the closest documents to a query; topK limits how many.

for a middle

Explain SearchRequest's query/topK/threshold and that filters run against metadata.

for a senior

Discuss both filter forms (DSL vs FilterExpressionBuilder), push-down to the store, threshold tuning, and type/metric gotchas.

for a principal

Reason about multi-tenant isolation via filters, recall/precision tuning at scale, and how filter push-down affects store choice and index design.

## `similaritySearch` overloads `VectorStore` exposes: - `List<Document> similaritySearch(String query)` — quick path, uses store defaults for topK/threshold. - `List<Document> similaritySearch(SearchRequest request)` — full control. ## `SearchRequest` Built via `SearchRequest.builder()`: - `.query(String)` — the natural-language query. The store embeds it with the same `EmbeddingModel` used for ingestion (this is why models must match). - `.topK(int)` — max number of nearest neighbors to return (default 4). - `.similarityThreshold(double)` — a floor in [0,1]; results scoring below it are dropped. Default `SIMILARITY_THRESHOLD_ALL` (0.0) returns everything up to topK. Higher = stricter. - `.filterExpression(...)` — a metadata predicate, either a `String` DSL or a `Filter.Expression`. Results are `Document`s ordered by descending similarity, each with `getScore()` (normalized so higher = more similar, regardless of the store's underlying distance metric). ## Metadata filtering — two forms **1. Textual DSL** (portable string, parsed by `FilterExpressionTextParser`): ``` country == 'BG' && year >= 2020 genre in ['comedy','documentary'] && rating > 4 NOT(status == 'archived') ``` Supported: `==`, `!=`, `>`, `>=`, `<`, `<=`, `IN`, `NIN`, boolean `AND`/`&&`, `OR`/`||`, `NOT`, and parentheses for grouping. String literals use single quotes. **2. Programmatic** via `FilterExpressionBuilder`: ```java var b = new FilterExpressionBuilder(); Filter.Expression exp = b.and( b.eq("country", "BG"), b.gte("year", 2020)).build(); ``` Builder methods: `eq`, `ne`, `gt`, `gte`, `lt`, `lte`, `in`, `nin`, `and`, `or`, `not`. ## Push-down, not post-filter The filter is translated into the store's native query — for `PgVectorStore` it becomes a SQL `WHERE` against the JSONB metadata column combined with the ANN vector operator; for Redis it becomes a RediSearch tag/numeric filter. This means filtering is done **in the database**, so `topK` is applied *after* the filter and results are both relevant and correct. It is not a naive fetch-then-filter in Java (which would silently return fewer than topK). ## Gotchas - **Metadata keys must exist at ingest time.** Filtering on a key you never stored simply excludes those docs. - **Type sensitivity**: `year >= 2020` needs `year` stored as a number, not the string `"2020"`; mismatched types can filter everything out. - **Threshold is metric-dependent conceptually but normalized**: Spring AI normalizes score so higher is better; still, a good threshold (e.g. 0.7) is empirical per model/corpus. - **topK vs recall**: too-small topK can miss relevant chunks in RAG; too-large floods the prompt with noise and tokens. - **Empty results are valid**: a strict threshold + narrow filter can legitimately return an empty list — handle it. - **Same EmbeddingModel required**: query and stored docs must be embedded by the same model or scores are meaningless. ## When to use Use `SearchRequest` (not the string overload) whenever you need multi-tenant isolation, category scoping, freshness windows, or quality control via threshold — i.e., essentially all production RAG retrieval.

  • Is the metadata filter applied in the database or in Java after fetching?
    In the store. Spring AI translates the Filter.Expression into the store's native query (SQL WHERE on JSONB for pgvector, RediSearch filter for Redis), so filtering happens server-side and topK is applied to the filtered set — not a fetch-then-filter in application code.
  • What does similarityThreshold do and how do you pick a value?
    It drops results whose normalized similarity score is below the floor, trading recall for precision. There's no universal value — you tune it empirically per model and corpus (often ~0.6-0.8), balancing missing relevant chunks against admitting noise.

saying these in an interview costs you the question

  • Believing filters are applied in Java after fetching topK (would under-fill results)
  • Filtering on numeric metadata stored as strings
  • Assuming a fixed 'correct' similarityThreshold across all models
  • Forgetting the query must be embedded by the same model as the stored docs

context