skip to content

Which parts of a search request should never be answered by vector similarity?

level: juniorimportance: must knowfreq 55%

answer

  1. similarity ranks, predicates decide
  2. no notion of less-than or count
  3. filter during retrieval, not after
  4. permissions are never a soft signal
  5. numbers written into text do not compare

basics

~20 s

Anything with a truth condition: numeric and date constraints, sorting, counting, permission scoping, and exact identifier lookups. Vector similarity ranks by how alike two texts are, and "alike" cannot express "price under 800,000" or "only records this user may see".

solid answer

~50 s

Embeddings encode what text is *about*, so they are the right tool for the fuzzy, meaning-bearing part of a request and the wrong tool for anything with a hard boundary. A real-estate query like "walkable neighbourhood near good schools, three beds, under $800k, sorted by newest" has two halves. The first half has no keyword form at all and is exactly what semantic search is for. The second half — bedroom count, a price ceiling, an ordering, and the requirement that only listed properties appear — are constraints and operations, and similarity has no notion of *less than*, of *count*, or of *authorised*. A vector that is "near" a $900k listing is still a $900k listing. The standard design is to parse the request, push constraints into structured metadata filters applied by the store, use vector similarity only for the semantic phrase, and apply sorting and permissions in the query engine, never in the ranking.

go deeper

for a junior

Be able to say that vector search finds text that means something similar, and that hard requirements — a price limit, a date range, who is allowed to see a record — must be handled by ordinary filters on stored fields.

for a middle

Explain why: similarity is continuous and always returns something, while a constraint is a binary predicate with a sharp edge. Describe splitting a request into filters plus a semantic phrase, and why the filter must be applied during retrieval rather than after.

for a senior

Bring the operational angle — post-filtering that empties result pages, tenant isolation enforced from the authenticated identity, and the fact that very selective filters interact badly with approximate indexes and need deliberate handling.

for a principal

Own the boundary as an architectural rule: the vector store is one participant in a query plan, not the query engine. Decide where predicate evaluation, ordering and authorisation live across the search stack so that correctness never depends on a ranking score.

## Similarity is not a predicate A vector search answers one question: which stored vectors are closest to this one. Closeness is a continuous, relative notion — it always returns *something*, and there is no answer that means "nothing qualifies". Constraints are the opposite: they are binary predicates with a sharp edge. Under $800,000 is either true or false. A listing at $801,000 is semantically indistinguishable from one at $799,000, and any embedding will place them side by side, because in meaning-space they *are* the same thing. Asking similarity to enforce a threshold is asking a continuous measure to do a discrete job. The same argument rules out several other operations: - **Sorting.** Ranking by relevance and ordering by price or date are different orderings over the same set. Similarity cannot produce the second one; it can only be overridden by it. - **Counting and aggregation.** "How many three-bed listings came on the market this week?" needs a complete, exact set, and approximate nearest-neighbour search deliberately does not guarantee completeness. - **Exact identifier lookup.** A reference number or an internal ID denotes one record. You want a key lookup, not the ten nearest strings. - **Access control.** Whether a user may see a record is a hard authorisation decision. Encoding it as a similarity signal is both wrong and a security defect, because "close enough" leaks. - **Recency and freshness rules.** "Only listings from the last 30 days" is a range predicate, not a shade of meaning. ## The split-the-query design The production pattern is to treat a search request as a compound object and route each part to the mechanism that can answer it: 1. **Extract the structured part.** Either from explicit UI controls — a price slider, a bedroom dropdown, a date range — or by parsing a natural-language query into filters. UI controls are strictly more reliable and are why filter panels have not disappeared. 2. **Store those attributes as metadata beside every vector.** Price, bed count, listing date, region, tenant, visibility. This is why vector stores universally support attaching structured metadata to a record. 3. **Apply the filters as part of retrieval**, not after it. Retrieving the nearest 50 vectors and then discarding those over budget is the classic bug: if 48 of the 50 are too expensive you return two results, and if all 50 are, you return an empty page while thousands of matching listings exist. The store must restrict the search to the qualifying subset. 4. **Use similarity only for the residual phrase** — the part with no keyword or predicate form, such as "walkable neighbourhood near good schools", which is genuinely a meaning query and is where semantic search earns its place. 5. **Apply ordering and permissions in the query engine**, deterministically, after or during retrieval — never as a soft ranking signal. ## Why the mistake is so common Semantic search demos beautifully. Typing a whole sentence and getting sensible results feels like the system understood the sentence, including the numbers in it. It did not: it noticed that the sentence is *about* budget-conscious family housing. In evaluation this looks fine, because plausible results are returned and nobody checks the prices. It surfaces in production as "the filter doesn't work", which is exactly what is happening. A second version of the mistake is trying to fix it by writing the constraint into the indexed text — appending "price: 795000" to each document so the query can "match" it. Numbers embedded as text do not support comparison; 795000 and 7950000 are similar strings and the model has no arithmetic. This never works and it wastes a rebuild to discover. ## Saying so in an interview The strong answer states the boundary as a principle rather than a list: **similarity ranks, predicates decide.** Then give the concrete shape — parse the request, push predicates to the store as metadata filters applied during the search, keep vector similarity for the meaning-bearing remainder, and handle sorting, counting and authorisation in the query engine. Being able to draw that line is one of the clearest signals that someone has built search rather than only read about it. It is also worth naming the honest cost: highly selective filters interact badly with approximate indexes, because restricting to a tiny subset can force the search to scan far more of the index or to miss qualifying items. That interaction is a real engineering concern in filtered vector search and is worth flagging even when the fix belongs to the store's configuration.

  • Why is applying a price filter after retrieving the top 50 vectors a bug rather than an optimisation?
    Because the filter can eliminate most or all of the retrieved set. If 48 of the 50 nearest listings exceed the budget you show two results, and if all of them do you show an empty page, even though thousands of qualifying listings exist — they simply were not among the nearest 50. The filter has to constrain the search itself so that the nearest neighbours are drawn from the qualifying subset.
  • Can you not just include the numeric attributes in the embedded text so the model matches them?
    No. Embeddings have no arithmetic and no ordering over numerals; as text, 795000 and 7950000 are similar strings. The model can at best learn a fuzzy association between price ranges and phrasing, which produces confidently wrong results near any boundary. Numeric attributes belong in structured metadata where comparison operators actually exist.
  • Where does multi-tenant isolation belong in a vector search pipeline?
    In the retrieval predicate, enforced by the store, and derived from the authenticated request rather than from anything the user typed. A tenant identifier stored as metadata on every vector and applied as a mandatory filter is the correct shape. Filtering after retrieval, or relying on tenant text making documents dissimilar, is a data-leak waiting to happen.

saying these in an interview costs you the question

  • Thinks a numeric constraint can be expressed as semantic similarity
  • Filters results after retrieval instead of during it
  • Writes prices or dates into the embedded text to make them matchable
  • Uses a similarity threshold as an access-control mechanism
  • Expects nearest-neighbour search to return exact counts

context