skip to content

What problem do Weaviate's named vectors solve, and what do they cost?

level: seniorimportance: should knowfreq 44%

answer

  1. one object, several embedding spaces
  2. aspects diluted by concatenation
  3. query names the space it searches
  4. each name is another index to build
  5. cost multiplies, properties do not

basics

~20 s

Named vectors let one object carry several independent embeddings — different properties, different models, or different modalities — each with its own configuration, chosen at query time by name. The cost is one stored vector and one ANN index per name.

solid answer

~50 s

Instead of a single vector per object, you declare a list of named vector configurations on the collection, each with its own vectorizer and its own `source_properties`. An object then stores one vector per name, and a query picks which space to search with `target_vector="..."`. This solves problems a single embedding handles badly: a product whose title and long review should be searchable separately, a document that needs both a general-purpose model and a fine-tuned one during a migration, or a record with text and image embeddings side by side. Some names can be module-generated while others are self-provided, so a hybrid pipeline is expressible in one collection. The cost is real and multiplies: each name means another embedding call at import, another vector's worth of memory and disk, and another HNSW graph to build and hold. Three named vectors on a large collection is roughly three times the vector-side footprint, so add names because a query needs that space, not because the property exists.

code

python · 24 lines
python
from weaviate.classes.config import Configure, Property, DataType

client.collections.create(
    name="Product",
    properties=[
        Property(name="title", data_type=DataType.TEXT),
        Property(name="review", data_type=DataType.TEXT),
    ],
    vectorizer_config=[
        Configure.NamedVectors.text2vec_openai(
            name="title_vec", source_properties=["title"]
        ),
        Configure.NamedVectors.text2vec_openai(
            name="review_vec", source_properties=["review"]
        ),
    ],
)

products = client.collections.get("Product")
res = products.query.near_text(
    query="waterproof hiking boot",
    target_vector="title_vec",
    limit=5,
)

go deeper

for a junior

Know that a Weaviate object can hold more than one vector, that each has a name, and that a query picks one with target_vector.

for a middle

Explain why a single concatenated embedding dilutes short fields, and how per-name source_properties fixes it. State that each name adds its own vector and its own index.

for a senior

Quantify the cost — an embedding call, a vector and an ANN graph per name — and use named vectors deliberately, for example to run an incumbent and a candidate model over the same corpus during a migration.

for a principal

Own the boundary decision. Named vectors are for aspects of one entity; separate collections are for entities with different lifecycles and cardinalities. Set a policy so names are added for a query that needs them, not per property.

## The single-vector limitation By default an object gets one vector, built from everything vectorizable about it. That works when the object is essentially one piece of prose. It works badly when an object has several distinct textual aspects. Concatenating a 12-word product title with a 900-word review and embedding the result produces a vector dominated by the review — a title-shaped query retrieves poorly, because the title's contribution was averaged away. The same problem shows up with a document that has an abstract and a body, or a person record with a bio and a list of skills. The usual workarounds are bad. Splitting into separate collections duplicates properties and breaks filtering across aspects. Storing the object twice with different vectorizable sets doubles writes and creates a deduplication problem at read time. ## What named vectors change A collection can declare several vector configurations, each with a name, its own vectorizer, and its own set of source properties. One object then holds one vector per name. Nothing about properties, filters or multi-tenancy changes — this is purely about how many spaces the object lives in. At query time you name the space: `target_vector="title_vec"` sends the search into that index only. The filter, the limit and the returned properties behave exactly as before; only the geometry differs. ## The uses that justify it **Aspect separation.** Title versus body, summary versus full text, name versus description. Each aspect gets an embedding that is not diluted by the others, and the application chooses which aspect a given query should match. **Model migration.** Run the incumbent model and the candidate model as two names over the same corpus. Both are populated by the same import, so you can evaluate relevance on identical data, shadow-read the candidate, and cut over by changing one argument rather than rebuilding a collection. **Mixed provenance.** One name can be produced by a module while another is self-provided from your own fine-tuned encoder. That is the practical way to adopt a custom model incrementally without giving up the convenience of server-side embedding for the general case. **Multiple modalities.** A record with both descriptive text and an image can carry a text embedding and an image-capable embedding under different names, letting the application search whichever matches the user's input. **Different index settings per space.** Because each name carries its own configuration, a rarely-queried aspect can use a cheaper index or on-disk storage while the hot aspect stays optimised, rather than one setting covering everything. ## The cost, stated plainly Every name multiplies the vector side of the collection: - **Import work.** One embedding call per name per object. Three names is three times the inference spend and roughly three times the import wall-clock, assuming the embedding backend is the bottleneck — which it usually is. - **Storage and memory.** One float array per name per object, and one ANN graph per name. For an HNSW index the graph itself is a substantial fraction of the footprint, not a rounding error on top of the vectors. - **Build time.** Each graph is constructed independently, so index build cost scales with the number of names too. - **Operational surface.** More configuration to get right, more places for a dimension mismatch, and a query bug now includes "searched the wrong target_vector" as a failure mode — which returns plausible-looking but wrong results rather than an error. Properties are cheap; named vectors are not. The discipline is to add a name when a query genuinely needs to search that space, and to delete names that no query targets. ## Writing self-provided named vectors When a name is self-provided, the insert passes a mapping from name to vector rather than a single list, so one write can populate several spaces at once — including a mix where some names are filled by modules and others by your own values. ## Choosing between named vectors and separate collections Use named vectors when the aspects belong to the *same entity*: same id, same properties, same filters, same tenant. Use separate collections when the things being embedded are genuinely different entities with different lifecycles — chunks versus documents, for instance, where the cardinalities differ by an order of magnitude and you want to delete or re-chunk one without touching the other.

  • What goes wrong if a query omits target_vector on a collection with several named vectors?
    There is no single default space to search, so the request cannot be resolved and fails. The instructive part is that the alternative — silently picking one — would be worse: you would get well-formed results from the wrong geometry with no signal that anything was wrong. Treat target_vector as a required argument in any client wrapper you write.
  • How would you use named vectors to migrate to a new embedding model without downtime?
    Add the candidate model as a second name so a single import populates both spaces over identical data. Evaluate relevance on both, shadow-read the candidate in production to compare rankings under real traffic, then switch target_vector on the read path. Once you are satisfied, drop the old name to reclaim its index and storage.
  • When would you prefer separate collections over named vectors?
    When the things being embedded are different entities rather than aspects of one — document-level records and their chunks, for example, where cardinality differs by orders of magnitude and you want independent lifecycles for re-chunking, deletion and retention. Named vectors assume one id, one property set and one filter surface.

saying these in an interview costs you the question

  • Thinking named vectors are free because the object is stored once
  • Assuming a query without target_vector searches all spaces and merges results
  • Adding a named vector per property by default
  • Believing all named vectors must use the same vectorizer module
  • Expecting named vectors to change filtering or tenancy behaviour

context