When do Qdrant named vectors beat separate collections per embedding?
answer
- several embeddings, one point
- own size and metric per name
- id and payload shared once
- pick one name per query
- lifecycle decides, not modality
basics
~20 sNamed vectors let one Qdrant point carry several embeddings under different names, each with its own size and distance, sharing one id and one payload. Use them when the same entity is searched by different representations; use separate collections when the data has different lifecycles.
solid answer
~50 sInstead of a single `VectorParams`, you pass a dict: `vectors_config={"text": models.VectorParams(size=768, distance=models.Distance.COSINE), "image": models.VectorParams(size=512, distance=models.Distance.DOT)}`. A point then supplies a dict of vectors, and queries select one with `using="image"` in `client.query_points`. The win is that id and payload are stored once: one product, one metadata record, retrievable by its description or by its picture, with the same filters and no cross-collection join. The costs are real too — each named vector builds and holds its own index, so memory adds up, and you can search only one named vector per query. Prefer separate collections when the embeddings have genuinely different lifecycles: different refresh cadence, different retention, or one that you expect to rebuild for a model upgrade without touching the other. Since Qdrant 1.7 a point may omit some named vectors, so sparsely populated modalities do not force you to fabricate placeholders.
code
python · 29 linesfrom qdrant_client import QdrantClient, models
client = QdrantClient(url="http://localhost:6333")
client.create_collection(
collection_name="products",
vectors_config={
"text": models.VectorParams(size=768, distance=models.Distance.COSINE),
"image": models.VectorParams(size=512, distance=models.Distance.DOT),
},
)
client.upsert(
collection_name="products",
points=[
models.PointStruct(
id=1,
vector={"text": [0.1] * 768, "image": [0.2] * 512},
payload={"sku": "A-100", "in_stock": True},
)
],
)
hits = client.query_points(
collection_name="products",
query=[0.2] * 512,
using="image",
limit=5,
)go deeper
Know that vectors_config can take a dict of names, so one point can hold more than one embedding, and that a query names which vector to search.
Explain that each named vector has its own size and distance while id and payload are shared, and that queries target exactly one name via using.
Argue the tradeoff from operations: shared payload and no id synchronization versus per-name index memory and a collection-wide rebuild when one representation's dimension changes.
Own this as a schema commitment — decide up front whether these representations will be versioned, retained and scaled as one unit, because that choice sets the cost of every future encoder upgrade.
## What named vectors are By default a collection stores one vector per point. Named vectors change that: `vectors_config` becomes a mapping from name to `VectorParams`, and each name gets its own dimensionality and its own distance metric. ``` vectors_config={ "text": VectorParams(size=768, distance=Distance.COSINE), "image": VectorParams(size=512, distance=Distance.DOT), } ``` A point in that collection carries `vector={"text": [...], "image": [...]}` alongside its single id and single payload. At query time you pick which one to search against — in the modern client, `client.query_points(collection_name="products", query=vec, using="image", limit=10)`. That is the whole mechanism. What matters in an interview is when you reach for it. ## The case for one collection The shared parts are the point: **one id, one payload.** If the same real-world entity is retrievable through several representations, splitting it across collections means duplicating its metadata and reconciling ids by hand. Concrete cases where named vectors win: - **Multimodal entities.** A product with a description embedding and a photo embedding. Search by text or by image, get back the same product record with the same price, stock and category fields. - **Multiple text encoders.** A cheap fast model and an expensive accurate one over the same documents, so you can route by query importance, or A/B them against live traffic without maintaining two ingest pipelines and two metadata copies. - **Different granularities.** A title embedding and a full-body embedding on the same document, chosen per query type. - **Late-interaction models.** A named vector can be declared as a multivector with `multivector_config=models.MultiVectorConfig(comparator=models.MultiVectorComparator.MAX_SIM)`, letting one point hold a matrix of token vectors scored by max-similarity. Sparse representations get their own `sparse_vectors_config` parameter alongside the dense one. Operationally the shared payload has another benefit: metadata is written once. Updating a product's price is one `set_payload` call, not one per collection with the risk of divergence. ## The case for separate collections Named vectors are not free, and the deciding question is usually lifecycle rather than modality. - **Independent rebuilds.** If you expect to re-embed the text side for a model upgrade while leaving images alone, separate collections let you build the replacement and swap it behind an alias. With named vectors in one collection, a dimension change on one name means recreating the whole collection, including the vectors that did not change. - **Independent retention or deletion.** Different TTLs, different tenants' data, or one representation that is regenerated nightly while the other is stable. - **Independent scaling.** Separate collections can be sized, sharded and hosted differently. Named vectors share the collection's configuration and its nodes. - **Very different cardinality.** If one representation exists for 2% of your entities, a separate small collection may be tidier than a mostly-empty named vector — though optional vectors soften this. ## Optional vectors A point does not have to provide every configured named vector. You can upsert a point with only `"text"` populated and add `"image"` later with `client.update_vectors`. Searching `using="image"` simply will not return points that lack an image vector. This matters for real catalogues where coverage of a modality is partial — you neither block ingestion nor pollute the index with zero vectors, which would otherwise cluster spuriously and surface as bogus nearest neighbours. ## Costs to state plainly Each named vector maintains its **own** index structure. Two named vectors is roughly two indexes' worth of memory and two indexes' worth of build work at ingest — the point is shared, the vector storage is not. So the payload is deduplicated but the expensive part is not, and "I'll just add a third encoder" is a capacity decision. You also search one named vector per query. There is no single call that scores against `"text"` and `"image"` simultaneously and blends them; combining representations means issuing queries and fusing results yourself, or using the client's prefetch-and-rerank query composition. Candidates who assume named vectors automatically give them multimodal fusion have the wrong mental model. ## A decision rule Ask two questions. *Do these vectors describe the same entity, with the same metadata and the same access control?* If no, separate collections. *Will they be rebuilt, retained and scaled together?* If no, separate collections — the migration pain dominates. When both answers are yes, named vectors are the cleaner model, and they save you an entire class of id-synchronization bugs.
- Can one query score against two named vectors at once and blend the scores?No — a query selects a single named vector with `using`. To combine representations you run more than one retrieval and fuse the results yourself, or compose a query that prefetches candidates with one vector and reranks them with another. Expecting automatic multimodal fusion from named vectors alone is the common misreading.
- What does adding a third named vector cost you?Roughly another index's worth of memory and another index's worth of build work on every ingest, because each named vector maintains its own structure. What you do not pay twice is id and payload storage. Treat each additional name as a capacity decision, not a free schema field.
- You need to upgrade only the text encoder to a new dimension. How does that play out with named vectors?Badly, compared with separate collections. Vector size is fixed per name at creation, so changing the text dimension means recreating the collection and reloading everything, including the image vectors that did not change. If you foresee independent rebuild cycles, that is the argument for splitting them up front.
saying these in an interview costs you the question
- Thinks all named vectors in a collection must share one dimension
- Expects a single query to search several named vectors and merge the scores
- Believes every point must supply every configured named vector
- Assumes named vectors save index memory, not just payload storage
- Chooses named vectors purely on modality, ignoring rebuild and retention differences