skip to content

In Weaviate, when do you configure no vectorizer and supply vectors yourself?

level: seniorimportance: must knowfreq 62%

answer

  1. Weaviate as store, not embedder
  2. every insert and every query carries floats
  3. near_text has nothing to call
  4. dimension checked, model identity is not
  5. model swap becomes your migration

basics

~20 s

Configure no vectorizer when you must control the embedding model yourself — pinning a version, reusing embeddings computed elsewhere, or keeping an external provider out of the query path. You then supply a vector on every insert and on every search.

solid answer

~50 s

Creating the collection with `Configure.Vectorizer.none()` (or, on recent clients, `Configure.Vectors.self_provided()`) tells Weaviate it is a store, not an embedder. Every `data.insert()` must pass `vector=[...]`, and every search must use `near_vector` with a vector your own code produced — `near_text` has no module to call and fails. You take on three obligations: dimension must match what the collection was built with, the *same model and version* must be used for both import and query, and re-embedding after a model change is your migration to run. In exchange you get the things server-side vectorization cannot give — one embedding pipeline shared across several systems, your own batching and GPU scheduling, no third-party dependency on the read path, and the freedom to use a model no Weaviate module wraps. Most teams that already have an embedding service pick this; teams treating Weaviate as the whole retrieval stack pick a module.

code

python · 18 lines
python
from weaviate.classes.config import Configure, Property, DataType

client.collections.create(
    name="Doc",
    properties=[Property(name="text", data_type=DataType.TEXT)],
    vectorizer_config=Configure.Vectorizer.none(),
)

docs = client.collections.get("Doc")
docs.data.insert(
    properties={"text": "Waterproof hiking boot"},
    vector=my_encoder.encode("Waterproof hiking boot").tolist(),
)

res = docs.query.near_vector(
    near_vector=my_encoder.encode("boots for wet trails").tolist(),
    limit=5,
)

go deeper

for a junior

Know that a collection can be created with no vectorizer, and that you then pass a vector on insert and search with near_vector instead of near_text.

for a middle

Explain the full contract: vector on every write, vector on every query, fixed dimension, and near_text unavailable. Name a concrete reason to choose it, such as reusing embeddings you already compute.

for a senior

Emphasise the unenforced invariant — Weaviate checks dimension, not model identity — and describe how you defend it: pinned model versions, recorded metadata, and relevance evaluation as a regression test.

for a principal

Frame it as an architectural boundary. Owning embeddings gives you one pipeline, one model policy and no third party on the read path; the price is a re-embedding migration you must plan and budget every time the model changes.

## What "no vectorizer" means A Weaviate collection created with `Configure.Vectorizer.none()` behaves like a plain vector store. It holds objects with properties, it holds a vector per object, it builds an ANN index over those vectors, and it answers nearest-neighbour queries — but it never calls an embedding model. Everything about how the numbers were produced is your problem, and your privilege. ## The contract you accept **Every write carries a vector.** `data.insert(properties={...}, vector=[...])` — omit it and the object is stored without a vector and is invisible to vector search, silently. There is no error telling you the object will never be retrievable semantically. **Every read supplies a vector.** `query.near_vector(near_vector=[...], limit=k)` is the search path. `near_text` cannot work: there is no module to turn the string into a vector, and the call errors. Candidates often miss this and describe a collection with `none()` that still answers text queries. **Dimension is fixed.** The collection's vector dimension is established by what you insert first; a later vector of a different length is rejected. Switching from a 768-dimension model to a 1536-dimension one is not a configuration change, it is a new collection or a new named vector. **Model identity is unenforced.** This is the sharp edge. Nothing in Weaviate checks that the query vector came from the same model as the stored vectors — only that the arity matches. Embed your corpus with one model and your queries with another of the same width and every query still returns k results, ranked by a similarity that means nothing. There is no error, no warning, only quietly poor relevance. Treat the model identifier and version as part of the collection's contract: record it in collection metadata or in your deployment config, and assert it in the service that issues queries. ## Why teams choose it *Model control.* Providers update hosted models. If your relevance evaluation was run against a specific model version, you want that version pinned, and you want the pin to live in your own dependency management rather than in a database module. *Reuse.* In most real architectures the text is already being embedded — for a reranker, a classifier, a cache, another index. Letting Weaviate embed it again doubles inference spend on identical input. *Latency and blast radius.* Server-side vectorization puts a third-party HTTP call in front of every text query. With your own vectors the query path is Weaviate alone, and a provider outage degrades ingestion rather than taking search down. *Models no module wraps.* A fine-tuned in-house encoder, a multi-vector or sparse representation, a domain-specific model — if there is no module for it, self-provided vectors are the only path. *Batching economics.* Your own pipeline can batch thousands of texts per inference call, deduplicate, cache by content hash, and retry on your own schedule. Per-object server-side embedding cannot. ## Why teams choose the module instead Because it removes an entire service. For a prototype, an internal tool, or a team without embedding infrastructure, `text2vec-openai` plus an API key is a working semantic search in an afternoon, with no risk of the query and corpus models drifting apart — the module guarantees they match. That guarantee is genuinely valuable, and it is precisely the guarantee you give up. ## The migration you now own Changing models is the moment the choice bites. With self-provided vectors, upgrading means: stand up the new model, re-embed the whole corpus, write the new vectors into a new collection or a new named vector, evaluate relevance on both, cut traffic over, then drop the old. Weaviate helps with the mechanics (aliases, named vectors, replication) but the re-embedding cost and the correctness of the cutover are yours. Budget for it explicitly — a corpus of tens of millions of objects is hours of inference and a real bill. ## A hybrid worth knowing The choice is not all-or-nothing per collection. Named vectors let one collection carry, say, a module-generated vector for a general-purpose text field and a self-provided vector from a fine-tuned model, each searchable by name. That is a reasonable staging ground during a model migration: write both, compare, then retire one.

  • What happens if you embed the corpus with one model and the queries with a different model of the same dimension?
    Every query still succeeds and still returns k results — Weaviate validates dimension, not provenance. The rankings are meaningless because the two models place the same concepts in different regions of the space. There is no error to catch it, so relevance evaluation and an asserted model identifier in your query service are the only defences.
  • Can you call near_text on a collection created with Configure.Vectorizer.none()?
    No. near_text needs a module to turn the query string into a vector, and that collection has none, so the call fails. You search it with near_vector using a vector your own code produced. This is the most common surprise for people moving from a module-backed collection to self-provided vectors.
  • How do you migrate a self-provided-vector collection to a new embedding model?
    Re-embed the corpus with the new model and write it into a new target — a new collection, or a new named vector on the existing one — then evaluate relevance on both before cutting traffic over. Vectors are never converted in place, and mixing old and new vectors in one index corrupts ranking, so the cutover has to be atomic from the reader's point of view.

saying these in an interview costs you the question

  • Expecting near_text to work without a vectorizer module
  • Assuming Weaviate detects that query and corpus used different models
  • Inserting objects without a vector and wondering why search misses them
  • Thinking a model upgrade is a configuration change rather than a re-import
  • Believing dimension can be changed on an existing collection

context