skip to content

How do you choose a Weaviate vectorization strategy for a production workload?

level: principalimportance: should knowfreq 38%

answer

  1. who owns the model, and where inference runs
  2. hosted, self-hosted, self-provided
  3. the matching guarantee you give up
  4. residency and existing pipeline decide fastest
  5. named vectors make it a two-way door

basics

~20 s

Decide who owns the embedding model. A hosted module is fastest to ship but puts a third party in the query path; a self-hosted inference container keeps data and cost in-house but adds an operated service; self-provided vectors give full model control at the price of owning re-embedding.

solid answer

~60 s

There are three shapes and the choice is mostly organisational. **Hosted module** (`text2vec-openai` and peers): no infrastructure, query and corpus models are guaranteed to match, but every import and every text query depends on an external endpoint, spend is per-token and unbounded, and text leaves your network. **Self-hosted module** (`text2vec-transformers` against your own inference container, pointed at by `TRANSFORMERS_INFERENCE_API`): data stays inside, cost becomes capacity you already pay for, and you keep the query-corpus guarantee — but you now operate a GPU or CPU inference service, size it, and keep it available or writes fail. **Self-provided vectors** (`Configure.Vectorizer.none()`): total control over model, version, batching and caching, and Weaviate stops being an inference client — but nothing enforces that queries use the same model as the corpus, and every model change is a full re-embed you must plan. I would ship a prototype on a hosted module, move to self-provided once relevance is being tuned or an embedding pipeline already exists elsewhere, and use named vectors to run both during the transition.

go deeper

for a junior

Know the three options exist — a hosted embedding module, a self-hosted inference container, or supplying vectors yourself — and that the collection's choice is set when it is created.

for a middle

Compare them on concrete axes: infrastructure to run, where the text goes, how cost scales, and what the query path depends on. Be able to name a workload that clearly suits each.

for a senior

Reason from the workload. Traffic shape, corpus size, residency rules and whether an embedding pipeline already exists should drive the answer, and you should describe how you would migrate without downtime.

for a principal

Frame it as ownership of the model and set the trigger conditions for changing position. Name the guarantee a hosted module gives you for free, the invariant you must enforce yourself once you give it up, and how named vectors keep the door two-way.

## Frame the decision correctly The question is not "which module is best". It is: **who owns the embedding model, and where does inference run?** Everything else — cost, latency, data residency, migration pain — falls out of that answer. Three positions are available, and mature systems often occupy more than one at a time. ## Option 1: hosted module Weaviate calls a third-party embedding API for you. Configured in one line, credentialed by a request header, working immediately. *Strengths.* Zero infrastructure. Access to strong general-purpose models without an ML team. And a guarantee that is easy to undervalue: because the same module embeds both the corpus and the query, the two can never drift apart. That entire class of silent relevance bug does not exist here. *Costs.* Every import and every text query is an external call, so the provider's availability is your availability on both paths, and their p99 is inside yours. Spend is per-token and scales with traffic rather than with data size, which makes it hard to cap. Your text goes to a third party, which may simply be disqualifying. And model versions move under you unless the module lets you pin one — a relevance evaluation from six months ago may no longer describe the system you are running. ## Option 2: self-hosted inference container Weaviate calls a model server you run — `text2vec-transformers` pointed at your own container via `TRANSFORMERS_INFERENCE_API`. *Strengths.* Text never leaves your network. Cost becomes hardware you already budget rather than a per-token meter. You keep the query-corpus matching guarantee. The model version is pinned by the image you deploy, so it changes only when you deploy. *Costs.* You are operating an inference service: sizing it for import bursts (which are far spikier than query load), keeping it available (its downtime is failed writes and failed text queries), monitoring it, and paying for accelerators that sit idle between import runs. On CPU, throughput may be low enough that bulk import becomes a scheduling problem. This option trades a vendor bill for engineering time, and only pays when you have the team to spend it. ## Option 3: self-provided vectors Weaviate stores and searches; your pipeline embeds. `Configure.Vectorizer.none()`. *Strengths.* One embedding pipeline for every consumer — the vector store, the reranker, the classifier, the cache — so identical text is embedded once. Batching, deduplication by content hash, retry policy and GPU scheduling are yours to optimise, and they are where the real cost savings live at scale. Any model works, including fine-tuned and multi-vector ones no module wraps. No third party on the read path. The model version is a dependency in your own build. *Costs.* The matching guarantee disappears: Weaviate validates dimension, not provenance, so a query embedded with a different model of the same width returns confident nonsense with no error. You must enforce that invariant yourself with pinned versions, recorded metadata and relevance regression tests. And you own the migration — changing models means re-embedding the whole corpus, evaluating, and cutting over. ## The dimensions that actually decide it - **Data sensitivity.** If text cannot leave the network, option 1 is out before any other argument is heard. - **Existing pipeline.** If something already embeds this text, embedding it again in the database is duplicated spend and a second model version to keep in sync. Choose option 3. - **Model specificity.** A fine-tuned domain encoder forces option 3. - **Team.** No ML operations capacity means option 1 or 3-with-a-managed-inference-provider, not option 2. - **Traffic shape.** High query volume against a small corpus makes per-query external calls expensive and puts weight on options 2 and 3. Bulk ingest of a large corpus with light querying inverts that. - **Relevance maturity.** The moment you are running relevance evaluations and comparing models, you need the model pinned and swappable on your own terms — option 3. ## The migration path, and why it is cheap These are not one-way doors. A collection's vectorizer is fixed at creation, but named vectors let one collection carry a module-generated vector and a self-provided vector simultaneously. So the transition is: add a self-provided named vector, backfill it from your new pipeline, evaluate both spaces on the same corpus, shadow-read, switch `target_vector`, then drop the old name. No dual-write application logic, no separate collection, no downtime. ## What I would actually do Start on a hosted module — the fastest path to knowing whether semantic search solves the problem at all, and the cheapest thing to throw away. Set a trigger for the move: the first of (a) an embedding pipeline appears elsewhere in the architecture, (b) relevance tuning starts, (c) embedding spend becomes a line item someone asks about, or (d) a data-residency requirement lands. Then migrate to self-provided vectors through a named vector, and treat the model identifier as a versioned contract from that day on. Self-hosted inference is the right middle only when residency rules them out of hosted APIs and the team can genuinely run the service.

  • What single fact most often settles this decision before any tradeoff analysis?
    Whether an embedding pipeline already exists in the architecture. If the same text is being embedded for a reranker, a classifier or another index, letting Weaviate embed it again doubles inference spend and creates a second model version to keep in sync. Self-provided vectors are then the obvious answer regardless of the other arguments.
  • Why is self-hosted inference the least commonly correct answer of the three?
    Because it takes on the operational burden without the main benefit of full ownership. You run and size a GPU service, absorb its availability as your write availability, and still have Weaviate calling it per object — yet you gain none of the batching, caching or cross-system reuse that a real embedding pipeline provides. It fits mainly when residency rules out hosted APIs and the team can operate it.
  • How do you keep a self-provided-vector deployment from silently drifting into a model mismatch?
    Treat the model identifier and version as part of the collection's contract: record it in metadata or deployment config, assert it in the service that issues queries so a mismatch fails loudly, and run a fixed relevance evaluation set as a regression test in CI. Weaviate checks dimension only, so no database-level error will ever catch this for you.

saying these in an interview costs you the question

  • Treating it as a pure cost question rather than an ownership question
  • Ignoring that a hosted module puts a third party on the read path
  • Assuming self-hosted inference is free because the hardware exists
  • Believing the decision is irreversible once the collection is created
  • Overlooking that a hosted module guarantees query and corpus models match

context