skip to content

When would you configure a Weaviate collection with a flat vector index instead of HNSW?

level: middleimportance: should knowfreq 48%

answer

  1. Set per collection, not per query
  2. Brute force versus navigable graph
  3. Exact recall, linear query cost
  4. Think vectors per shard, not per corpus
  5. One option switches over a threshold

basics

~20 s

Choose flat for small collections, typically a few thousand objects, and for per-tenant shards. It scans every vector, so recall is exact and there is no graph to build or hold in memory, but latency grows linearly as the collection grows.

solid answer

~50 s

Weaviate sets the vector index per collection with `Configure.VectorIndex.hnsw()`, `.flat()` or `.dynamic()`. Flat is brute force: it compares the query against every vector in the shard, so results are exact, imports have no graph-building cost, and memory is just the vectors — optionally shrunk further with binary quantization, the quantizer flat supports. The price is that query time scales linearly with object count. HNSW is the opposite trade: a graph that gives sub-linear search on millions of vectors, at the cost of build time, extra memory for the graph and approximate recall you tune. The decision is about size per shard, not size overall. In a multi-tenant collection each tenant is its own shard, so thousands of small tenants are usually better served by flat than by thousands of half-empty graphs. `dynamic` splits the difference: it starts flat and converts to HNSW once the shard crosses a threshold.

code

python · 11 lines
python
import weaviate
import weaviate.classes.config as wvcc

client = weaviate.connect_to_local()
client.collections.create(
    name="TenantDoc",
    vector_index_config=wvcc.Configure.VectorIndex.flat(
        quantizer=wvcc.Configure.VectorIndex.Quantizer.bq(),
    ),
)
client.close()

go deeper

for a junior

Know that a Weaviate collection picks its vector index type at creation, and that flat means an exact scan of every vector while HNSW is an approximate graph search.

for a middle

Explain the trade in both directions: flat gives exact recall, zero build cost and small memory but linear query time; HNSW gives sub-linear search at the cost of build time, memory and tuned approximate recall.

for a senior

Reason from vectors per shard rather than per corpus, tie the choice to the multi-tenant shard model and the node's memory budget, and note that the type is fixed at creation so a wrong call means a re-import.

for a principal

Own the fleet-level call: model the tenant size distribution, decide where dynamic's threshold sits, and set the policy for when a growing collection is migrated rather than left to degrade.

## The setting A Weaviate collection carries one vector index configuration, chosen at creation through `weaviate.classes.config.Configure.VectorIndex`. The three options are `hnsw()`, `flat()` and `dynamic()`. Each takes its own parameters — the distance metric, cache sizes, and an optional quantizer — and the choice of *type* is effectively permanent for the collection, so it is a design decision rather than a knob. ## What flat does A flat index stores the vectors and nothing else. A query compares the query vector against every vector in the shard and returns the true nearest neighbours. Three consequences follow. **Recall is exact.** There is no approximation, so no recall tuning, no surprising misses, and no divergence between what you tested and what production returns. **There is nothing to build.** Imports write vectors and stop. No graph construction means fast ingest, no background index building, and no build parameters to get wrong. **Memory is just the vectors.** Flat supports binary quantization, which compresses each vector aggressively; because the index also rescores candidates against fuller information, a compressed flat index can stay small and still return sensible results. The cost is the obvious one: work per query is proportional to the number of vectors in the shard. At a few thousand objects that is microseconds-to-milliseconds and invisible. At a few million it is unacceptable. ## What HNSW buys and charges HNSW builds a navigable graph so that a query touches a small, roughly logarithmic fraction of the data. That is what makes million-scale collections searchable in milliseconds. You pay in three currencies: build time during import, memory for the graph structure on top of the vectors, and approximation — recall is now a tuned quantity, controlled by search-time effort, rather than a guarantee. Weaviate's HNSW configuration exposes `max_connections` and `ef_construction` at build time and `ef` (or the dynamic-ef range) at search time. ## Size per shard is the real variable The mistake candidates make is reasoning about total corpus size. Weaviate indexes per shard, and in a multi-tenant collection each tenant gets its own shard with its own vector index. A SaaS product with 5,000 customers averaging 2,000 documents each has ten million vectors in total but only two thousand per index. Building five thousand HNSW graphs for that shape wastes memory on graph overhead and gives each tenant no meaningful speedup, because two thousand brute-force comparisons are already fast. Flat is the better fit, and it keeps the per-tenant memory footprint small enough that many tenants can stay resident. The same logic applies to small reference collections — a product taxonomy, a set of canned responses, a few thousand FAQ entries — where exactness matters more than throughput. ## The dynamic option `Configure.VectorIndex.dynamic()` starts a shard on a flat index and converts it to HNSW once it crosses a configured object threshold. It is aimed squarely at the multi-tenant case where most tenants stay small but a few grow large: small tenants keep exact search and low memory, big tenants get the graph automatically. The conversion is a background operation, so the deployment needs asynchronous indexing enabled for it to work. ## How to decide Ask how many vectors sit in one shard at steady state and at the 95th percentile of your tenants. If that number is in the low thousands and stays there, flat is simpler, exact and cheaper. If it is in the hundreds of thousands or millions, HNSW. If the distribution is wildly skewed — the usual SaaS shape — dynamic. Then measure: brute force at your dimensionality and shard size either meets the latency budget or it does not, and that measurement beats every rule of thumb. ## Failure modes Picking flat for a collection that grows into millions produces latency that degrades smoothly and invisibly until it does not; since the index type is not something you flip in place, the fix is a new collection and a re-import. Picking HNSW for thousands of tiny tenants produces a memory bill dominated by graph overhead and a heap that grows with tenant count rather than with data.

  • What does the dynamic index type need from the deployment to work?
    Asynchronous indexing must be enabled on the server, because the flat-to-HNSW conversion happens in the background once a shard crosses the configured threshold. Without it the shard cannot build the graph out of band. It is worth confirming before you rely on dynamic in a multi-tenant deployment, otherwise you have effectively chosen flat with an unpleasant surprise waiting at scale.
  • Can you switch a collection from flat to HNSW after it is created?
    Not as an in-place edit of an existing flat or HNSW collection — the index type is fixed at creation, and changing it means creating a new collection with the desired configuration and re-importing. The `dynamic` type exists precisely because that migration is painful: it makes the transition automatic per shard rather than a manual re-import.
  • How does the choice interact with how much memory a Weaviate node needs?
    HNSW holds the graph in addition to the vectors, so per-shard overhead is real and multiplies by shard count. Flat holds only vectors, and binary quantization shrinks those substantially. In a multi-tenant deployment with thousands of small active tenants, that difference decides how many tenants a node can keep resident before you have to start deactivating them.

saying these in an interview costs you the question

  • Says flat is always slower, so never use it
  • Reasons from total corpus size instead of per-shard size
  • Thinks the index type can be toggled on a live collection
  • Assumes HNSW recall is exact like brute force
  • Believes the index is chosen per query rather than per collection

context