In LanceDB's create_index, what do num_partitions and num_sub_vectors control?
answer
- Two knobs, two different tradeoffs
- One is clustering, one is compression
- Fixed at build time, rebuild to change
- Roughly square root of row count
- Must divide the embedding dimension evenly
basics
~20 snum_partitions sets how many clusters the vectors are split into, so it governs how much data one query touches per probe. num_sub_vectors sets how many pieces each vector is chopped into for compression, so it governs index size and how lossy stored distances are.
solid answer
~50 sBoth are build-time knobs on `create_index(index_type="IVF_PQ", ...)` and neither can be changed without rebuilding. `num_partitions` is the number of k-means clusters: more partitions mean fewer vectors scanned per probe and faster queries, but each partition holds less, so a fixed probe count sees a smaller slice of the data and recall drops unless you probe more. A common starting point is roughly the square root of the row count, aiming for a few thousand vectors per partition. `num_sub_vectors` splits each embedding into that many chunks, each compressed to a small code — so it sets bytes per row in the index and how much distance information is lost. More sub-vectors means more fidelity and a bigger index; fewer means aggressive compression and worse ordering. It must divide the embedding dimension evenly, and keeping dimension divided by num_sub_vectors a multiple of 8 keeps the distance kernels efficient.
code
python · 8 lines# 768-dim embeddings, ~65k rows
tbl.create_index(
index_type="IVF_PQ",
metric="cosine",
num_partitions=256, # ~ sqrt(65_000), ~250 vectors per partition
num_sub_vectors=96, # 768 / 96 = 8 dims per chunk, SIMD friendly
replace=True,
)go deeper
Recall which parameter is which: num_partitions is how many clusters the data is split into, num_sub_vectors is how finely each vector is compressed. Both are set when the index is built.
Explain each tradeoff in a sentence and state the divisibility constraint on num_sub_vectors. Know the square-root-of-rows starting point and why partitions that are too small train badly.
Show a tuning method: baseline recall, change one build parameter at a time, and record index size and build time alongside latency. Know that PQ loss is fixed by fidelity or refinement, not by probing harder.
Own the economics. Index footprint decides whether it stays resident in cache, which on object-storage deployments dominates both latency and request cost, and rebuild time sets how often the shape can be revisited at all.
## What create_index actually builds `create_index` on a LanceDB table with `index_type="IVF_PQ"` builds a two-stage structure. The IVF stage partitions the vector space by running k-means over a sample of your vectors and assigning every row to its nearest centroid. The PQ stage replaces each stored vector with a compact code instead of the full float array. `num_partitions` configures the first stage, `num_sub_vectors` the second. `metric` — `l2`, `cosine` or `dot` — is fixed at build time too, and must match what you later pass to `.metric()` at query time. Because both parameters are baked into the built artifact, getting them wrong means a rebuild, which on a large table is a real batch job. That is why interviewers ask: the cost of the mistake is not a config reload. ## num_partitions Think of it as the coarseness of the first filter. With N rows and P partitions, an average partition holds N/P vectors, and a query that probes p partitions scores roughly p × N/P candidates. - **Raise P**: each partition is smaller, so a probe is cheaper and queries are faster. But the query's slice of the corpus shrinks with it, so with the probe count held constant recall falls — the true neighbour is more likely to sit just over a boundary you did not visit. - **Lower P**: each probe scans more vectors, giving better recall per probe but slower queries, and at the limit you are back to something close to a scan. The usual heuristic is P ≈ √N, tuned so partitions land in the low thousands of vectors. Two failure modes bracket it. Set P far too high on a small table and k-means has too few points per cluster to learn anything useful — training is poor and the partitioning is close to arbitrary. Set P too low on a huge table and the index barely narrows the search. Build cost also scales with P: more centroids means a longer training pass. ## num_sub_vectors Product quantization splits each embedding into equal chunks and encodes each chunk as the id of the nearest entry in a small learned codebook — one byte per chunk in the usual 8-bit configuration. So `num_sub_vectors` is, directly, the compressed size of a row in the index. A 768-dimension float32 embedding is 3072 bytes raw. With `num_sub_vectors=96`, the code is about 96 bytes — roughly a 32× reduction. That compression is the whole point on LanceDB, where the index is expected to stay cached while the full vectors live on disk or in object storage. The cost is fidelity. Distances computed from codes are approximations, and the fewer sub-vectors you use, the coarser each chunk's codebook is relative to the data it must represent, so the ordering of near-tied candidates degrades. Raising `num_sub_vectors` improves that ordering and enlarges the index proportionally. **Hard constraint**: the embedding dimension must be divisible by `num_sub_vectors` — 768 works with 96, 64 or 48; it does not work with 100. Beyond divisibility, keeping dimension ÷ num_sub_vectors a multiple of 8 lets the distance computation use wide SIMD operations, which is a measurable throughput difference rather than a stylistic preference. ## How the two interact They trade off against different things and should be tuned separately. `num_partitions` trades query latency against how much of the corpus a query sees; `num_sub_vectors` trades index footprint against how accurate stored distances are. A common mistake is compensating for aggressive PQ compression by probing more partitions — that costs latency without fixing the real problem, which is that the compressed distances are too lossy to order candidates correctly. The right remedy for PQ loss is either more sub-vectors or a refinement step that rescores candidates against the full-precision vectors. ## Practical procedure Start from √N partitions and a sub-vector count that divides your dimension cleanly with an 8-aligned chunk size. Build. Measure recall@k against a flat-scan baseline and record build time and index size. If recall is low but latency is fine, probe more or add refinement before touching the build parameters. If latency is the problem, raise partitions and re-measure recall. Change one variable per build — with a rebuild costing minutes to hours, sweeping blindly is expensive.
- What goes wrong if num_sub_vectors does not divide the embedding dimension?Product quantization needs equal-width chunks, so the build rejects a sub-vector count that leaves a remainder — 768 with 100 sub-vectors has no valid split. Pick a divisor instead, and prefer one where dimension ÷ num_sub_vectors is a multiple of 8 so the distance kernels stay SIMD-friendly. If your model's dimension is awkward, reducing dimensionality at the embedding stage is cleaner than fighting the index.
- Why does raising num_partitions sometimes make recall worse rather than better?More partitions means each one holds fewer vectors, so a query probing a fixed number of partitions sees a smaller slice of the corpus and is likelier to miss a true neighbour sitting just across a cluster boundary. Partition count and probe count have to move together: if you double partitions for latency, re-measure recall and raise the probe count to compensate.
saying these in an interview costs you the question
- Thinking these can be tuned at query time
- Copying num_partitions from a tutorial regardless of row count
- Setting num_sub_vectors without checking it divides the dimension
- Believing more partitions always improves recall
- Compensating for lossy compression by probing more partitions