What does the FAISS index_factory string "IVF4096,PQ64" build?
answer
- Comma-separated pipeline, read left to right
- Transform, then coarse quantizer, then encoding
- First number is nlist
- Second number is bytes per vector
- Arrives untrained on purpose
basics
~20 sIt builds an IVFPQ index: a coarse quantizer that splits the space into 4096 inverted lists, with each vector stored as a product-quantized code of 64 one-byte values instead of the raw floats. Both stages must be trained before adding data.
solid answer
~40 s`faiss.index_factory(d, "IVF4096,PQ64")` composes an index from comma-separated components read left to right: optional preprocessing (`PCA80`, `OPQ64_256`), then the coarse quantizer (`IVF4096`), then the per-vector encoding (`PQ64`). Here 4096 is `nlist`, the number of Voronoi cells the vectors are partitioned into, and 64 is the number of sub-quantizers, so each vector becomes 64 bytes at the default 8 bits per code. The factory is equivalent to constructing `IndexIVFPQ(quantizer, d, 4096, 64, 8)` with a flat coarse quantizer, but far less error-prone. Two constraints bite: `d` must be divisible by 64, and the index arrives with `is_trained` False — you must call `train()` on a representative sample before `add()`. `nlist` is normally sized around 4x to 16x the square root of the collection size, with enough training vectors per centroid to make the clustering meaningful.
code
python · 13 linesimport numpy as np
import faiss
d = 128
index = faiss.index_factory(d, "IVF4096,PQ64", faiss.METRIC_L2)
print(index.is_trained) # False - both stages must be learned
xt = np.random.random((200000, d)).astype('float32')
index.train(xt) # learns 4096 centroids and the PQ codebooks
index.add(xt)
index.nprobe = 16
D, I = index.search(xt[:5], 10)go deeper
Know that index_factory takes a string describing the index and that some indexes must be trained before you add vectors to them.
Be able to decompose the string component by component, say that 4096 is nlist and 64 is bytes per vector, and state the divisibility rule and the training requirement.
Show you can size nlist from the collection size and the available training sample, compute the memory saving out loud, and name what the compression costs in accuracy.
Own the trade explicitly: argue when a corpus is small enough that compression buys nothing but a recall regression, and treat the factory string as a configuration artifact you version and re-evaluate as data grows.
## Why a factory string exists FAISS index classes compose. A realistic production index is a preprocessing transform wrapped around a partitioned index wrapped around an encoding, and building that by hand means instantiating a quantizer, passing it to an `IndexIVFPQ` constructor with four positional integers, and possibly wrapping the result again. `faiss.index_factory(d, description, metric)` takes a short string and does all of that, which is why nearly every FAISS example and benchmark is written in factory notation. It also makes an index configuration a piece of data — a string in a config file — rather than code. ## Reading the string left to right A factory string is a comma-separated pipeline: 1. **Optional vector transforms.** `PCA80` reduces the dimension to 80 before indexing. `OPQ64_256` applies a learned rotation that makes the data friendlier to a 64-sub-quantizer PQ, optionally reducing to 256 dimensions first. `L2norm` normalises. These stages are learned during `train()`. 2. **The coarse quantizer / partitioning stage.** `IVF4096` says: cluster the (possibly transformed) training vectors into 4096 centroids and route each stored vector to the nearest one. `IVF65536_HNSW32` says the same but uses an HNSW index rather than a flat scan to find the nearest centroid, which matters once `nlist` gets large. Omitting this stage entirely gives a non-partitioned, exhaustive index. 3. **The encoding.** `Flat` stores the raw float32 vectors. `PQ64` stores 64 product-quantized codes. `SQ8` stores one byte per dimension via scalar quantization. `PQ64x4fs` is the fast-scan SIMD variant with 4-bit codes. So `"Flat"` is a plain exact index, `"IVF4096,Flat"` partitions but does not compress, `"PQ64"` compresses but does not partition, and `"IVF4096,PQ64"` does both. ## What the two numbers mean `4096` is `nlist`: the number of inverted lists, or Voronoi cells. It controls how much of the dataset a query touches — a search visits `nprobe` of those 4096 lists rather than all of them, which is where the speedup comes from. Larger `nlist` means shorter lists and a faster scan per list, at the cost of a more expensive coarse assignment and a larger centroid table (`nlist * d * 4` bytes). The standard sizing heuristic is between 4 and 16 times the square root of the number of vectors: about 4k–16k lists for 1M vectors, about 65k for 16M. There is a second constraint on `nlist` that is easy to miss — clustering needs enough training points per centroid to be meaningful, on the order of tens of vectors per centroid, so a huge `nlist` on a small training sample produces poorly placed centroids and FAISS will warn about it. `64` in `PQ64` is the number of sub-vectors the dimension is split into. Each sub-vector is replaced by the id of its nearest entry in a learned 256-entry codebook, so at the default 8 bits that is exactly one byte per sub-vector: 64 bytes per stored vector regardless of the original dimension. The hard constraint is divisibility — `d % 64` must be 0, so a 768-dim embedding works with `PQ64` (12 dims per sub-vector) or `PQ96` (8 dims) but not with `PQ50`. ## Memory arithmetic This is the payoff and the thing worth being able to compute out loud. Raw float32 storage is `4 * d` bytes per vector — 3072 bytes at d=768. `PQ64` is 64 bytes of codes plus an 8-byte int64 id in the inverted list, plus small per-list overhead: roughly 72–80 bytes, a compression ratio near 40x. On top of that sits the centroid table, `nlist * d * 4` bytes, which for 4096 lists at d=768 is about 12.6 MB — negligible against the codes. ## Training is mandatory An index built from this string comes back with `is_trained` False, and calling `add()` on it raises. Both the IVF centroids and the PQ codebooks are learned from data, so you must call `train(xt)` on a sample that looks like the real distribution first. A sample drawn from one tenant, one language, or one time window produces centroids that fit that slice and cluster the rest badly — the resulting recall loss looks like a mysterious quality regression rather than a configuration error. ## What it costs you PQ is lossy. Distances are computed against the reconstructed codes, so the ranking is approximate even before the IVF partitioning skips lists. That is a deliberate trade: you accept a recall hit to fit 20M vectors in a couple of gigabytes instead of sixty. When the loss is unacceptable, the usual escape hatch is to re-rank the shortlist against exact vectors — the `RFlat` suffix and `IndexRefineFlat` do this — but that requires storing the full vectors somewhere and gives back the memory you just saved. Being explicit about that trade, rather than presenting factory strings as free wins, is what a strong answer sounds like.
- How would you pick nlist for a 10M-vector collection?Start from the heuristic of 4x to 16x the square root of the count — for 10M that is roughly 12k to 50k lists — then check you have enough training data, on the order of tens of vectors per centroid, or the clustering is meaningless. Larger nlist shortens each list but makes the coarse assignment costlier, so past about 65k lists people switch the quantizer to an HNSW variant such as IVF65536_HNSW32.
- Why can't you use PQ50 on 768-dimensional embeddings?Product quantization splits the vector into equal sub-vectors, so the number of sub-quantizers must divide the dimension. 768 is not divisible by 50, and construction fails. Valid choices at d=768 include PQ64, PQ96, PQ128 or PQ192. If the dimension is awkward, put a transform in front — PCA or OPQ — to reduce it to a friendly size first.
- What does the leading OPQ component in a string like "OPQ64_256,IVF4096,PQ64" contribute?It is a learned linear transform applied before quantization. It rotates and, here, reduces the data to 256 dimensions so that the variance is spread evenly across the 64 sub-vectors PQ will carve out, which makes each codebook fit its slice better. It costs extra training time and a matrix multiply per query, and buys back some of the accuracy PQ gives away.
- What is the difference between "IVF4096,Flat" and "IVF4096,PQ64"?Both partition the vectors into 4096 lists and scan only the lists nearest the query, so both are approximate in the same way. The difference is storage: IVF4096,Flat keeps the raw float32 vectors, so distances within a scanned list are exact, while IVF4096,PQ64 stores 64-byte codes and computes approximate distances. The first trades only latency, the second also trades memory for accuracy.
saying these in an interview costs you the question
- Reads 4096 as the number of bytes or the vector dimension
- Thinks a factory index can be used without calling train()
- Assumes any PQ size works regardless of the dimension
- Says the string is only a shorthand with no constraints
- Claims PQ compression is lossless