skip to content

Qdrant offers scalar, product, and binary quantization — what does each trade?

level: seniorimportance: must knowfreq 66%

answer

  1. three modes, three compression ratios
  2. int8 is the safe default
  3. one bit per dimension needs high dimensions
  4. codebooks cost build time
  5. the compressed copy is the one that should stay resident

basics

~20 s

Scalar quantization stores int8 components for about 4x compression with small accuracy loss and is the safe default. Product quantization compresses far harder (up to 64x) at real accuracy cost and slower builds. Binary quantization keeps one bit per dimension — 32x, fastest, but only viable for high-dimensional embeddings and only with rescoring.

solid answer

~50 s

All three are set through `quantization_config` on `create_collection` or `update_collection`, and all three keep the original vectors so results can be rescored. **Scalar** — `models.ScalarQuantization(scalar=models.ScalarQuantizationConfig(type=models.ScalarType.INT8, quantile=0.99, always_ram=True))`. Each float32 component becomes an int8, roughly 4x smaller, with accuracy loss usually under a percent. `quantile` clips outliers so the int8 range is not wasted on a few extreme components. This is the default choice for most workloads. **Product** — `models.ProductQuantization(product=models.ProductQuantizationConfig(compression=models.CompressionRatio.X16, always_ram=True))`, with ratios from X4 to X64. It compresses much harder but distances become noticeably approximate and indexing is significantly slower. **Binary** — `models.BinaryQuantization(binary=models.BinaryQuantizationConfig(always_ram=True))`. One bit per dimension, about 32x smaller and by far the fastest distance computation, but it only holds up on high-dimensional embeddings and needs oversampling plus rescoring. `always_ram=True` in each keeps the compressed copy in memory even when the originals live on disk — which is the whole point of quantizing.

code

python · 13 lines
python
from qdrant_client import QdrantClient, models

client = QdrantClient(url="http://localhost:6333")
client.update_collection(
    collection_name="docs",
    quantization_config=models.ScalarQuantization(
        scalar=models.ScalarQuantizationConfig(
            type=models.ScalarType.INT8,
            quantile=0.99,
            always_ram=True,
        )
    ),
)

go deeper

for a junior

Know that Qdrant can store a compressed copy of vectors, that the three modes are scalar, product and binary, and that the originals are kept so results can be re-checked.

for a middle

Explain the compression ratios and their mechanisms: int8 per component, codebook indices per subvector, one bit per dimension — and that always_ram controls whether the compressed copy stays in memory.

for a senior

Show that you would pick from a measured memory budget and validate recall against exact search before and after, and that the useful pairing is originals on disk with the quantized copy pinned in RAM.

for a principal

Own the cost model: quantization is how vector RAM spend scales sub-linearly with corpus growth, so set the policy for which collections quantize, at what recall floor, and who signs off when accuracy is traded for machine cost.

## The shape of the feature Quantization in Qdrant is a *storage and distance-computation* optimization layered onto an existing collection. The original float32 vectors are still kept; the quantized copy is an additional, much smaller representation used to traverse the HNSW graph quickly. Because the originals survive, Qdrant can re-rank the approximate candidates using exact distances — the mechanism that makes aggressive compression usable at all. You attach it at creation: `client.create_collection(collection_name="docs", vectors_config=models.VectorParams(size=1536, distance=models.Distance.COSINE), quantization_config=models.BinaryQuantization(binary=models.BinaryQuantizationConfig(always_ram=True)))` or later with `client.update_collection(collection_name="docs", quantization_config=...)`, which makes the optimizer rebuild segments in the background. You never call a separate training step; codebooks and ranges are computed by the optimizer as segments are built. ## Scalar quantization `ScalarQuantizationConfig` currently supports `type=models.ScalarType.INT8`. Every float32 component is mapped linearly onto one signed byte, so the vector shrinks about 4x. Two knobs matter. `quantile` (for example 0.99) decides the value range that maps onto the int8 scale. Embedding components are usually near-Gaussian with a few extreme outliers; if the range is set by those outliers, the entire useful middle is squeezed into a handful of levels. Clipping at the 99th percentile sacrifices fidelity for rare extremes and buys resolution where the mass of the distribution actually lives. `always_ram=True` pins the quantized vectors in memory. This is what turns quantization into a RAM strategy: the 4x-smaller copy stays resident and serves traversal, while the originals can be pushed to disk. Accuracy loss is typically small — often under one percent recall for common text embeddings — which is why scalar is the recommended starting point. Speed also improves, because int8 distance computation is cheaper and far more cache-friendly than float32. ## Product quantization `ProductQuantizationConfig` takes a `compression` value from `models.CompressionRatio`: X4, X8, X16, X32, X64. Vectors are split into subvectors, each replaced by an index into a learned codebook, so the stored form is a handful of bytes regardless of the original dimensionality. The trade is sharper than with scalar. Distances become genuinely approximate — neighbourhood structure survives, exact ordering does not — and index building is markedly slower because codebooks must be computed. In exchange, memory drops by an order of magnitude or more, which can be the difference between a collection fitting in RAM and not. Use it when memory is the binding constraint, the collection is large, and you have measured that rescoring recovers enough recall. Do not reach for it as a default: on mid-sized collections, scalar quantization plus more RAM is usually the better engineering answer. ## Binary quantization `BinaryQuantizationConfig` reduces each component to a single bit, roughly 32x compression. Distance computation becomes bitwise work, which is dramatically faster than float arithmetic — this is the mode people choose for latency, not only for memory. The catch is that it only works when the embedding space tolerates it. High-dimensional embeddings (roughly 1024 dimensions and up) from models whose components are well-centred retain enough signal in their sign pattern; low-dimensional embeddings do not, and recall collapses. Binary quantization is also effectively unusable without oversampling and rescoring: you retrieve a wide candidate set with cheap bit distances, then re-rank with the original vectors. Validate it, always, on your own embeddings before committing. "Works great for 1536-dimensional embeddings" is a statement about particular models, not a law. ## always_ram and the storage picture Each config carries `always_ram`. Left false, the quantized vectors may be memory-mapped like everything else; set true, they are kept resident. The canonical high-value configuration is `VectorParams(..., on_disk=True)` for the originals plus quantization with `always_ram=True`: traversal touches only the small in-memory copy, and disk is read only during rescoring of the final candidates. Inverting that — originals in RAM, quantized on disk — gets you the accuracy loss without the memory saving. ## Choosing Start from the memory budget. Compute raw size as points x dimensions x 4 bytes and compare it to the RAM you are willing to buy. If raw fits with headroom, quantization is optional and scalar is a cheap latency win. If you are 3-5x over, scalar with originals on disk usually closes the gap. If you are an order of magnitude over, choose between product and binary by dimensionality and latency needs — and in every case measure recall against exact ground truth before and after, because the whole point is to know exactly how much accuracy you traded away.

  • What does the quantile setting in scalar quantization actually buy you?
    Resolution where the data is. Embedding components are roughly bell-shaped with rare extremes; if the int8 range is stretched to cover those extremes, most components collapse into a few levels. Setting `quantile=0.99` clips the outer one percent so the byte range covers the bulk of the distribution, trading fidelity on rare extreme components for much better precision on typical ones. It usually improves recall rather than hurting it.
  • Why does binary quantization work for 1536-dimensional embeddings but fail on 384-dimensional ones?
    Because a sign bit per dimension retains a usable fraction of the information only when there are many dimensions to aggregate over. At 1536 dimensions the bit pattern still separates neighbours from non-neighbours well enough for a candidate set; at 384 there is far too little signal left and the candidate set no longer contains the true neighbours, so rescoring has nothing good to re-rank. Always validate on your own model rather than trusting a dimension threshold.
  • If you set always_ram=False on the quantized vectors while the originals stay in RAM, what have you achieved?
    Essentially the worst of both. You pay quantization's accuracy loss and its build cost, but the memory saving never materialises because the full-precision vectors are still resident, and traversal may now fault to disk for the compressed copy. The intended pairing is the reverse: originals on disk via `on_disk=True`, quantized copy pinned with `always_ram=True`, so hot traversal is in memory and disk is touched only when rescoring.

saying these in an interview costs you the question

  • Says quantization discards the original vectors
  • Treats product quantization as the default choice
  • Expects binary quantization to work on low-dimensional embeddings
  • Enables quantization without measuring recall against exact search
  • Leaves always_ram false and expects a memory saving

context