skip to content

A FAISS GPU index OOMs though the vectors fit in device RAM. What is consuming memory?

level: seniorimportance: should knowfreq 42%

answer

  1. vectors are not the whole budget
  2. something is reserved before you add data
  3. ids cost bytes per vector too
  4. query batch size allocates workspace
  5. load-triggered OOM means working set

basics

~20 s

Vectors are only part of the budget. StandardGpuResources pre-reserves a scratch pool, ids cost 8 bytes each by default, IVFPQ precomputed tables scale with nlist, and each search allocates workspace proportional to the query batch. Shrink the pool, narrow ids, use half precision, or batch smaller.

solid answer

~50 s

Count everything that lives on the device, not just the vectors. `StandardGpuResources` reserves a scratch pool up front — often around a gigabyte — before any data is added; `res.setTempMemory(n_bytes)` shrinks it and `res.noTempMemory()` removes it at the cost of slow per-search allocations. Ids default to 64 bits per vector, so `indicesOptions = faiss.INDICES_32_BIT` or `faiss.INDICES_CPU` reclaims real space on a large index. For IVFPQ, `usePrecomputedTables` builds a table that scales with `nlist` times the code layout and can dwarf the codes themselves on a large-nlist index. Finally, search allocates transient workspace proportional to the query batch size, the number of probed lists and k — so an OOM that appears only under load is usually a batch that grew, not an index that grew. Fix in that order: measure, shrink the pool, narrow ids, enable half precision, then reduce the batch or shard across GPUs.

code

python · 13 lines
python
import faiss

res = faiss.StandardGpuResources()
res.setTempMemory(256 * 1024 * 1024)   # shrink the pre-reserved pool

d, nlist, m = 768, 65536, 64
config = faiss.GpuIndexIVFPQConfig()
config.device = 0
config.indicesOptions = faiss.INDICES_32_BIT   # halve id storage
config.useFloat16LookupTables = True           # half-precision tables
config.usePrecomputedTables = False            # large at high nlist

index = faiss.GpuIndexIVFPQ(res, d, nlist, m, 8, faiss.METRIC_L2, config)

go deeper

for a junior

Know that a GPU index has to hold everything in device memory at once, and that vectors are not the only thing stored there — ids and scratch space count too.

for a middle

Be able to itemise the budget: scratch pool, vectors or codes, ids, coarse centroids and precomputed tables, plus per-query workspace, and name the settings that shrink each one.

for a senior

Demonstrate the diagnosis: measure device memory at rest and under representative batch and concurrency, separate steady-state from working-set failures, then order remedies by their cost to accuracy and latency.

for a principal

Own the capacity model — memory per vector at a chosen index type, headroom for query working set and growth, and the threshold at which sharding across devices beats compressing further.

## The device memory budget has five line items When a FAISS GPU index runs out of memory, the instinct is to divide vector count by device capacity and conclude it should have fit. That calculation is wrong because it counts one of five consumers. **1. The scratch pool.** Creating `StandardGpuResources()` immediately reserves a temporary memory pool sized from the device's total memory — on modern cards this is on the order of one to two gigabytes. Its purpose is to avoid `cudaMalloc` on the query path, which is slow and synchronising. Its side effect is that `nvidia-smi` shows a large allocation before you have added a single vector. Control it with `res.setTempMemory(256 * 1024 * 1024)` before cloning, or eliminate it with `res.noTempMemory()` — but understand the trade: with no pool, every search allocates and frees, and throughput can fall sharply. **2. The vectors or codes.** An IVFFlat index stores `n * d * 4` bytes of float32 vectors. An IVFPQ index stores `n * m` bytes of codes instead, which is why it exists — a hundred million 768-dimensional float vectors are roughly 300 GB uncompressed and a few gigabytes as PQ codes. **3. The ids.** Every vector carries an id, 8 bytes by default. At a hundred million vectors that is 800 MB of pure bookkeeping — frequently larger than the PQ codes it accompanies. `indicesOptions = faiss.INDICES_32_BIT` halves it when your ids fit in 32 bits; `faiss.INDICES_CPU` moves the table to host memory entirely, so the device stores codes only and result ids are resolved on the host. **4. Auxiliary tables.** The coarse quantizer's centroids are `nlist * d * 4` bytes — small at nlist=1024, not small at nlist=262144. IVFPQ's `usePrecomputedTables` option builds a table whose size scales with `nlist` and the code layout; it speeds up queries and can quietly become the single largest allocation on a large-nlist index. Turning it off is often the cheapest way to make an index fit. **5. Per-query workspace.** Search allocates transient buffers proportional to the number of queries in the batch, the number of lists probed and k: candidate distances, per-list scan buffers, and the output distance and id arrays. This is the item that explains the most confusing failure mode — an index that has been stable for weeks OOMs the day someone raises the batch size or the probe count. The index did not grow; the working set did. ## Diagnosing it Measure device memory at three points: after creating the resources object (that is the pool), after `add` completes (pool plus index), and during a representative search at production batch size (plus workspace). The gap between the second and third readings is your query working set, and it is the one nobody budgets for. If the failure only appears under concurrency, remember that concurrent searches each need their own workspace out of the same pool. ## The order to fix things 1. **Shrink the scratch pool** if it is oversized for the workload — the cheapest win, though watch latency afterwards. 2. **Narrow the ids** with 32-bit or host-side storage; on a large index this is often hundreds of megabytes for no accuracy cost. 3. **Enable half precision** — `useFloat16LookupTables` on IVFPQ, and the half-precision flat storage option — trading a small amount of accuracy for a large amount of space. 4. **Drop precomputed tables** if `nlist` is large; you lose some query speed and regain a lot of memory. 5. **Reduce the query batch** or cap concurrency, which fixes load-triggered OOMs specifically. 6. **Change index type** — IVFFlat to IVFPQ is the step change, since it compresses the dominant line item. 7. **Shard across GPUs** when the index genuinely exceeds one device, which is a capacity decision rather than a tuning one. ## What separates a senior answer A weaker candidate proposes "use a bigger GPU" or "compress the vectors" and stops. A senior answer enumerates the consumers, distinguishes the *steady-state* footprint from the *per-query* footprint — because they fail differently, one at startup and one under load — and orders the remedies by cost to accuracy and latency rather than reaching for the most invasive one first.

  • Why does the OOM sometimes only appear under production load and not during your capacity test?
    Steady-state footprint and query working set are separate budgets. Search allocates buffers proportional to batch size, probed lists and k, and concurrent searches each need their own. A single-query test measures only the steady state, so an index that fits at rest can fail the moment real batch sizes and concurrency arrive. Test at production batch size and concurrency, not with one query.
  • What is the downside of calling res.noTempMemory() to reclaim space?
    The scratch pool exists so searches never call into the device allocator on the hot path. Without it every search allocates and frees device memory, which is slow and synchronising, so throughput can drop sharply and latency becomes far less predictable. Prefer shrinking the pool with setTempMemory to a size the workload actually needs rather than removing it outright.
  • On a large-nlist IVFPQ index, why might turning off precomputed tables be the cheapest fix?
    The precomputed table's size scales with the number of inverted lists and the code layout, so at a large nlist it can exceed the PQ codes it is meant to accelerate. Disabling it costs some query speed but returns memory proportional to the largest single auxiliary allocation — usually a better trade than compressing further or buying a bigger device.

saying these in an interview costs you the question

  • Sizing the GPU from vector bytes alone
  • Not knowing the resources object reserves memory before any data is added
  • Forgetting that ids cost 8 bytes per vector by default
  • Assuming query batch size has no memory cost
  • Reaching for a bigger GPU before measuring where memory went

context