skip to content

How do you move a FAISS index onto a GPU, and what does StandardGpuResources do?

level: middleimportance: must knowfreq 62%

answer

  1. one handle object per device
  2. scratch pool reserved up front
  3. clone call takes device ordinal
  4. keep the Python reference alive
  5. serialise only from the CPU form

basics

~20 s

Create a faiss.StandardGpuResources() object, then call faiss.index_cpu_to_gpu(res, device, index). The resources object owns the GPU scratch memory and cuBLAS handles for that device, and you must keep a Python reference to it for as long as the index lives.

solid answer

~40 s

FAISS keeps GPU state in a `StandardGpuResources` object: it pre-reserves a scratch memory pool on one device and holds the stream and cuBLAS handles the kernels use. The usual flow is `res = faiss.StandardGpuResources()` then `gpu_index = faiss.index_cpu_to_gpu(res, 0, cpu_index)`, which copies the CPU index's structures (centroids, codes, vectors) into device memory. The classic Python bug is letting `res` go out of scope — the wrapper does not keep it alive for you, and the index then references freed resources and segfaults, so hold it in a module-level or object field. Everything after that is the normal FAISS API: `train`, `add`, `search`. To persist, convert back with `faiss.index_gpu_to_cpu(gpu_index)` first, because `write_index` only serialises CPU indexes.

code

python · 17 lines
python
import faiss
import numpy as np

d, nlist = 128, 1024
xb = np.random.random((100000, d)).astype('float32')

quantizer = faiss.IndexFlatL2(d)
cpu_index = faiss.IndexIVFFlat(quantizer, d, nlist)

res = faiss.StandardGpuResources()          # must outlive gpu_index
gpu_index = faiss.index_cpu_to_gpu(res, 0, cpu_index)
gpu_index.train(xb)
gpu_index.add(xb)

D, I = gpu_index.search(xb[:1024], 10)

faiss.write_index(faiss.index_gpu_to_cpu(gpu_index), "ivf.faiss")

go deeper

for a junior

Know that FAISS GPU support is opt-in code you write: you create a resources object and clone the index onto a device. Nothing happens automatically just because the machine has a GPU.

for a middle

Be ready to write the three-line clone yourself, explain that the resources object owns a pre-reserved scratch pool and CUDA handles, and name the round trip back to CPU that serialisation requires.

for a senior

Show you have debugged this in production: the dangling-resources segfault, the transient double memory cost during a clone, and building on GPU while shipping a CPU-serialised artifact so serving hosts stay flexible.

for a principal

Own the deployment shape — where indexes are built versus served, whether GPU is a build-time accelerator or a serving dependency, and what the fallback path is when GPU capacity is unavailable or costs spike.

## What the GPU side of FAISS actually is FAISS is a library, not a server, and its GPU support is a parallel set of index implementations that live in device (GPU) memory. You do not point a config file at a GPU; you either build a GPU index class directly or you clone a CPU index onto a device. Nothing about the search API changes — the same `train`, `add`, `search` calls apply — but the data now lives in GPU RAM and the distance computations run as CUDA kernels. ## StandardGpuResources `faiss.StandardGpuResources()` is the handle for one GPU's runtime state. It holds: - a **temporary scratch pool** pre-allocated on the device, reused by every search so kernels do not hit `cudaMalloc` on the hot path (device allocation is slow and synchronising); - the CUDA **stream** the index's kernels are enqueued on; - the **cuBLAS handle** used for the matrix multiplications behind distance computation. Because the pool is reserved up front, GPU memory usage jumps the moment you create the resources object, before you have added a single vector. You can shrink it with `res.setTempMemory(n_bytes)` or disable it with `res.noTempMemory()` (which trades memory for slower, allocation-heavy searches). One resources object per GPU is the normal pattern; sharing one across many indexes on the same device is fine and preferred, since they then share the same pool. ## Cloning an index onto a device ```python res = faiss.StandardGpuResources() gpu_index = faiss.index_cpu_to_gpu(res, 0, cpu_index) ``` The third argument is the CPU index; the second is the device ordinal (`faiss.get_num_gpus()` tells you how many are visible). The clone copies the index's contents into device memory — so the CPU copy still exists, and you now hold two copies of the data until you drop one. For a multi-GPU spread there is `faiss.index_cpu_to_all_gpus(cpu_index, co=..., ngpu=...)`. You can also skip the CPU step and construct a GPU index directly, e.g. `faiss.GpuIndexIVFFlat(res, d, nlist, faiss.METRIC_L2, config)`, which avoids ever materialising a CPU copy. That matters when the dataset only just fits: cloning needs both copies resident at the moment of the copy. ## The reference-lifetime trap The SWIG-generated Python wrapper does **not** take ownership of the resources object when you pass it in. If `res` is a local variable in a function that returns the index, Python garbage-collects it, the C++ index keeps a dangling pointer, and the next search crashes the interpreter — usually with no useful traceback, which is why this bug eats hours. Store `res` alongside the index (an attribute on the object that owns the index is the simplest fix). The same rule applies to the list of resources you build for a multi-GPU setup. ## Getting data back off the GPU `faiss.index_gpu_to_cpu(gpu_index)` produces an ordinary CPU index with the same contents. This is required before `faiss.write_index(...)`, because serialisation is defined for CPU index types only. The practical deployment pattern that follows: build or train on GPU (which is where the speedup is largest), convert down, write the file, and then let each serving replica decide whether to load it on CPU or clone it back onto its own device. It also gives you a fallback — the same artifact serves on a CPU-only host if a GPU box is unavailable. ## What to say in an interview Name the three moving parts — resources object, `index_cpu_to_gpu`, `index_gpu_to_cpu` — and then show operational awareness: the scratch pool is pre-reserved, the resources object must outlive the index, and persistence goes through the CPU form. Candidates who have only read the README describe the clone call and stop there.

  • What actually happens if the StandardGpuResources object is garbage-collected while the index is still in use?
    The C++ index keeps a raw pointer to freed resources — the stream, cuBLAS handle and scratch pool are gone. The next search dereferences that pointer and the process typically segfaults or reports an opaque CUDA error, with no Python traceback pointing at the cause. The fix is structural: store the resources object on whatever object owns the index so their lifetimes match.
  • Can you write a GPU index straight to disk with write_index?
    No. Serialisation is implemented for CPU index types only, so you call `faiss.index_gpu_to_cpu(gpu_index)` first and write that. It is a useful discipline anyway: the on-disk artifact is device-independent, so the same file can be loaded on a CPU-only host or cloned onto a different number of GPUs later.
  • When would you construct a GpuIndex class directly instead of cloning a CPU index?
    When the dataset only just fits in device memory, since cloning briefly requires the CPU and GPU copies to coexist, and when you want to set GPU-specific options at construction — a config object lets you pin the device, choose the indices storage mode, and enable half-precision paths that the generic clone call would otherwise decide for you.

saying these in an interview costs you the question

  • Thinking a GPU index can be written directly with write_index
  • Creating StandardGpuResources as a local and letting it be collected
  • Assuming index_cpu_to_gpu moves data rather than copying it
  • Believing FAISS picks up a GPU automatically if CUDA is installed
  • Creating a new resources object per query instead of reusing one

context