skip to content

Which FAISS index types have GPU implementations, and what happens when one does not?

level: middleimportance: should knowfreq 40%

answer

  1. only a subset is ported
  2. flat and IVF families, not graphs
  3. cloning an unsupported type errors out
  4. config struct carries device and precision
  5. top-k has a hard ceiling

basics

~20 s

FAISS implements a subset of its indexes on GPU — flat brute force, IVF with flat, scalar-quantized or product-quantized codes, and binary flat. Cloning a CPU index type with no GPU counterpart raises an error instead of silently falling back to CPU.

solid answer

~40 s

The GPU side of FAISS mirrors only part of the CPU catalogue: `GpuIndexFlatL2` / `GpuIndexFlatIP` for exact search, `GpuIndexIVFFlat`, `GpuIndexIVFScalarQuantizer` and `GpuIndexIVFPQ` for the inverted-file family, and `GpuIndexBinaryFlat` for Hamming search. Graph indexes built for CPU do not clone: `faiss.index_cpu_to_gpu` on an `IndexHNSWFlat` raises rather than quietly running on CPU, which is the behaviour you want — a silent fallback would look like a mysterious 20x regression. Each GPU class takes a config object (`GpuIndexIVFPQConfig`, for instance) carrying the device ordinal, how ids are stored, and half-precision options such as `useFloat16LookupTables`. GPU search also caps top-k in the thousands, so very large k values must be served on CPU. Plan the index type around GPU support up front rather than discovering the gap after building.

code

python · 11 lines
python
import faiss

d, nlist, m = 128, 4096, 32
res = faiss.StandardGpuResources()

config = faiss.GpuIndexIVFPQConfig()
config.device = 0
config.useFloat16LookupTables = True
config.indicesOptions = faiss.INDICES_32_BIT

index = faiss.GpuIndexIVFPQ(res, d, nlist, m, 8, faiss.METRIC_L2, config)

go deeper

for a junior

Remember that only some FAISS indexes run on GPU — flat and the IVF family — and that asking for an unsupported one raises an error rather than working slowly.

for a middle

Name the supported classes, explain that each takes a config object controlling device, id storage and half-precision paths, and say why an explicit error beats a silent CPU fallback.

for a senior

Show you design around the limits: index type chosen for what fits in device memory, top-k ceiling accounted for in the retrieval funnel, and PQ parameters verified against the GPU implementation before a build runs.

for a principal

Own the consequence for architecture — whether GPU is a hard serving dependency or a build accelerator, and whether the retrieval design survives being pushed back onto CPU when GPU capacity is unavailable.

## The GPU catalogue is a subset FAISS's CPU library has dozens of index classes; the GPU library re-implements a deliberately small set, because each one is a hand-written CUDA implementation rather than a compiled-down version of the CPU code. What exists: - **`GpuIndexFlat`** (with `GpuIndexFlatL2` and `GpuIndexFlatIP`) — exact brute-force search. This is where the GPU wins most spectacularly, because brute force is a dense matrix multiply and that is exactly what the hardware is built for. - **`GpuIndexIVFFlat`** — inverted file with uncompressed vectors in each list. - **`GpuIndexIVFScalarQuantizer`** — inverted file with scalar-quantized codes. - **`GpuIndexIVFPQ`** — inverted file with product-quantized codes, the memory-efficient option for large datasets. - **`GpuIndexBinaryFlat`** — brute-force Hamming distance over binary codes. A CPU index type outside that list has no GPU counterpart, and `faiss.index_cpu_to_gpu` throws. This is a design choice worth defending in an interview: the alternative — silently keeping the index on CPU — would produce a system that appears to work while delivering none of the speedup you provisioned hardware for. ## Config objects Every GPU index class takes an optional config struct, and the fields are where the real decisions live: - `device` — which GPU ordinal the index lives on. - `indicesOptions` — how the vector ids are stored. `faiss.INDICES_64_BIT` is the faithful default; `faiss.INDICES_32_BIT` halves id memory when your ids fit in 32 bits; `faiss.INDICES_CPU` keeps the id table in host memory entirely, trading a lookup for device RAM. - `useFloat16LookupTables` (IVFPQ) — keeps the distance lookup tables in half precision, cutting scratch usage and often speeding search, at a small accuracy cost. - `usePrecomputedTables` (IVFPQ) — precomputes part of the distance decomposition. Faster queries, but the table itself is sizeable and scales with `nlist`, so on a large index it can be the thing that pushes you into an out-of-memory error. Constructing directly, e.g. `faiss.GpuIndexIVFPQ(res, d, nlist, m, 8, faiss.METRIC_L2, config)`, gives you these knobs explicitly; cloning picks defaults for you (with `GpuClonerOptions` / `GpuMultipleClonerOptions` exposing similar fields). ## Limits that bite **Top-k is bounded.** The GPU selection kernels support k up to a few thousand (2048 in current versions); ask for more and the search is rejected. Rerank pipelines that pull 10,000 candidates from the vector stage have to either lower k, shard the retrieval, or run that stage on CPU. Check this before designing the retrieval funnel, not after. **Code parameters are constrained.** GPU IVFPQ does not accept every combination of sub-quantizer count and bits-per-code that the CPU implementation does — byte-aligned 8-bit codes are the well-trodden path. If you plan an exotic PQ configuration, verify it constructs on GPU before committing to it. **Everything must fit.** There is no paging: a GPU index holds all its codes and ids in device memory. That constraint, not raw compute, is usually what decides the index type — IVFPQ exists on GPU precisely because IVFFlat's uncompressed vectors stop fitting. ## How to choose Work backwards from the dataset. If it is small enough that exact search is affordable, `GpuIndexFlat` is both the fastest and the simplest thing on a GPU, and it needs no training at all. If it is not, you are in the IVF family, and the choice between flat, scalar-quantized and PQ codes is a memory-versus-accuracy decision measured against your device's capacity. If your design calls for a graph index, accept that it is a CPU serving path and use the GPU for building or for offline batch scoring instead. ## What interviewers listen for The strong answer names the supported families rather than reciting every class, states plainly that unsupported types raise instead of falling back, and connects the config fields to what they cost — id width, half-precision tables, precomputed tables — rather than treating them as decoration.

  • What does the indicesOptions setting on a GPU index config actually trade?
    It decides where and how the vector ids live. INDICES_64_BIT stores full 64-bit ids on the device; INDICES_32_BIT halves that when your ids fit, saving 4 bytes per vector — meaningful at a hundred million vectors; INDICES_CPU keeps the table in host memory so the device holds only codes, at the cost of a host-side lookup on every result. It is a pure memory-versus-indirection trade.
  • Why is exact flat search the case where GPUs help most?
    Brute-force search is a dense matrix multiply between the query batch and the whole database, which maps directly onto the hardware's throughput and needs no branching or irregular memory access. Approximate indexes replace that arithmetic with lookups and gathers, which use the GPU far less efficiently — so the speedup over CPU is largest exactly where there is the most arithmetic to do.
  • Your reranking design wants the vector stage to return 10,000 candidates. What does that mean for GPU serving?
    GPU search caps k in the low thousands, so a single call cannot return 10,000. Options are to lower the candidate count and lean harder on the reranker, split retrieval across shards and merge the per-shard top-k, or serve that stage on CPU where k is unbounded. Deciding this before building the index avoids a rewrite.

saying these in an interview costs you the question

  • Assuming every FAISS index type has a GPU version
  • Expecting a silent CPU fallback when the clone is unsupported
  • Thinking a GPU index can page vectors from host memory on demand
  • Ignoring that top-k is capped on GPU search
  • Treating the config object as optional boilerplate with no memory impact

context