How do you decide whether a FAISS deployment justifies GPUs at all?
answer
- three separate axes, not one decision
- must fit, no paging, no swap
- batchable work or nothing
- building may be the real bottleneck
- cost per query at target recall
basics
~20 sDecide from workload shape, not hardware envy. GPUs win where there is bulk parallel work — index building, clustering, batch scoring, high-QPS batched retrieval — and where the index fits in device memory. Low-QPS single-query serving on a tuned CPU index is usually cheaper and simpler.
solid answer
~50 sFrame it as three separate questions. **Does it fit?** A GPU index must hold its codes, ids and auxiliary tables in device memory with headroom for per-query workspace; if not, you are buying a multi-device sharded deployment, which changes the cost and failure model. **Is there enough parallel work?** GPU FAISS is a throughput engine — bulk workloads and batched high-QPS retrieval convert directly into value, while a trickle of single online queries mostly buys per-call overhead. **Is build time the real bottleneck?** Training a coarse quantizer and adding hundreds of millions of vectors is often where the GPU earns its keep, even when serving stays on CPU. That split — build on GPU, `index_gpu_to_cpu`, serve on CPU — is a very common and defensible answer. Then compare cost per query against a tuned CPU fleet, and weigh the operational tax: driver and CUDA version pinning, scarcer capacity, and no paging when memory runs out.
go deeper
Know that GPUs help FAISS most on bulk parallel work and require the whole index to fit in device memory — they are not a general speed switch for every deployment.
Be able to separate the axes: capacity (does it fit), throughput (is the work batchable), and build time, and name the CPU-serving case as a legitimate outcome.
Argue from measurement — achievable QPS at your recall target and batch size on each option — and design the split deployment plus a tested CPU fallback for when GPU capacity is unavailable.
Own the whole tradeoff: cost per query including utilisation, the operational tax of driver and capacity coupling, the hard memory ceiling with no swap, and the measurable triggers that would reverse the decision.
## Ask three questions, in order ### 1. Does the index fit? GPU FAISS has no paging: everything the index needs must live in device memory simultaneously — the vectors or compressed codes, the id table, the coarse centroids, any precomputed tables, plus headroom for per-query workspace at your real batch size and concurrency. Work out that number before anything else, because it determines whether the conversation is about one GPU or about a sharded multi-device deployment with its own redundancy requirements and its own failure mode (a lost shard silently removes part of the corpus from every result). If a compressed index fits comfortably, GPU is on the table. If fitting requires aggressive quantization that costs recall you cannot afford, a CPU deployment with far more RAM per dollar may serve the same recall target more cheaply. ### 2. Is there enough parallel work? GPUs are throughput machines, and FAISS's GPU indexes are written for large query batches. Score your workload honestly: - **Bulk offline work** — corpus-against-corpus matching, deduplication, nightly re-scoring, evaluation sweeps — is the ideal fit. Millions of queries, no latency constraint, batches as large as memory allows. - **High-QPS retrieval that can be batched** works well, with a dynamic batching layer trading a few milliseconds of latency for a large throughput multiple. - **Low-QPS single-query serving with a tight p99** is the weak case. Per-call transfer and launch overhead dominate, and a well-tuned CPU inverted-file index answers in a millisecond or two with no batching machinery to build and operate. The honest principal answer often ends "not for serving" — and that is a stronger answer than reflexively provisioning GPUs. ### 3. Is building the bottleneck? This is the case people most often miss. Training a coarse quantizer over a large sample and adding hundreds of millions of vectors is heavy, parallel, latency-insensitive work — exactly what a GPU is for. If your pain is a nightly rebuild that takes eight hours, a GPU may compress it to under an hour without any change to how you serve. `faiss.index_gpu_to_cpu` then `faiss.write_index` gives you a device-independent artifact that CPU serving hosts load unchanged. Build on GPU, serve on CPU is a mature, common shape and it decouples the two decisions entirely. ## The cost comparison to actually run Compare **cost per query at your recall target**, not peak queries per second. That means: measure achievable QPS on the GPU configuration at the recall you require and at your production batch size, do the same for a tuned CPU configuration on comparably priced instances, and divide by hourly cost. GPU instances carry a large premium, so a 10x throughput advantage is not automatically a win. Include utilisation — a GPU idle two thirds of the day is paid for around the clock, whereas a CPU fleet scales down. ## The operational tax Beyond price, GPUs add real friction that belongs in the decision: - **Version coupling.** Driver, CUDA runtime and FAISS build must agree; upgrades become coordinated events rather than a package bump. - **Capacity scarcity.** GPU instance types are harder to obtain and to autoscale than general-purpose CPU instances, in most environments. - **Hard memory ceiling.** Exceeding device memory is a failure, not a slowdown — there is no swap. Growth planning must be explicit. - **Narrower index catalogue.** Only part of FAISS's index selection has a GPU implementation, so a GPU commitment constrains future index choices. - **A fallback path is mandatory.** Serving must degrade to CPU when GPU capacity is unavailable, which means keeping the CPU-serialised artifact and testing that path. ## What a principal-level answer sounds like It separates the axes — capacity, throughput, build time — instead of treating "use GPUs" as one decision. It quotes a cost-per-query method rather than a benchmark headline. It names the split deployment (build on GPU, serve on CPU) as a first-class option rather than a compromise. And it states the conditions under which it would revisit the choice — corpus growth past device memory, a change in traffic shape from batchable to strictly online, or a recall target that compression can no longer meet. There is no single right answer here, and an interviewer asking this is checking whether you can hold several constraints at once and still commit to one.
- What does the build-on-GPU, serve-on-CPU pattern look like in practice?Train the coarse quantizer and add vectors on a GPU index, convert with faiss.index_gpu_to_cpu, then faiss.write_index to produce a device-independent artifact. Serving hosts load that file on CPU. You get the GPU's advantage where it is largest — bulk build work — while serving stays on cheap, plentiful, easily autoscaled hardware, and the two decisions stop being coupled.
- Why compare cost per query rather than peak queries per second?Peak QPS ignores both price and utilisation. A GPU instance costing several times a CPU instance needs a proportionally larger throughput advantage to win, and a GPU idle for much of the day is still paid for hourly while a CPU fleet scales down. Measure achievable QPS at your required recall and production batch size on each option, divide by hourly cost, and compare that.
- What would make you revisit a decision to serve FAISS on CPU?Corpus growth that pushes the working set past what CPU memory serves at acceptable latency; a traffic shift toward batchable bulk retrieval where GPU throughput converts directly into savings; a recall target that now requires probing far more lists per query; or a rebuild window that has grown past its schedule. Each maps to one of the three axes, so the trigger is measurable rather than a matter of taste.
saying these in an interview costs you the question
- Provisioning GPUs from a benchmark headline without a cost-per-query comparison
- Assuming a GPU lowers p99 latency for single online queries
- Forgetting a GPU index cannot page or swap when it outgrows the device
- Overlooking build and training as the workload GPUs help most
- Committing to GPU serving with no tested CPU fallback path