How do you persist a FAISS index with write_index and read_index, and what is not saved?
answer
- two module-level functions, not a method
- index objects are not picklable
- trained state travels with the file
- vectors and ids only, no payloads
- memory-map when it will not fit
basics
~20 sfaiss.write_index(index, path) serialises the whole index — trained centroids, codebooks, codes and ids — and faiss.read_index(path) restores it ready to search, with no retraining. What it never stores is your documents or metadata; FAISS keeps only vectors and 64-bit ids.
solid answer
~40 sPersistence is two calls: `faiss.write_index(index, "vectors.faiss")` and `faiss.read_index("vectors.faiss")`. The file carries everything the index needs, including the learned coarse centroids and PQ codebooks, so a restored IVF index has `is_trained == True` and can be searched immediately — training is a build-time cost you pay once. FAISS objects are not picklable, so this is the supported route. The important gap is that FAISS stores vectors and integer ids and nothing else: the mapping from id to document text, URL or metadata lives in your own store, and it must be written, versioned and deployed as one unit with the index file, or you will serve confident results pointing at the wrong documents. For indexes too large to hold in RAM, `read_index` accepts `faiss.IO_FLAG_MMAP` to memory-map instead of loading.
code
python · 16 linesimport faiss
import numpy as np
d = 64
xb = np.random.random((10_000, d)).astype('float32')
index = faiss.IndexIVFFlat(faiss.IndexFlatL2(d), d, 64)
index.train(xb)
index.add(xb)
faiss.write_index(index, "vectors.faiss")
loaded = faiss.read_index("vectors.faiss")
print(loaded.ntotal) # 10000
print(loaded.is_trained) # True - centroids came back with the file
loaded.nprobe = 8
D, I = loaded.search(xb[:5], 10)go deeper
Know the two calls, faiss.write_index and faiss.read_index, and be able to say that the file holds vectors and ids but not your documents or metadata.
Explain that trained state travels in the file, so a reloaded IVF or PQ index must not be retrained, and that FAISS objects are not picklable. Mention IO_FLAG_MMAP for indexes larger than RAM.
Show the deployment discipline: atomic write-then-rename, immutable versioned artifacts, a load-time consistency check between ntotal and the id mapping, and periodic rebuilds rather than indefinite in-place mutation.
Own the artifact contract — how index files are built, versioned against corpus and embedding-model versions, validated for recall before promotion, and rolled back — so that a bad index never becomes an outage nobody can attribute.
## The API FAISS persistence is deliberately minimal: ``` faiss.write_index(index, "vectors.faiss") index = faiss.read_index("vectors.faiss") ``` `write_index` serialises the concrete index type and all of its state into a single binary file; `read_index` reconstructs the right subclass from that file. There is no `index.save()` method, and FAISS index objects are not picklable — attempting to pickle one fails or produces something that will not survive a round trip, so the module-level functions are the only supported path. ## What the file contains Everything the index needs to answer queries: - the index type and dimension, - for IVF indexes, the trained coarse-quantizer centroids and the inverted lists, - for PQ indexes, the learned codebooks and every stored code, - for HNSW indexes, the graph links, - any pre-transform such as an OPQ rotation matrix, - the 64-bit ids, when the index is wrapped in `IndexIDMap` or uses `add_with_ids`. Because the trained parameters are included, a reloaded IVF or PQ index reports `is_trained == True` and must never be trained again. Re-training would discard the centroids the stored codes were computed against and silently corrupt every result. This is why the standard production shape is: train and build offline, write the file, and have serving processes only read it. ## What the file does not contain FAISS is a similarity-search library, not a database. It stores vectors and ids. It does not store the text a vector was embedded from, any metadata, any filterable attributes, or any notion of a document. That external mapping — id to payload — is your responsibility, and the operational consequence is sharper than it sounds: if the index file and the id-to-document store drift apart by even one insertion, searches return plausible distances attached to the wrong content, and nothing errors. Ship them as one versioned artifact, generated by the same build, and validate on load (for example, that `index.ntotal` matches the row count of the mapping) before the process starts serving. The file also implicitly depends on the embedding model that produced the vectors. A new model version means new vectors, which means a rebuild — record the model version alongside the index file so that a mismatch is detectable rather than silent. ## Loading large indexes `read_index` loads the whole structure into process memory by default, which is exactly what you want for an index that comfortably fits. When it does not, pass a flag: ``` index = faiss.read_index("vectors.faiss", faiss.IO_FLAG_MMAP) ``` This memory-maps the file instead of copying it, so pages are demand-loaded by the OS and several processes on the same host can share them. The tradeoff is that queries touching cold pages pay disk latency, so it suits large IVF indexes where a query scans a small fraction of the data, and suits it badly when random access is broad. ## Deployment hygiene A few habits prevent most production incidents around index files: - **Write atomically.** Write to a temporary path and rename into place, so a reader never opens a half-written file. Rename on the same filesystem is atomic; copying over a live file is not. - **Treat the file as immutable.** Never mutate a file that a serving process has open. Publish a new file, then swap. - **Version the artifact.** Name or tag it with the corpus snapshot and the embedding model version, and store the measured recall next to it, so a regression is attributable. - **Check load time and memory.** Reading a multi-gigabyte index takes real seconds and doubles nothing but does allocate the full structure; size container memory limits against the file size plus overhead, and account for the load in readiness probes. - **Rebuild rather than patch.** FAISS supports adding vectors to a loaded index, but a long-lived, repeatedly-mutated index drifts away from the distribution its centroids were trained on. Periodic full rebuilds keep recall predictable. ## The mental model to carry A FAISS index file is a compiled artifact: expensive to produce, cheap to load, and meaningless without the id mapping and the embedding model it was built with. Treating it that way — built by a pipeline, versioned, validated, swapped atomically — is what separates a FAISS deployment that behaves from one that mysteriously returns the wrong documents after a deploy.
- After read_index restores an IVF index, do you need to call train() again?No, and you must not. The coarse centroids and any PQ codebooks are serialised with the index, so is_trained comes back True. Calling train() again refits the centroids while the stored codes still refer to the old ones, which corrupts results silently rather than raising. Training is strictly a build-time step; serving processes only read.
- How do you keep a FAISS index and its document metadata in sync across deploys?Build them together and ship them as one versioned artifact — the index file plus the id-to-document store, tagged with the corpus snapshot and embedding-model version. Validate on load that ntotal matches the mapping size before the process serves traffic, and swap atomically by writing to a temporary path and renaming. Drift between the two produces wrong documents with entirely plausible distances.
- When is faiss.IO_FLAG_MMAP the right way to load an index?When the index is large relative to available RAM and each query touches only a small slice of it, which is the typical IVF pattern at modest nprobe. Memory-mapping lets the OS demand-load pages and lets several processes on a host share them. It is a poor fit when access is broad and random, since every cold page becomes disk latency inside the query.
saying these in an interview costs you the question
- Expecting FAISS to store the source text or metadata alongside vectors
- Pickling index objects instead of using write_index
- Retraining an index after loading it from disk
- Overwriting a live index file in place instead of renaming a new one
- Assuming ids are preserved without IndexIDMap or add_with_ids