skip to content

In FAISS, why wrap an index in IndexIDMap, and what does that change?

level: middleimportance: should knowfreq 52%

answer

  1. Default ids are positions, not keys
  2. Flat indexes reject add_with_ids
  3. Wrapper keeps a parallel id table
  4. IDMap2 adds the reverse lookup
  5. Ids are int64, so map strings yourself

basics

~20 s

FAISS numbers vectors sequentially from zero in insertion order, so a flat index cannot return your application's identifiers. IndexIDMap wraps an index so add_with_ids attaches your own 64-bit integer ids, which search then returns instead of positions, at the cost of an extra id table in memory.

solid answer

~50 s

By default a FAISS index labels vectors by their ordinal position: the first vector added is 0, the second is 1, and `search` returns those positions in the id matrix. That is useless if your rows have database keys, so you wrap the index — `faiss.IndexIDMap(faiss.IndexFlatL2(d))` — and call `add_with_ids(xb, ids)` with an int64 array. The wrapper keeps a vector of your ids parallel to the internal positions and translates on the way out. `IndexIDMap2` additionally maintains the reverse map, which is what makes `reconstruct(my_id)` work. Two limits matter. Ids must be 64-bit integers, so UUIDs and string keys need your own external mapping table. And removal, via `remove_ids` with a selector such as `faiss.IDSelectorBatch`, only works when the wrapped index actually implements deletion — IVF does, HNSW does not. The factory spellings `"IDMap,Flat"` and `"IDMap2,Flat"` build the same thing.

code

python · 14 lines
python
import numpy as np
import faiss

d = 64
xb = np.random.random((1000, d)).astype('float32')
ids = np.arange(100000, 101000).astype('int64')

index = faiss.IndexIDMap2(faiss.IndexFlatL2(d))
index.add_with_ids(xb, ids)

D, I = index.search(xb[:1], 3)
print(I[0])                       # ids in the 100000 range, not 0,1,2
vec = index.reconstruct(100005)   # reverse map: IDMap2 only
print(vec.shape)

go deeper

for a junior

Know that FAISS labels vectors 0, 1, 2 in insertion order by default, and that IndexIDMap is how you get your own ids back out of a search.

for a middle

Explain that the wrapper stores a parallel int64 id table, that IDMap2 adds the reverse map needed by reconstruct, and that IVF indexes already accept ids natively.

for a senior

Show you have run this in production: the external key-to-int64 table is a persisted artifact, deletions depend on the wrapped index supporting removal, and HNSW simply does not.

for a principal

Own the identity model as a system concern — where the key mapping lives, how it is backed up and rebuilt, and what the deletion story is before the index type is chosen.

## The default: positional ids FAISS is a library over arrays, not a database. When you `add(xb)` to an index, each row is assigned the next sequential integer starting at 0, and `search` returns those integers in `I`. If you added your corpus in a known order you can map back by position — but that arrangement is brittle. It breaks the moment you add a second batch out of order, delete something, rebuild the index from a different query, or shard the corpus across processes. The symptom people hit first is more direct: calling `add_with_ids` on a plain `IndexFlatL2` raises an error, because flat indexes have nowhere to put ids. That is the moment `IndexIDMap` enters. ## What the wrapper does `IndexIDMap` is a thin index that owns another index. It forwards `add_with_ids(xb, ids)` to the child's plain `add`, and appends your ids to an internal `id_map` vector in the same order. On `search`, it calls the child, gets back internal positions, and rewrites them into your ids before returning. `ntotal`, `d` and the metric come from the child. The cost is one int64 per vector — 8 bytes — plus the indirection. That is negligible next to raw vectors, but it is not free at very large scale on a compressed index where the codes themselves are only 32–96 bytes. `IndexIDMap2` extends this with a hash map from your id back to the internal position. That reverse direction is what `reconstruct(id)` needs: without it, the wrapper knows how to translate outward but not inward. Use `IndexIDMap2` whenever you want to fetch a stored vector back by its application id; use plain `IndexIDMap` when you only ever search. ## Which indexes need it Not all of them. `IndexIVFFlat`, `IndexIVFPQ` and the other IVF variants implement `add_with_ids` natively — they already store an id alongside each code inside the inverted lists, so wrapping them in `IndexIDMap` is redundant and just adds a second copy of your ids. The wrapper is for indexes that store vectors positionally: the flat family, and graph indexes. A useful mental check: if the index type needs to remember an id per entry to do its own job, it supports `add_with_ids`; if it is a bare array or a graph over positions, it does not. ## The int64 constraint FAISS ids are signed 64-bit integers, and `-1` is reserved to mean "no result" in the search output. Real applications rarely have int64 primary keys for documents — they have UUIDs, URLs, or composite keys. There is no facility in FAISS for that: you maintain the mapping yourself, typically a dict or a small database table from your key to a monotonically increasing int64, and you persist it alongside the index. Losing that table is equivalent to losing the index, which is why it belongs in the same backup and the same deployment artifact. A related trap is packing meaning into ids — for example, using the high bits as a tenant id. It works, and people do it, but the ids are also used internally in some index types and the packing must survive every rebuild, so it is a maintenance burden that a plain lookup table avoids. ## Deletion `remove_ids(selector)` deletes entries matching a selector: `faiss.IDSelectorBatch(ids)` for an explicit set, `faiss.IDSelectorRange(lo, hi)` for a contiguous span. It returns the number of entries removed. Through an `IndexIDMap` this is delegated to the wrapped index, so it works only when the child supports removal. IVF indexes do — they can drop an entry from an inverted list cheaply. `IndexHNSWFlat` does not implement removal at all, because deleting a node would break the graph's navigability; deletions there mean rebuilding, or maintaining a tombstone set in your application and filtering results after search. A subtlety worth mentioning: on indexes where removal compacts storage, internal positions shift, which is exactly why you should never rely on positional ids in a mutable index. The id map absorbs that churn. ## Putting it together The practical pattern for a small mutable corpus is: `IndexIDMap2` over `IndexFlatL2`, application keys mapped to int64 through a table you own and persist, `add_with_ids` for inserts, `remove_ids` with an `IDSelectorBatch` for deletes, and `reconstruct` when you need the stored embedding back. For a large IVF-based corpus, skip the wrapper entirely and use the IVF index's native `add_with_ids`. Knowing which of those two shapes applies — and why the wrapper exists at all — is the actual content of this question.

  • Your document keys are UUID strings. How do you index them in FAISS?
    FAISS ids are int64 only, so you keep your own bidirectional table: assign each UUID a monotonically increasing integer, add vectors with those integers, and translate search results back on the way out. Persist that table with the index — in the same artifact or the same database — because losing it makes the index unusable. Hashing UUIDs into int64 is tempting but risks collisions that silently return the wrong document.
  • Do you need IndexIDMap when using IndexIVFFlat?
    No. IVF indexes store an id alongside every entry in their inverted lists, so they implement add_with_ids natively and accept your int64 ids directly. Wrapping one in IndexIDMap just adds a redundant second id table and 8 more bytes per vector. The wrapper is for indexes that identify vectors by storage position, such as the flat family.
  • What is the difference between IndexIDMap and IndexIDMap2?
    IndexIDMap keeps only the forward direction — internal position to your id — which is all search results need. IndexIDMap2 additionally maintains the reverse map from your id to the internal position, which is what reconstruct(id) requires to fetch a stored vector back. IDMap2 costs a little more memory; choose it when you need lookup by id, not only search.

saying these in an interview costs you the question

  • Assumes search returns application keys by default
  • Calls add_with_ids on a plain flat index
  • Thinks FAISS accepts string or UUID identifiers
  • Expects reconstruct to work on plain IndexIDMap
  • Believes every index type supports remove_ids

context