What does LanceDB do when you search a table with no vector index?
answer
- No index is a legal state here
- Search still answers, it just scans
- Cost scales with the row count
- Exactness is what indexing spends
- Flat scan is your free ground truth
basics
~20 sIt runs an exhaustive brute-force scan, comparing the query against every vector in the table. Results are exact, but latency grows linearly with row count. Calling create_index builds an approximate IVF_PQ index that trades that exactness for sub-linear search.
solid answer
~50 sLanceDB does not index automatically. Until you call `create_index`, every `search()` is a full scan of the vector column — exact nearest neighbours, latency proportional to rows. That is genuinely fine at small scale: Lance stores vectors columnar on disk, so a scan of a few tens of thousands of rows is fast, and you get perfect recall for free. Past that, scan time becomes the query time and you build an index, typically `index_type="IVF_PQ"`, which clusters vectors into partitions and stores compressed codes so a query touches a fraction of the data. The moment you do, results become **approximate**: the same API returns neighbours that may differ from the flat-scan answer, and nothing in the response tells you which mode ran. So before indexing, capture a flat-scan baseline for a sample of queries and measure recall against it afterwards.
code
python · 16 linesimport numpy as np
# exact top-k while the table is unindexed: use it as ground truth
queries = np.random.rand(200, 768).astype("float32")
truth = [
{r["id"] for r in tbl.search(q).limit(10).to_list()}
for q in queries
]
tbl.create_index(index_type="IVF_PQ", num_partitions=256, num_sub_vectors=96)
hits = sum(
len(truth[i] & {r["id"] for r in tbl.search(q).limit(10).to_list()})
for i, q in enumerate(queries)
)
print("recall@10:", hits / (len(queries) * 10))go deeper
Know that a LanceDB table works without any index and that search then compares against every row. Be able to say that create_index is what makes large tables fast.
Explain the tradeoff in both directions: the unindexed scan is exact and linear, the IVF_PQ index is approximate and sub-linear, and the API call looks identical either way.
Demonstrate the operational habit: capture flat-scan ground truth first, measure recall@k after building, and own a re-index or optimize cadence so post-build writes do not quietly return you to a scan.
Frame it as a cost model rather than a setting. Index size, cache residency and re-index frequency drive both the storage bill and the latency SLO, especially on object storage; decide the acceptable recall target with product before tuning anything.
## No index is the default state A freshly created LanceDB table has data but no vector index. Search still works — this surprises people coming from engines that refuse to answer until an index exists. LanceDB simply scans: it reads the vector column, computes the distance from the query to every stored vector, and keeps the best k. This is often called a flat or brute-force search. ## Why that is a reasonable default Lance is a columnar on-disk format, so a vector scan reads one tightly packed column rather than whole rows. At tens of thousands of vectors that scan completes in milliseconds to low tens of milliseconds on local disk, and you get two things an index cannot give you: perfect recall, and zero build cost. During development, while embeddings and chunking are still churning, that is exactly the tradeoff you want — no stale index to rebuild every time the data changes. It also means the index is a pure performance decision. Adding one never changes the API and never changes what a query means; it changes latency and answer quality. ## When the scan stops being acceptable Flat-scan cost is roughly linear in rows times dimensions, so it degrades predictably. The signals to move are: query latency tracking table growth one-for-one; p99 latency dominated by the scan rather than by embedding the query; or data living on object storage, where reading the whole vector column per query is expensive in both time and request count. LanceDB's design leans on the index staying small enough to cache, which matters most in exactly that object-storage case. ## What building an index changes `create_index` on the table builds, by default in practice, an IVF-style index — commonly `IVF_PQ`, which partitions the vectors by k-means and stores each vector as a compressed product-quantized code. Building it requires a training pass over the data, so it is not instantaneous, and it needs enough rows for the clustering to be meaningful; training a large number of partitions on a tiny table produces a bad index. After the build, a query probes only a handful of partitions and scores mostly against compressed codes. Latency drops sharply and stops tracking the row count so directly. What you have given up is exactness: the true nearest neighbour may sit in a partition the query did not probe, or may be misordered because the compressed distance is approximate. This is the central tradeoff of the whole product surface, and interviewers ask about it because engineers routinely ship an index without ever measuring what it cost them. ## Measuring the cost The practical discipline is simple. Before indexing, take a representative sample of query vectors — a few hundred — and record the exact top-k from the flat scan. That is your ground truth, and it is free precisely because the unindexed table is exact. After indexing, run the same queries and compute recall@k: the fraction of the true neighbours that the approximate search returned. Tune from there. Without that baseline, "the index made search worse" is a complaint you can never quantify or defend. ## Rows added after the index exists An index is built over the data present at build time. Rows written afterwards land in fragments the index does not cover. LanceDB does not silently drop them from vector results — it combines the indexed search with a scan over the unindexed portion — but that scan grows with every write, so query latency drifts upward the longer you go without maintenance. The fix is to fold new data into the index, either by re-running the index build or by running the table's optimize step, on a schedule matched to your ingest rate. Teams that ingest continuously and never re-index end up back at a brute-force scan while believing they have an index. ## Rules of thumb Small and static: no index, exact answers, less to operate. Large or growing, or stored remotely: index, and own the recall measurement and the re-index cadence as part of the deployment rather than as a one-off tuning exercise.
- How are rows inserted after create_index handled by a later vector search?They are not lost. LanceDB searches the index and also scans the fragments the index does not yet cover, merging both. Results stay complete, but that extra scan grows with every write, so latency drifts upward as ingest continues. Folding new data in — re-running the index build or the table's optimize step on a schedule — is required maintenance, not an optimization.
- At roughly what table size would you stop relying on the flat scan?There is no universal number; it depends on dimensionality, storage medium and latency budget. The honest answer is to measure: plot p95 query latency against row count and index when the scan starts to dominate your budget. Order of magnitude, local-disk tables of tens of thousands of vectors are comfortable unindexed, and object-storage-backed tables want an index much earlier because every query otherwise pulls the whole vector column.
saying these in an interview costs you the question
- Claiming search fails until an index exists
- Assuming LanceDB indexes automatically on write
- Believing an indexed search is still exact
- Never measuring recall against a flat-scan baseline
- Building the index once and never re-indexing after ingest