Vector
Stores built for nearest-neighbour search over embeddings rather than exact lookups, which is what makes semantic search and RAG possible. Interviewers use this area to check you understand approximate search — recall traded against latency — and how filtering and hybrid scoring change the picture.
on this pageshowhide
explore
- Pinecone20 questions
- Indexes and Namespaces5 questions
- Upsert and Query API6 questions
- Metadata Filtering4 questions
- Hybrid Sparse-Dense Search5 questions
- Weaviate23 questions
- Collections and Schema6 questions
- Vectorizers and Modules6 questions
- Queries and Filters6 questions
- Hybrid Search5 questions
- Qdrant18 questions
- Collections and Points6 questions
- Payload Filtering6 questions
- HNSW Config and Quantization6 questions
- Chroma17 questions
- Collections and Embedding Functions5 questions
- Querying and Filtering6 questions
- Persistence and Deployment6 questions
- FAISS16 questions
- Index Types in Practice5 questions
- Training and Recall Tuning5 questions
- GPU and Scaling6 questions
- LanceDB11 questions
- Tables and the Lance Format5 questions
- Search and Indexing6 questions
questions
105 · 6 sectionsIn Pinecone, what do the dimension and metric arguments to create_index lock in?
basics
~20 screate_index fixes an index's vector dimension and distance metric permanently. Every vector you upsert must have exactly that length or the write is rejected, and switching either setting means creating a new index and re-upserting the data.
What does a Pinecone vector record contain, and how should you batch upserts?
basics
~20 sA Pinecone record is an id string, a values list whose length must equal the index dimension, and optional metadata. index.upsert() takes a list of records, so send batches of roughly 100 to stay under the request size cap.
Why must a Pinecone index use the dotproduct metric for hybrid sparse-dense search?
basics
~20 sPinecone accepts sparse values only in indexes created with metric="dotproduct". Dot product is linear, so scaling the dense and sparse query vectors weights their contributions predictably; cosine and euclidean indexes reject sparse vectors, and the metric cannot be changed later.
In Pinecone, how do you upsert and query a record holding both dense and sparse vectors?
basics
~20 sA hybrid record carries the dense embedding in values and a sparse vector in sparse_values as {"indices": [...], "values": [...]}. The query passes vector= and sparse_vector= together, and Pinecone returns one ranked list with a single combined score per match.
What is a Pinecone namespace, and how does it change query behaviour?
basics
~20 sA namespace is a logical partition inside a Pinecone index. Each upsert and each query names one namespace, and a query searches only that partition — there is no cross-namespace search in a single call. Namespaces are created implicitly on first write.
In Weaviate, what is a collection and what does collections.create define?
basics
~20 sA collection is Weaviate's typed container for objects of one kind, much like a table. collections.create names it and declares its properties with data types, plus per-collection settings such as the vector index, replication and multi-tenancy.
In Weaviate's hybrid() query, what does the alpha parameter control?
basics
~20 salpha weights the two halves of a Weaviate hybrid search: alpha=0 is pure BM25 keyword search, alpha=1 is pure vector search, and values in between blend them. The v4 Python client's default of 0.7 leans toward the vector side.
In Weaviate, when do you use near_text, near_vector, or near_object?
basics
~10 snear_text sends a raw string and lets the collection's configured vectorizer embed it server-side. near_vector takes an embedding you computed yourself. near_object finds the neighbours of an object already stored, referenced by its UUID.
In Weaviate, what does setting a text2vec vectorizer on a collection do?
basics
~20 sA configured vectorizer makes Weaviate call the embedding model itself: you insert plain properties and the server produces the vector, and a text query is embedded server-side too. Without one you must supply every vector yourself.
Why is relying on Weaviate's auto-schema risky for a production collection?
basics
~20 sAuto-schema is on by default and invents property definitions from whatever object arrives first, so types are guessed from one sample, stray keys become permanent indexed properties, and the resulting schema differs between environments. Property types cannot be changed afterwards.
What are the parts of a Qdrant point, and which id types does it accept?
basics
~20 sA Qdrant point is the unit of storage: an id, one or more vectors, and an arbitrary JSON payload. Ids must be an unsigned integer or a UUID — arbitrary strings are rejected. Upserting an existing id replaces that point.
In Qdrant, how do you restrict a vector search to points matching a payload condition?
basics
~10 sPass a filter to the search call. In the Python client that is query_points(..., query_filter=models.Filter(must=[models.FieldCondition(key="lang", match=models.MatchValue(value="en"))])). Only points whose payload satisfies the filter come back, still ranked by vector distance.
In Qdrant, what do VectorParams size and distance fix at collection creation?
basics
~20 sVectorParams pins a Qdrant collection's vector dimension and its distance metric for the collection's lifetime. Every upserted point must match that size exactly, and neither value can be changed later — switching embedding models means creating a new collection.
In Qdrant, what do hnsw_config's m and ef_construct control, and what do they cost?
basics
~20 sm is how many graph links Qdrant keeps per vector (default 16); ef_construct is how wide the neighbour search is while building (default 100). Raising m costs RAM permanently, raising ef_construct costs indexing time only.
How do must, should, and must_not combine inside a Qdrant Filter?
basics
~20 sClauses are ANDed with each other: a point must satisfy every condition in must, at least one in should, and none in must_not. min_should raises the should bar from one to N. Nesting Filter objects inside should expresses OR-of-ANDs.
In Chroma, what happens to raw text passed to collection.add(documents=...)?
basics
~20 sChroma runs each document through the collection's embedding function and stores the resulting vector next to the text, its id and its metadata. ids are required and must be unique; documents, metadatas and embeddings are parallel lists.
In Chroma, what is the difference between EphemeralClient and PersistentClient?
basics
~20 sEphemeralClient keeps collections in memory only and loses everything when the process exits. PersistentClient(path="./chroma") writes to that directory and reloads it on the next run. Plain chromadb.Client() is ephemeral, which is why prototype data disappears.
In Chroma, when do you pass query_embeddings instead of query_texts to collection.query()?
basics
~20 squery_texts hands raw strings to Chroma, which embeds them with the collection's embedding function. query_embeddings takes vectors you already computed, so Chroma skips embedding entirely. Use it when the vector comes from elsewhere or is cached.
Why must a Chroma collection keep the same embedding_function for its whole life?
basics
~20 sDistances are only meaningful between vectors from the same model. Mixing embedding functions in one Chroma collection puts incomparable vectors in one index, so nearest-neighbour results become arbitrary — and if the dimensions happen to match, nothing errors.
In Chroma's query(), how do the where and where_document filters differ?
basics
~20 swhere filters on the structured metadata dictionary attached to each record, using operators like $eq, $in and $gte. where_document filters on the stored document text itself, with $contains and $not_contains substring matching. They are separate arguments and combine with AND.
In FAISS, what does IndexFlatL2 do and when is it the right index?
basics
~20 sIndexFlatL2 stores every vector uncompressed and compares a query against all of them, returning exact L2 nearest neighbours with no training step. It is the right choice for small collections and as a recall baseline, but its cost grows linearly with the dataset.
How do you move a FAISS index onto a GPU, and what does StandardGpuResources do?
basics
~20 sCreate a faiss.StandardGpuResources() object, then call faiss.index_cpu_to_gpu(res, device, index). The resources object owns the GPU scratch memory and cuBLAS handles for that device, and you must keep a Python reference to it for as long as the index lives.
What does the FAISS index_factory string "IVF4096,PQ64" build?
basics
~20 sIt builds an IVFPQ index: a coarse quantizer that splits the space into 4096 inverted lists, with each vector stored as a product-quantized code of 64 one-byte values instead of the raw floats. Both stages must be trained before adding data.
Why must a FAISS IVF index be trained before add(), and on what data?
basics
~20 sIVF and PQ indexes learn structure from data — k-means centroids for IVF, codebooks for PQ. Until train() runs, is_trained is False and add() raises. Train on a representative sample of the vectors you will actually search.
When do you shard a FAISS index across GPUs instead of replicating it?
basics
~20 sReplicate when the index fits on one GPU and you need more queries per second — each device holds a full copy and queries are split between them. Shard when the index does not fit on one GPU: each device holds a slice, every query touches all of them, and results are merged.
In LanceDB, how do you run a vector search and read the results as a DataFrame?
basics
~20 sCall table.search(query_vector), chain .limit(n) to say how many neighbours you want, and finish with .to_pandas(). Every returned row carries the table's own columns plus a _distance column holding that row's distance from the query vector.
In LanceDB, how do you create a table and where can its schema come from?
basics
~20 sdb.create_table(name, data=...) infers an Arrow schema from a pandas DataFrame, PyArrow table or list of dicts. Passing schema= with a PyArrow schema or a lancedb.pydantic LanceModel defines it explicitly and allows creating an empty table.
In LanceDB's create_index, what do num_partitions and num_sub_vectors control?
basics
~20 snum_partitions sets how many clusters the vectors are split into, so it governs how much data one query touches per probe. num_sub_vectors sets how many pieces each vector is chopped into for compression, so it governs index size and how lossy stored distances are.
What does LanceDB do when you search a table with no vector index?
basics
~20 sIt runs an exhaustive brute-force scan, comparing the query against every vector in the table. Results are exact, but latency grows linearly with row count. Calling create_index builds an approximate IVF_PQ index that trades that exactness for sub-linear search.
How does LanceDB versioning work, and how do you read an older table version?
basics
~20 sEvery write to a LanceDB table — append, update, delete, overwrite, index build — commits a new immutable version instead of mutating data in place. table.list_versions() enumerates them, table.checkout(n) pins the table object to version n for reading, and table.checkout_latest() returns to the newest.