In Pinecone, how do you upsert and query a record holding both dense and sparse vectors?
answer
- two vectors, one record, one id
- the sparse half is pairs, not a full array
- only the non-zero entries travel
- query sends both halves at once
- values plus sparse_values, indices and values
basics
~20 sA hybrid record carries the dense embedding in values and a sparse vector in sparse_values as {"indices": [...], "values": [...]}. The query passes vector= and sparse_vector= together, and Pinecone returns one ranked list with a single combined score per match.
solid answer
~40 sBoth vectors live on the same record in the same index — you do not maintain two indexes. On upsert, each record has an `id`, the dense embedding in `values` (whose length must equal the index `dimension`), optional `metadata`, and `sparse_values` shaped as `{"indices": [int, ...], "values": [float, ...]}` with the two lists the same length and the indices unique. Indices are opaque integer term ids from whatever sparse encoder you use; Pinecone never interprets them. At query time you send both halves — `index.query(vector=dense_q, sparse_vector=sparse_q, top_k=10, include_metadata=True)` — and the engine scores both parts of each candidate and returns one list ranked by a single `score` field. `top_k` applies to that combined ranking, so there is no client-side merging of two result sets and no second round trip.
code
python · 23 linesfrom pinecone import Pinecone
index = Pinecone(api_key="YOUR_KEY").Index("docs-hybrid")
index.upsert(
vectors=[
{
"id": "chunk-42",
"values": [0.02] * 768,
"sparse_values": {"indices": [7, 913, 20514], "values": [0.71, 0.44, 0.28]},
"metadata": {"doc": "handbook", "page": 12},
}
]
)
res = index.query(
vector=[0.02] * 768,
sparse_vector={"indices": [7, 20514], "values": [0.9, 0.5]},
top_k=5,
include_metadata=True,
)
for match in res["matches"]:
print(match["id"], match["score"])go deeper
Recall the two fields: values for the dense embedding and sparse_values with its indices and values lists, both on the same record.
Explain that the query carries both halves, that Pinecone returns one score and one ranked list, and that the sparse side stores only non-zero pairs with no declared dimension.
Show you know upsert replaces the whole record, so a partial re-write silently strips the sparse half, and describe how you would detect that in a running system.
Weigh the single-index blend against running separate dense and sparse retrieval: one round trip and one score versus independent tuning and independent scaling of each side.
## One record, two vectors The defining property of Pinecone's hybrid model is that the lexical and semantic representations of a chunk are two fields on the *same* record, in the *same* index, sharing the same `id`. There is no separate keyword index to keep in sync, and no client-side fusion of two ranked lists — the engine scores both halves of each candidate and hands you one list. ## The record shape A hybrid record is: - `id` — your string key, unchanged from a dense-only setup. - `values` — the dense embedding, a list of floats whose length must equal the index's `dimension`. This is mandatory in a dense index; you cannot upsert a record that has only `sparse_values`. - `sparse_values` — `{"indices": [...], "values": [...]}`. `indices` is a list of non-negative integers, `values` a list of floats of exactly the same length. The integers must be unique within a record; each pair means "term with this id has this weight in this document". - `metadata` — an optional flat dict, exactly as in a dense-only index, and it still drives filtering independently of scoring. The sparse side has no declared dimension. `create_index(dimension=...)` describes only the dense vector; the sparse index space is implicitly as wide as your encoder's vocabulary, and Pinecone stores only the non-zero entries you send. A 30,000-term vocabulary with 40 non-zero terms in a chunk costs 40 pairs, not 30,000 floats. That sparsity is the entire point of the representation. ## The query shape A hybrid query looks like a dense query with one extra argument: ``` index.query(vector=dense_q, sparse_vector={"indices": [...], "values": [...]}, top_k=10, include_metadata=True) ``` Pinecone computes the dot product of the dense halves and the dot product of the sparse halves and returns their sum as the match's single `score`. `top_k` cuts the combined ranking, not each side separately, so you never see "top 10 dense plus top 10 sparse" — you see the 10 best hybrid scores. Metadata filters apply the same way they do for a dense query, narrowing the candidate set before the ranking is returned. Note that both halves are optional *on the query*. Sending only `vector=` gives pure semantic search over the same index; sending a sparse vector whose values are all zeroed removes the lexical contribution. This is what makes A/B testing hybrid against dense-only cheap: same index, same records, different query payload. ## Consequences worth naming **Re-upsert semantics.** `upsert` replaces a record wholesale by id. If you upsert a record supplying only `values` and omit `sparse_values`, you drop the sparse half of that record. That is a real production bug: a re-embedding job that rewrites vectors without regenerating sparse values silently strips lexical matching from every record it touches, with no error and no obvious symptom other than gradually worse keyword recall. **Storage and cost.** Sparse values add to the record size and therefore to index storage, roughly in proportion to the average number of non-zero terms per chunk. Aggressive encoders (learned sparse models that expand a chunk to hundreds of weighted terms) cost meaningfully more than a plain BM25 encoding over the chunk's own tokens. **Sanity-checking the wiring.** `describe_index_stats()` tells you the vector count, which confirms records landed, but it will not tell you whether the sparse half is populated. Fetch a known record by id and look at whether `sparse_values` is present — that is the cheap check when hybrid results look suspiciously like dense-only results. **Alternative shape.** Pinecone also supports sparse-only indexes created with `vector_type="sparse"`. There the record carries just the sparse vector, you query a dense index and a sparse index separately, and you merge the two lists yourself. The single-index sparse-dense shape trades that flexibility for one round trip and one already-blended ranking. ## A checklist for the interview answer Same id, same index, two fields. Dense length equals the index dimension; sparse is index/value pairs with unique integer ids and no declared width. Query sends both; `top_k` applies to the blended ranking. Upsert replaces, so always write both halves together.
- What happens if a re-embedding job upserts records with only the dense values?Upsert replaces the record by id, so the omitted `sparse_values` are dropped and those records become dense-only. Nothing errors, and the index still returns results — lexical recall just quietly decays for the affected slice. The defence is to treat dense and sparse as one write unit in the ingestion pipeline, and to spot-check a fetched record for a populated sparse half after any bulk rewrite.
- Does the index `dimension` constrain the sparse vector at all?No. `dimension` describes only the dense embedding, and the dense list you upsert must match it exactly. The sparse side has no declared width: it is a set of integer term ids with weights, and Pinecone stores only the non-zero pairs you send. Its effective size is your encoder's vocabulary, which Pinecone never sees or validates.
- How do you A/B test hybrid against pure dense retrieval without building a second index?Query the same hybrid index twice: once with both `vector=` and `sparse_vector=`, once with `vector=` alone. The records are identical, so the only variable is the query payload, and you can route a fraction of live traffic to each arm and compare click-through or labelled relevance. That is a direct benefit of both vectors living on one record.
saying these in an interview costs you the question
- Thinks hybrid needs two indexes that must be kept in sync
- Sends the sparse vector as a dense array of mostly zeros
- Believes the index dimension applies to the sparse vector
- Assumes top_k is applied separately to dense and sparse results
- Upserts dense values alone and expects sparse values to persist