Your Gemini-embedded RAG index needs a new embedding model. What breaks?
answer
- Vectors are not portable between models
- No conversion function exists
- Dimension match does not imply comparability
- Backfill into a second index
- Cut over only after offline recall wins
basics
~20 sVectors from two different embedding models are not comparable, so the entire corpus must be re-embedded — there is no conversion. Plan a backfill into a second index and a cutover, not an in-place, gradual replacement.
solid answer
~50 sEach embedding model defines its own vector space, so a query vector from the new model has no meaningful geometric relationship to stored vectors from the old one, even when the dimension counts happen to match. There is no projection or conversion function; the only path is re-embedding the corpus. The safe operational shape is: stand up a second index with the new model's dimension count, backfill it while the old index continues serving, run an offline recall comparison on a fixed query set to confirm the new model is actually better on your data, then cut reads over and keep the old index for rollback. Because Gemini's embedding models are also parameterised by task type and output dimensionality, treat all three — model name, task type, dimension — as the index's identity and store them as metadata so a mismatch is detectable. Budget for the re-embed cost: it is one full pass over the corpus at input-token pricing.
go deeper
Know the headline rule: embeddings from different models cannot be compared, so switching models means re-embedding everything rather than converting what you already stored.
Explain why — each model defines its own vector space — and note that Gemini's task type and output dimensionality change the vector too, so all three settings define the index.
Walk the operational plan: second index, resumable backfill, dual-write, offline recall comparison, flagged cutover, old index retained for rollback, and the failure modes such as cached query vectors and partial backfills.
Own the economics and the policy: estimate the one-off re-embed spend against measured retrieval gains or a deprecation deadline, mandate index metadata and a standing evaluation set, and require one variable per migration so results stay interpretable.
## Why old and new vectors cannot mix An embedding model is trained to arrange text in a space of its own invention. Dimension 417 means whatever that training run made it mean; a different model, even from the same vendor and the same family, assigns entirely different semantics to its coordinates. Cosine similarity between a vector from model A and a vector from model B is a number, and it is meaningless. This surprises people because the artefacts look interchangeable — both are float arrays, and if both are 3072 long the database will happily accept them. Nothing errors. Search simply returns approximately random neighbours for any query embedded with the new model, or subtly wrong ones if the two models are related. There is no linear map you can apply to convert; learning one is a research exercise, not an ops step. The practical statement: **the model name is part of the index's schema.** So are the task type and the output dimensionality, since on Gemini's embedding API both change the vector for the same input text. Any one of the three changing invalidates the stored vectors. ## The migration shape 1. **Create a new index/collection** with the new model's dimension count. Do not reuse the old one; you want both to exist simultaneously. 2. **Backfill** by re-embedding every chunk with the new model and the same task type you intend to use at write time (`RETRIEVAL_DOCUMENT` for stored passages). This is a batch job: chunk-limited, rate-limited, resumable, idempotent by chunk ID. 3. **Dual-write** new and updated documents to both indexes for the duration, so the new index is not stale by the time the backfill ends. 4. **Evaluate offline** before switching any traffic: a fixed query set with known-relevant chunks, recall@k and MRR measured on both indexes. "Newer model" is a hypothesis, not a result — a newer general-purpose model can lose to an older one on a narrow domain, and this is where you find out. 5. **Cut reads over** behind a flag, ideally as a percentage ramp with retrieval-quality telemetry. 6. **Keep the old index** until you are confident, then delete it. Rollback is a flag flip only while it exists. ## Cost and time Re-embedding is a full pass over the corpus priced on input tokens, so the bill scales with total corpus size and is a one-off you can estimate precisely in advance. The wall-clock time is usually governed by your rate limit rather than compute: batch multiple texts per request, run bounded concurrency, and back off on rate-limit responses rather than hammering. Make the job resumable — a multi-hour backfill will be interrupted, and restarting from zero twice is how migrations slip a week. ## Things that quietly go wrong - **Dimension change smuggled in.** If the new model's default length differs from the old one, or you take the opportunity to set `output_dimensionality`, the new collection's schema must match, and every downstream consumer that hardcoded a vector width breaks. - **Task type drift.** The query path and the ingest path are often in different services. If only one is updated, you get a silent recall regression that looks like "the new model is worse". - **Partial backfill served as complete.** Cutting over at 90% means the missing 10% is invisible to search — no error, just documents that can never be found. Gate the cutover on a counted completion check, not on the job exiting. - **Chunking changed at the same time.** Changing chunk boundaries and the model in one step makes the evaluation uninterpretable. Change one thing per migration. - **Cached query vectors.** Any cache keyed only on query text now serves old-space vectors against a new-space index. Include the model, task type and dimension in the cache key. ## When it is worth doing Re-embedding costs real money and a real week. Justify it with numbers: a measured recall improvement on your evaluation set, a dimension reduction that solves a quantified RAM problem, or a deprecation deadline that forces your hand — model lifecycles are the common trigger, since providers retire older embedding models on a published schedule. "There is a newer model" is not a reason on its own. ## The lasting lesson Build the first index as though you will replace it: store model, task type and dimension as metadata on the collection; keep an evaluation set from day one; make ingestion re-runnable by chunk ID. Teams that do this migrate in days. Teams that do not discover that nobody remembers what the production vectors were generated with.
- Both models output 3072 dimensions. Can you keep the old vectors and only embed new documents with the new model?No. Matching dimension counts make the vectors storable side by side but not comparable — each model's coordinates mean different things. The index would contain two disjoint clouds, and a query embedded with either model would only ever retrieve well from its own half. The database will not complain, which is exactly why this bug survives into production.
- How do you keep search working during a multi-hour backfill?Keep the old index serving reads for the whole backfill, dual-write new and updated documents into both indexes so the new one does not go stale, and switch reads only after a counted completeness check confirms every chunk landed. Keep the old index alive afterwards so rollback is a flag flip rather than another backfill.
- How do you prove the new embedding model is actually better before cutting over?Offline evaluation on a fixed query set with known-relevant chunks: measure recall@k and MRR against both indexes and compare. Newer does not mean better on a narrow domain. If the numbers are within noise, the migration is only justified by a deprecation deadline or an infrastructure win such as smaller vectors.
- What else besides the model name invalidates stored Gemini embeddings?Changing the task type used at ingest, and changing output dimensionality. Both alter the vector produced for identical text, so either one silently splits an index into incomparable halves. Record all three — model, task type, dimension — as collection metadata and assert on them at write time.
Two embedding models are two cartographers who each drew the world on their own private grid: both maps have coordinates, but reading one map's coordinates on the other's grid puts you in the ocean.
saying these in an interview costs you the question
- Applying a linear transform to convert old vectors
- Believing equal dimensions make vectors comparable
- Re-embedding only new documents after a model switch
- Cutting over before the backfill is verifiably complete
- Changing chunking and the embedding model in one step