skip to content

How do you migrate a production vector index to a new OpenAI embedding model?

level: principalimportance: should knowfreq 36%

answer

  1. Different models, incomparable spaces
  2. No conversion, no partial migration
  3. Build alongside, never mutate in place
  4. Dual-write during the backfill window
  5. Prove recall improved before cutting over

basics

~20 s

Vectors from different embedding models are not comparable, so the whole corpus must be re-embedded. Build the new index alongside the old one, keep queries on the old index until the new one is complete and evaluated, then cut over and only afterwards delete the old vectors.

solid answer

~50 s

The constraint that drives everything: an embedding is only meaningful relative to the model and dimension width that produced it. There is no conversion between spaces, so switching from `text-embedding-ada-002` to `text-embedding-3-large`, or changing the `dimensions` value, means re-embedding every chunk. A partial migration produces an index where some vectors are simply unreachable by any query. The safe pattern is dual-index. Provision a second index, backfill it through the Batch API so bulk work does not starve live query traffic, and keep serving from the old one throughout. Before cutover, run a labelled evaluation set against both and compare recall@k — a newer model is not automatically better on your corpus, and this is the only step that proves it. Cut over behind a flag with the old index still warm so rollback is instant, then delete. What makes this cheap is having stored the model name and width with every vector from day one, so the mixed state is detectable rather than silent.

go deeper

for a junior

Understand the one rule that matters: vectors from two different embedding models cannot be compared, so changing the model means embedding the whole corpus again.

for a middle

Explain why a partial migration fails silently rather than loudly, and that dimension width is as breaking a change as the model itself.

for a senior

Walk the dual-index playbook — parallel index, offline backfill, dual-write during the window, shadow reads, flag-controlled cutover, delayed deletion — and name the evaluation metrics you gate on.

for a principal

Own the economics and the prevention: quantify re-embed tokens, double storage and days of wall-clock before proposing it, protect live rate-limit budget from the backfill, and mandate model-and-width metadata so the next migration is routine.

## Why this is a migration and not a config change An embedding model defines a coordinate system. Two models trained separately place "quarterly revenue" at unrelated points, and there is no published transform between OpenAI's embedding spaces — no projection matrix, no adapter. The same applies to a change of the `dimensions` parameter within one model: a 1536-wide vector and a 3072-wide vector cannot even be compared, since cosine similarity is undefined across differing lengths. So the moment you change the model or the width, every stored vector becomes dead weight. This is what makes the embedding model choice one of the stickier decisions in a retrieval system: it is cheap to make and expensive to revisit. ## The failure mode you are avoiding The dangerous version of this migration is the incremental one. A team switches the model in the ingestion code, deploys, and new documents start arriving as `text-embedding-3-large` vectors in an index full of `ada-002` vectors. Nothing errors — the dimensions may even match if widths coincide. Queries embedded with the new model retrieve only the new documents and score the old ones near-randomly; queries embedded with the old model do the reverse. Recall degrades quietly, users report "the search got worse", and because there is no exception anywhere the cause takes days to find. The structural defence is metadata: every vector row carries the model identifier and the dimension width, ingestion asserts they match the index's declared configuration, and a mismatch is a hard failure at write time rather than a soft failure at query time. ## The dual-index playbook **1. Decide with evidence, not novelty.** Before committing, embed a labelled evaluation set — queries paired with the chunks that should be retrieved — under both configurations and compare recall@k and MRR. Newer and larger is not automatically better on a specific corpus, and if the gain is inside the noise, the correct decision is not to migrate. **2. Provision a parallel index.** Separate namespace, collection or table, configured for the new width and metric. Do not mutate the live one. **3. Backfill offline.** Submit the corpus through the Batch API rather than the synchronous endpoint, so a multi-million-chunk job runs against separate queues and cannot consume the rate-limit budget your live query path depends on. Content-hash keys make the job resumable and idempotent. **4. Dual-write new documents.** From the moment the backfill starts, ingest writes to both indexes. Otherwise documents created during the backfill window exist in the old index only, and you cut over to a stale corpus. **5. Shadow-read before cutting over.** Route a sample of real production queries to both indexes, log the top-k from each, and compare — offline evaluation sets are small and synthetic, real query distributions are neither. This is also where you catch latency regressions from a wider vector. **6. Cut over behind a flag,** with the old index still populated and warm. A flag flip is an instant rollback; a deleted index is not. **7. Delete last.** Keep the old index for a defined soak period — long enough to be confident, short enough that you are not paying for two copies forever. ## Costs to put on the table before starting - **Re-embedding tokens.** Knowable exactly in advance: tokenise the corpus once and multiply by the model's rate. There is no output-token component. - **Double storage** for the duration. At 3072 dimensions and float32 a vector is 12 KB; ten million chunks held twice is real money in a memory-resident index. - **Wall-clock time.** The Batch API's completion window is up to 24 hours per job, and a large corpus may take several. Plan the migration in days, not hours. - **Ongoing serving cost** if the new width is larger: more index memory and more work per query, permanently. ## The judgment an interviewer is probing Three things. First, whether you understand that vector spaces are model-specific and that no conversion exists — the technical core. Second, whether you reach for a dual-index, shadow-read, flag-controlled cutover rather than an in-place mutation, which is generic migration discipline applied to a new substrate. Third, whether you insist on measuring before and after: the entire premise of the migration is that retrieval improves, and a migration that ships without a recall comparison has not demonstrated its own reason for existing. The strongest answers also mention the preventative half — recording model and width alongside every vector, and asserting them on write — because that is what turns this from an archaeology project into a scheduled job.

  • Could you keep old vectors and only embed new documents with the new model?
    No. The two sets occupy unrelated coordinate systems, so a query embedded with either model retrieves sensibly from only one half of the index and scores the other half near-randomly. Nothing errors, which is what makes it dangerous: recall degrades silently. The index must be homogeneous in model and width.
  • Does changing only the dimensions parameter, keeping the same model, avoid a full re-embed?
    No. Cosine similarity is undefined between vectors of different lengths, so a width change is as breaking as a model change. If you shortened client-side you could in principle re-truncate stored vectors and re-normalise, but going wider is impossible without new API calls, so treat any width change as a full migration.
  • How do you validate the new index before cutting over?
    Two layers. Offline, run a labelled evaluation set of queries and expected chunks against both indexes and compare recall@k and MRR. Online, shadow-read a sample of real production queries against both and diff the top-k, which surfaces distribution effects and latency regressions that a small curated set will not.
  • What single practice makes this migration dramatically cheaper next time?
    Storing the model identifier and dimension width as metadata on every vector, and asserting them at write time so a mismatch fails loudly. That turns a silent, hard-to-diagnose mixed index into an impossible state, makes the corpus self-describing, and lets tooling report exactly what needs re-embedding when the next model ships.

saying these in an interview costs you the question

  • Thinks old vectors can be converted or projected into the new space
  • Switches the model in ingestion and lets the index become mixed
  • Skips a recall comparison because the new model is newer
  • Deletes the old index at cutover, leaving no rollback
  • Runs the full backfill on the live endpoint during peak traffic

context