skip to content

How do Qdrant aliases let you swap in a rebuilt collection with no downtime?

level: seniorimportance: should knowfreq 38%

answer

  1. a name that points at a collection
  2. use it wherever a collection name goes
  3. one call, several operations
  4. atomic switch, no missing-name window
  5. snapshot copies, alias reveals

basics

~20 s

An alias is a name that resolves to a collection and can be used wherever a collection name is accepted. Applying the delete and create operations in one update_collection_aliases call repoints it atomically, so readers move to the rebuilt collection with no gap and the old one stays as rollback.

solid answer

~50 s

Point your application at an alias rather than a physical collection name. To roll out a re-embedded corpus you build `docs_v2` alongside the live `docs_v1`, load and verify it, then send a single `client.update_collection_aliases(change_aliases_operations=[...])` containing a `DeleteAliasOperation` for the alias on `docs_v1` and a `CreateAliasOperation` binding it to `docs_v2`. The operations in one call are applied atomically, so there is no instant where the alias is missing or ambiguous — queries in flight hit one collection or the other, never an error. Rollback is the same call with the arguments swapped, which is what makes this safe enough to do during business hours. Snapshots are the complementary tool: `client.create_snapshot(collection_name)` writes a point-in-time copy, and `client.recover_snapshot(collection_name, location=...)` restores it — typically into a new collection that you then alias into place. Aliases give you the atomic cutover; snapshots give you the copy to cut over to.

code

python · 19 lines
python
from qdrant_client import QdrantClient, models

client = QdrantClient(url="http://localhost:6333")

# after docs_v2 has been built and verified
client.update_collection_aliases(
    change_aliases_operations=[
        models.DeleteAliasOperation(
            delete_alias=models.DeleteAlias(alias_name="docs")
        ),
        models.CreateAliasOperation(
            create_alias=models.CreateAlias(
                collection_name="docs_v2", alias_name="docs"
            )
        ),
    ]
)

print(client.get_aliases())

go deeper

for a junior

Know that an alias is a second name for a collection and can be used in place of the collection name, so applications need not hard-code the physical name.

for a middle

Explain that update_collection_aliases applies its list of operations atomically, so delete-plus-create in one call repoints the name with no failing window.

for a senior

Walk the whole rollout: build, load, verify recall, shadow-read, switch, keep the old collection for rollback — and raise the in-flight-writes problem before being asked.

for a principal

Own the convention that nothing addresses a physical collection name directly, so model upgrades and restores are routine operations rather than coordinated deploys, and set the retention policy for superseded collections and snapshots.

## The problem aliases solve A Qdrant collection's vector size and distance are fixed at creation, so a new embedding model, a changed chunking strategy or a different metric all mean building a **new collection**. If your application hard-codes the collection name, the switch is a code deploy timed against a data load — and a rollback is another deploy. Aliases decouple the two. An alias is simply an alternative name that resolves to a collection. Anywhere the API takes a collection name — search, upsert, scroll, retrieve — it will accept an alias. Your service reads and writes `docs`; `docs` currently means `docs_v3`. ## The atomic switch Alias changes go through one call that takes a **list** of operations: ``` client.update_collection_aliases( change_aliases_operations=[ models.DeleteAliasOperation(delete_alias=models.DeleteAlias(alias_name="docs")), models.CreateAliasOperation( create_alias=models.CreateAlias(collection_name="docs_v2", alias_name="docs") ), ] ) ``` The operations in a single call are applied together. That is the whole point: doing the delete and the create as two separate calls leaves a window — however short — in which the alias does not exist and every query fails. Bundling them means a request either resolves to the old collection or to the new one. There is no third outcome to handle. The available operations are `CreateAliasOperation`, `DeleteAliasOperation` and `RenameAliasOperation`. You can inspect the current state with `client.get_aliases()` for all aliases or `client.get_collection_aliases(collection_name)` for one collection's, which is the first thing to check when someone asks which physical collection is live. ## The rollout sequence 1. Create `docs_v2` with the new vector configuration. 2. Load it — bulk upsert with deterministic ids, as a background job with no traffic on it. 3. Verify: exact point count against the source, collection status settled, and a recall check against a held-out query set. This is the step people skip and regret, because a new embedding model can produce a collection that looks fine on counters and is worse on relevance. 4. Optionally shadow-read: issue a sample of live queries against `docs_v2` explicitly and compare results with `docs_v1`. 5. Switch the alias in one call. 6. Keep `docs_v1` for a defined period. It costs storage; it buys an instant rollback. 7. Delete `docs_v1`. A collection can carry more than one alias, which lets you keep a stable `docs` for readers plus a pinned `docs_v2` name for the tools that need to address a specific generation. ## The writer problem The alias switch is atomic for readers; live writers need a plan of their own. If your ingest pipeline is writing continuously, points that land in `docs_v1` between your last sync and the cutover are lost from `docs_v2`. The usual answers are to dual-write to both collections during the migration window, to pause ingest briefly around the switch, or to run a catch-up pass over everything modified since the load started before flipping. Raising this unprompted is what distinguishes a senior answer — otherwise the alias story sounds cleaner than it is. ## Snapshots A snapshot is a point-in-time copy of a collection's data, produced on the node and written to that node's snapshot storage. - `client.create_snapshot(collection_name)` creates one. - `client.list_snapshots(collection_name)` enumerates them. - `client.recover_snapshot(collection_name, location=...)` restores from a snapshot, where the location is a file URI on the node or a URL it can fetch. - `client.delete_snapshot(...)` removes one. - `client.create_full_snapshot()` captures the whole storage rather than a single collection. The critical property is *point in time*: writes accepted after the snapshot are not in it. That is a feature for backup and a hazard for migration — recover a week-old snapshot and you are missing a week of ingest unless you replay it. Snapshots serve three jobs. **Backup and restore** is the obvious one. **Environment cloning** — snapshot production, recover into staging — gives you a realistic dataset without re-running embeddings, which is often the expensive part. **Migration** moves a collection between clusters or versions: snapshot, transfer, recover into a new collection name, verify, alias into place. In a distributed deployment, snapshots are taken per node and per shard, so a whole-collection restore means dealing with each shard's snapshot rather than one file; do not describe it as a single portable artifact for a sharded collection. ## Putting the two together The pattern to state in an interview: **snapshots create the copy, aliases decide who sees it.** A restore never overwrites live data in place — you recover into a new collection, check it, and then repoint the alias in one atomic call. The old collection stays until you are sure. Applied consistently, this makes both model upgrades and disaster recovery boring, which is the goal.

  • Why not just delete the old alias, then create the new one in a second call?
    Because between the two calls the alias name does not resolve to anything, and every query in that window fails. Both operations belong in one `update_collection_aliases` call precisely so they apply together and a request sees either the old collection or the new one.
  • Ingest is running continuously while you build the replacement collection. What breaks?
    Writes that land in the old collection after your bulk load started are missing from the new one, so the cutover silently loses recent data. Handle it by dual-writing to both collections during the migration window, pausing ingest around the switch, or running a catch-up pass over everything changed since the load began.
  • You restore a week-old snapshot after an incident. What should you expect?
    Exactly the collection as it stood when the snapshot was taken — a week of subsequent writes are not in it. Recover into a new collection rather than over the live one, replay the missing ingest from your source of truth, verify counts, and only then move the alias.

saying these in an interview costs you the question

  • Performs the alias delete and create as two separate calls
  • Thinks an alias can only be used for search, not for writes
  • Deletes the old collection immediately after switching the alias
  • Assumes a snapshot includes writes that arrived after it was taken
  • Restores a snapshot over the live collection instead of into a new one

context