In Qdrant, what does create_payload_index buy you and what breaks without it?
answer
- separate from the vector index
- typed per field
- the planner needs a number
- correctness fine, strategy blind
- costs RAM and write throughput
basics
~20 sA payload index is a per-field, typed structure separate from the vector index. It makes filters selective and lets the query planner estimate how many points a filter matches. Without it results are still correct, but filtering degrades to scanning candidates and the planner picks blindly.
solid answer
~50 s`create_payload_index(collection_name, field_name, field_schema=...)` builds a secondary index over one payload field, entirely separate from the HNSW vector index. You declare the type — `PayloadSchemaType.KEYWORD`, `INTEGER`, `FLOAT`, `BOOL`, `GEO`, `DATETIME`, `UUID`, or a full `TextIndexParams` for tokenised text — and Qdrant builds the matching structure over existing and future points. Two things follow. First, resolving a filter becomes a lookup rather than a payload check per candidate. Second, and more important, the planner gains a **cardinality estimate**: it can tell whether a filter matches ten points or ten million, which is what lets it choose between an exact scan of the matching subset and a filtered graph traversal. Without an index the filter still returns correct results, but slowly and with the planner unable to make that choice. Indexes cost memory and write throughput, so index the fields you actually filter on, not every key.
code
python · 15 linesfrom qdrant_client import QdrantClient, models
client = QdrantClient(url="http://localhost:6333")
client.create_payload_index(
collection_name="articles",
field_name="lang",
field_schema=models.PayloadSchemaType.KEYWORD,
)
client.create_payload_index(
collection_name="articles",
field_name="published_at",
field_schema=models.PayloadSchemaType.DATETIME,
)go deeper
Know that filterable payload fields need their own index created with create_payload_index, and that the field_schema declares the field's type.
Explain that the payload index is separate from the vector index and that its real value is giving the query planner a cardinality estimate, not just faster matching.
Diagnose a slow filtered query by comparing get_collection's payload_schema against the fields the query filters on, and weigh index memory and write cost against query benefit.
Own index provisioning as part of collection creation so environments cannot drift, and set the policy for which payload fields earn an index as the schema grows.
## Two independent indexes A Qdrant collection has a vector index (HNSW over the embeddings) and, optionally, payload indexes over individual payload fields. They are built, configured and stored separately. Creating one does not touch the other, and dropping a payload index does not affect vector search. ## Creating one ``` client.create_payload_index(collection_name="articles", field_name="lang", field_schema=models.PayloadSchemaType.KEYWORD) ``` The call is asynchronous with respect to the data already in the collection: Qdrant backfills existing points and then keeps the index up to date on every upsert. `delete_payload_index` removes it. `get_collection` reports the configured payload schema, which is how you check in production whether the index you *think* exists actually does. ## Choosing the schema type The type is not cosmetic — it determines which conditions can be answered efficiently: - **keyword** — exact string matching, for `MatchValue` / `MatchAny` / `MatchExcept`. The bread-and-butter index for categorical fields. - **integer** / **float** — `Range` conditions and, for integers, exact lookups. `IntegerIndexParams` exposes `lookup` and `range` flags so you can build only the half you need. - **bool**, **uuid**, **datetime** — the specialised equivalents; `datetime` backs `DatetimeRange`. - **geo** — required for `GeoRadius`, `GeoBoundingBox` and `GeoPolygon` conditions. - **text** — built via `TextIndexParams` with a tokenizer (`TokenizerType.WORD`, `PREFIX`, `WHITESPACE`, `MULTILINGUAL`), token length bounds and `lowercase`. This is what makes `MatchText` behave as a token-level full-text match rather than a naive scan. Several of the params objects carry extra flags: `on_disk=True` keeps the index off the heap at the cost of I/O, `is_tenant=True` marks a keyword field as the tenant partition key, and `is_principal` marks the field data is predominantly ordered by. ## Nested and dotted paths Indexes follow payload paths. `author.country` indexes a nested object field; `items[].price` indexes a field inside an array of objects. If you filter on a nested path, index that exact path — indexing the parent does nothing for it. ## What actually breaks without an index Nothing breaks in the correctness sense. Qdrant will honour a filter on an unindexed field and return exactly the right points. What you lose: 1. **Speed.** Evaluating the condition means reading payload for candidates instead of consulting a compact structure. On a large collection this dominates query time. 2. **Cardinality estimation.** This is the subtle and more damaging loss. Qdrant's filtered-search planner decides between exhaustively scanning the matching subset and traversing the graph with the filter applied, and it makes that call from an estimate of how many points the filter matches. Without an index it has no estimate, so it cannot reliably pick the cheap strategy — a highly selective filter that should have been answered by scanning a few hundred points instead runs as a general search. 3. **Full-text semantics.** `MatchText` against a field with no text index does not give you tokenised matching with your chosen tokenizer. ## The costs Each index consumes memory proportional to the field's cardinality and adds work to every upsert, since the index has to be maintained. High-cardinality fields — a unique document id, a free-text title indexed as keyword — are the expensive case and often the least useful, because a filter that matches one point is better served by fetching by id. Index the fields that appear in filters at query time, check the payload schema in `get_collection` against the filters your code actually issues, and drop the ones nobody queries. ## Operational habits Create payload indexes as part of collection provisioning, in the same code path that creates the collection, so environments cannot drift. When a filtered query is slow in production and fast in staging, the first thing to compare is the payload schema — a missing index on one environment is the most common cause. And when you add a new filterable field to your payload, adding the index is part of the same change, not a follow-up ticket.
- Does creating a payload index on a populated collection require re-uploading the points?No. Qdrant backfills the index over the points already stored and then maintains it on subsequent upserts. The build runs in the background, so a large collection will show improving filtered-query performance rather than an instant switch, and `get_collection` reports the field once it is part of the payload schema.
- Which payload fields are a bad candidate for an index?Very high-cardinality fields you rarely filter on — free text stored as keyword, per-document unique ids, timestamps at nanosecond precision. They cost memory proportional to cardinality and slow every upsert, while a filter matching one or two points is usually better served by a direct id lookup or a `HasIdCondition`.
- How would you check in production whether a filter's field is indexed?Call `get_collection(name)` and inspect `payload_schema`, which lists each indexed field with its type. Comparing that map against the fields your query code actually filters on catches the most common performance regression — an index that exists in one environment and not another, or a new payload field whose index was never added.
saying these in an interview costs you the question
- Believing an unindexed filter returns wrong or partial results
- Indexing every payload key by default
- Indexing the parent object instead of the nested path
- Expecting MatchText to tokenise without a text index
- Thinking the payload index is part of the HNSW structure