In Weaviate's hybrid(), what does query_properties=['title^2','body'] change?
answer
- keyword side only
- the caret is a field weight
- a list, not an addition
- stored vectors ignore it
- tune it at alpha=0 first
basics
~20 squery_properties configures only the keyword half of a Weaviate hybrid query: it restricts BM25 matching to the listed properties and the ^ suffix boosts a property's weight. The vector half is unaffected — it always uses the object's stored vector.
solid answer
~40 s`query_properties` scopes and weights the **BM25 side** of `hybrid()`. Listing `["title^2", "body"]` means keyword matching happens against only those two properties, with a term found in `title` counting roughly twice as much as the same term in `body`. Two consequences follow. First, any property you leave out stops contributing keyword matches entirely — a common cause of "the exact phrase is in the document but it never ranks". Second, the boost does **not** reach the vector half: the vector score comes from the object's stored embedding, so no `^` value can make semantic search prefer titles. If you want title text to dominate semantically you have to change what gets vectorized, not the query. Omitting `query_properties` lets the keyword half search the collection's searchable text properties.
code
python · 18 linesimport weaviate
from weaviate.classes.query import MetadataQuery
client = weaviate.connect_to_local()
articles = client.collections.get("Article")
res = articles.query.hybrid(
query="connection pool exhaustion",
alpha=0.4,
query_properties=["title^2", "body"],
limit=5,
return_metadata=MetadataQuery(score=True, explain_score=True),
)
for obj in res.objects:
print(obj.properties["title"], obj.metadata.score)
print(obj.metadata.explain_score)
client.close()go deeper
Know that the list names which properties the keyword part of a hybrid query searches and that ^2 makes matches in that property count more. Be able to read the syntax back correctly.
Explain that the list replaces the default property set, so omitting a property silently removes its keyword matches, and that the vector half is untouched by anything in this argument.
Show the diagnostic instinct: when a verbatim match will not rank, check property coverage and tokenization before touching boost values, and validate boosts at alpha=0 before trusting them under production fusion settings.
Frame it as the query-time versus ingest-time split — lexical weighting is tunable per request, semantic emphasis requires changing what is vectorized and re-embedding, which is a migration with real cost.
## What the argument controls `hybrid()` accepts `query_properties` as a list of property names, each optionally suffixed with `^` and a number. It answers two questions about the **keyword half** of the hybrid query: 1. **Which properties are searched** for the query's terms. 2. **How much a match in each property is worth** relative to the others. `["title^2", "body"]` means: search `title` and `body` only, and treat a term occurrence in `title` as worth about twice a term occurrence in `body`. A property with no suffix carries the default weight of 1. This is the field-weighting lever you reach for when your objects have a short high-signal property and a long low-signal one. A query term appearing in a 6-word title is far stronger evidence of aboutness than the same term buried in a 4,000-word body, and without a boost, BM25's own length normalization only partly compensates. ## The exclusion effect is the bigger gotcha The list is not additive on top of a default — it is a **replacement**. If you pass `["title^2", "body"]` and your collection also has an `abstract` property, keyword matches in `abstract` now contribute nothing. The document may still surface via the vector half, and it may still rank respectably, which is what makes this hard to spot: nothing errors, results still come back, and only the ordering is quietly wrong. The classic symptom is a user searching a phrase they can see verbatim in a document that will not come to the top. When you omit `query_properties`, the keyword half searches the collection's searchable text properties. That default is usually what you want early on; you narrow it when you have learned which properties carry signal. ## What it does not touch The vector half is completely outside `query_properties`' reach. A hybrid query has one query embedding compared against each object's stored vector. That stored vector was produced at ingest time, from whatever text the vectorizer was configured to include. So: - Boosting `title^5` cannot make the semantic side care more about titles. - Excluding a property from `query_properties` does not remove its text from the object's embedding. - If a property is excluded from vectorization at the schema level, no query-time argument brings it back into the vector score. The practical rule: `query_properties` is a **query-time** lever on lexical matching; what the vector represents is an **ingest-time** decision. Confusing the two leads to hours of fruitless boost tuning against a semantic ranking that was never listening. ## Boost values are relative, not absolute `^2` is a multiplier within the keyword scoring, not a guarantee of position. Its effect competes with everything else BM25 weighs — term rarity, term frequency, field length — and then the whole keyword contribution is scaled again by `1 - alpha` during fusion. So a boost that looks decisive in a pure keyword query can be almost invisible in a hybrid query with a high alpha. Tune boosts at `alpha=0` first, where you can see the keyword ranking on its own, then re-check at your production alpha. Also: boosts interact with fusion strategy. Under a fusion mode that reads only rank positions, a boost that improves a document's keyword *score* without moving its keyword *rank* has no effect at all on the fused order. ## A property that never matches If a listed property produces no keyword matches no matter what you search, the usual causes are: - The property is not searchable — a text property configured without a searchable index cannot participate in BM25 at all, and listing it in `query_properties` will not change that. - The property is not a text type. Numbers, booleans and dates are matched by filters, not by keyword scoring. - Tokenization mismatch — the query token and the indexed token differ because of casing, punctuation, or how the property's tokenizer split the text. A hyphenated identifier or an email address is the usual culprit. ## Debugging a boost Request `return_metadata=MetadataQuery(score=True, explain_score=True)` on the hybrid call and read `object.metadata.explain_score`, which breaks out what the keyword and vector halves contributed. If the keyword contribution for a document you expected to win is zero, the problem is coverage or tokenization, not the boost value — no amount of `^` tuning fixes a property that was never searched.
- A document contains the searched phrase verbatim but never ranks well. How does query_properties enter the diagnosis?Check whether the property holding that phrase is in the list at all. query_properties replaces the default set rather than extending it, so an omitted property contributes no keyword score and the document can only surface through the vector half. Confirm by running the query at alpha=0 with and without the property listed. If it still fails when listed, move on to tokenization and whether the property is searchable.
- Can you use a boost to make the semantic half favour titles over body text?No. The vector half compares one query embedding against each object's stored vector, and that vector was built at ingest time from whatever the vectorizer was configured to include. query_properties never touches it. Changing the semantic emphasis is an ingest-time change — what text goes into the embedding — followed by re-vectorizing the affected objects.
- Why might a boost that clearly helps at alpha=0 do nothing in production?Two reasons compound. First, the keyword contribution is scaled by 1 minus alpha during fusion, so at a vector-leaning alpha the whole keyword ranking has limited influence. Second, under a fusion strategy that reads only rank positions, a boost that raises a document's keyword score without changing its keyword rank contributes nothing. Verify the change at your production alpha and fusion setting, not in isolation.
saying these in an interview costs you the question
- Thinking query_properties adds to the default set rather than replacing it
- Expecting a ^ boost to influence the vector similarity score
- Assuming a listed property is searchable even without a searchable index
- Believing boosts guarantee position rather than adjusting keyword weight
- Confusing query_properties with a filter that excludes objects