Why does boosting the title field ten times often make search relevance worse?
answer
- only the ratio between weights carries meaning
- field contributions are not on one scale
- short fields already score high before boosting
- consider best-field instead of summing fields
basics
~20 sBecause a boost multiplies that field's score contribution, and field contributions are not on a comparable scale. A ten-times title boost lets a weak, incidental title match outscore a document that matches the query strongly everywhere else.
solid answer
~50 sPer-field boosts are multipliers on a field's score contribution, and those contributions were never comparable to begin with. A title is short, so length normalisation already inflates its score, and title terms are often rarer, so their IDF is higher too. Multiplying an already-large number by ten means a document whose title merely contains one query term beats a document matching the entire phrase in the body. Boosts are also purely **relative**: scaling every field's weight by the same factor changes nothing, so only the ratios matter — which is why "boost of 10" is meaningless without saying what it is ten times bigger than. In practice you start near 1, move in small steps against a fixed judgment set, and decide deliberately whether to **sum** field scores or take the **best** field's score, because summing rewards documents that match many fields shallowly.
code
text · 5 lines# sum-of-fields: rewards matching a little of the query in many fields
score = 3.0 * bm25(title, q) + 1.0 * bm25(body, q)
# best-field: one field must explain the whole query
score = max(3.0 * bm25(title, q), 1.0 * bm25(body, q))go deeper
Recall that a field boost is a multiplier on that field's contribution and that only the ratio between field weights matters, so scaling them all changes nothing.
Explain why an unboosted title already scores high — short-field length normalisation plus higher IDF terms — and why stacking a large multiplier on top silences the other fields entirely.
Demonstrate a tuning method rather than instincts: a held-out judgment set, a mechanical search over a small weight space, and a deliberate choice between summed, best-field and pooled combination.
Own the config's long-term health. Argue for few, documented weights with owners and review dates over a growing pile of per-complaint multipliers nobody can justify a year later.
## What a per-field boost actually multiplies In a multi-field search, each field produces its own relevance contribution from the query terms it matched, and those contributions are combined into one number. A per-field boost is a coefficient in that combination: title weight 3 means the title's contribution is tripled before combining. It is not a percentage of the final score, not a guarantee of position, and not a statement about how much the user cares about titles. The first consequence is that boosts are **scale-free**. If the score is `w_title * s_title + w_body * s_body`, then multiplying both weights by the same constant multiplies every document's score by that constant and changes no ordering at all. Only the ratio carries information. A candidate who says "I set the title boost to 10" without naming the other weights has said nothing. ## Why field scores are not comparable Ranking functions in the BM25 family normalise term frequency by field length relative to the average length of that field. A title is a handful of tokens, so a single occurrence of a query term represents a large share of the field and scores near the saturation ceiling. A body is hundreds of tokens, so the same single occurrence is diluted. Title terms also tend to be more distinctive, which raises their inverse document frequency. The upshot is that an *unboosted* title contribution is already systematically larger than an unboosted body contribution for the same query. So a ten-times title boost is not "titles matter ten times more than bodies" — it is roughly "titles matter ten times more than something that already mattered several times more". At that ratio the body has effectively stopped participating in ranking. A document that mentions one query term in a long title now outranks a document that matches the whole query as a phrase in the body, which is precisely the complaint the boost was added to fix. ## Sum of fields versus best field The combination rule matters as much as the weights. Two common shapes: ``` # sum-of-fields: rewards documents matching a bit of the query in many fields score = w_title * s_title + w_body * s_body + w_tags * s_tags # best-field: one field must carry the match score = max(w_title * s_title, w_body * s_body, w_tags * s_tags) ``` For a query like *red running shoes*, sum-of-fields happily rewards a product with "red" in the colour field, "running" in the category field and "shoes" in the type field — no single field means what the user asked, yet the sum is large. Best-field prefers the product whose title is "Red Running Shoes", which is what the user meant. Conversely, a person-name query spanning `first_name` and `last_name` genuinely spans fields, and a cross-field combination that pools the query terms across fields is the right shape there. ## BM25F: the principled version The field-weighting problem was solved properly a long time ago by BM25F, which weights and length-normalises each field's term frequency **before** applying the saturation and IDF, rather than scoring each field independently and adding the results: ``` tf_pooled(t) = sum over fields f of w_f * tf(t, f) / (1 - b_f + b_f * len_f / avg_len_f) score(q) = sum over terms t of idf(t) * tf_pooled(t) / (k1 + tf_pooled(t)) ``` Because saturation is applied once to the pooled frequency, a term appearing in three fields cannot be counted three times at full value, and one field's boost cannot escape the saturation curve. This is why field-pooled scoring behaves much more sanely under aggressive weights than naive weighted summation does. ## How to pick the numbers Hand-tuning by eye is how ranking configs become archaeology. The disciplined loop is: 1. Fix a set of queries with known good results — the judgment set — before you touch anything. 2. Restrict yourself to a handful of fields. Three or four weights you can defend beat fifteen you cannot. 3. Search the weight space mechanically — a coarse grid, or coordinate ascent that optimises one weight at a time and repeats — instead of nudging numbers after each complaint. 4. Re-run the search whenever the analyzer chain or the corpus composition changes, because the optimum moves when field lengths and term statistics move. And watch for overfitting: a weight vector tuned on thirty queries from one stakeholder will win on those thirty queries and lose on the tail. Hold out queries you never tune against. ## The symptom-driven anti-pattern The most common real-world path to a broken config is one boost per bug report. Someone complains that a document ranks too low, a boost is added for whatever property that document has, and the config grows by one unexplained multiplier. Six months later nobody can say why `tags` has weight 7, and every attempt to fix a new complaint breaks two old ones. Weights should come out of a measured search over a judgment set, and each one should carry a written rationale.
- When is summing field scores wrong and best-field right?Summing rewards a document that matches a few query terms shallowly across many fields; best-field rewards a document where one field explains the whole query. For "red running shoes", a product whose title contains all three terms should beat one with each term scattered across colour, category and type fields — that is best-field behaviour. Cross-field pooling is the third option when the query legitimately spans fields, such as a first name and a last name.
- How do you actually choose the weight values?Fix a judgment set of queries with known good results, then search the weight space mechanically — a coarse grid or coordinate ascent over three or four fields — instead of hand-editing after each complaint. Keep the weights few; a config with fifteen tuned fields is overfitted to whichever queries were in the room. Re-run the search when the analyzer or corpus changes, because the optimum moves with field lengths and term statistics.
Field weights are like the sliders on a graphic equaliser: pushing one to maximum doesn't make the mix louder, it just makes everything else inaudible.
saying these in an interview costs you the question
- Treats a boost as a percentage of the final score
- Believes doubling every field weight changes the ranking
- Ignores that short fields already score high before boosting
- Tunes weights by eye with no held-out judgment set
- Adds a new field boost for every individual bug report