skip to content

After adding fq=category:books, a Solr facet on category returns only that one bucket — why, and how do you keep the others?

level: seniorimportance: should knowfreq 56%

answer

  1. ask what document set the facet counts over
  2. the filter already removed the other values
  3. filters can be named, then ignored
  4. tag the fq, exclude the tag on that facet

basics

~20 s

Facets are computed over the documents remaining after all filter queries, so filtering on category leaves only that value to count. Tag the filter and exclude it from that facet, with fq={!tag=cat} plus facet.field={!ex=cat}, to restore the full list.

solid answer

~50 s

Solr computes facet counts over the base document set, which is the query intersected with every `fq`. Once `fq=category:books` is applied, no document outside that category survives, so the `category` facet can only report one bucket — the counts are correct, they are just counted over the wrong set. The fix is filter exclusion: name the filter with a local param, `fq={!tag=cat}category:books`, and tell the facet to ignore it, `facet.field={!ex=cat}category`. The excluded facet is then computed over a set that still honours `q` and every other `fq`, so a user sees the alternative categories while their other filters still narrow the counts. In the JSON Facet API the same thing is expressed as `domain: { excludeTags: "cat" }` on that facet. Facets on *other* fields should keep the category filter applied, which is why exclusion is per facet rather than global.

code

bash · 10 lines
bash
# without exclusion: the category facet collapses to one bucket
q=laptop&fq=category:books&facet=true&facet.field=category

# with tag/ex: the category facet ignores its own filter, the brand filter still applies
q=laptop
&fq={!tag=cat}category:books
&fq={!tag=brand}brand:acme
&facet=true
&facet.field={!ex=cat}category
&facet.field={!ex=brand}brand

go deeper

for a junior

Know that facet counts are computed over the documents left after the query and all filter queries, so filtering on a field naturally collapses that field's own facet.

for a middle

Be able to write the tag and exclude local params correctly and explain that the excluded facet still honours the query and the other filters. Knowing the JSON Facet domain excludeTags equivalent is expected.

for a senior

Show that you have built a faceted UI: exclusion applied per facet rather than globally, awareness that each distinct exclusion set costs another base document set, and that distributed facet counts are approximate unless you refine. Have an opinion on docValues and field type for facet performance.

for a principal

Decide how much facet accuracy the product actually needs and what it costs. Exact counts across many shards mean refinement round trips on every page; approximate counts are usually fine for navigation and unacceptable for reporting, and that distinction should shape the API you expose.

## How Solr computes a facet count A Solr request produces a base `DocSet`: the documents matching `q`, intersected with the `DocSet` of every `fq`. Faceting then counts term occurrences within that base set. This is the whole explanation for the symptom — after `fq=category:books`, every surviving document has `category:books`, so a terms facet on `category` has exactly one non-empty bucket. Nothing is broken; the facet is answering a question about a set the filter already narrowed. ## Multi-select faceting and filter exclusion What a faceted UI actually wants is: "count categories as if the category filter were not applied, but with all the *other* filters applied." Solr expresses this with tagging and exclusion, using local params. ``` q=laptop fq={!tag=cat}category:books fq={!tag=brand}brand:acme facet=true facet.field={!ex=cat}category facet.field={!ex=brand}brand facet.field=price_range ``` `{!tag=cat}` labels the filter. `{!ex=cat}` on a facet asks Solr to compute that facet over a base set built without the tagged filter. The `category` facet therefore still respects `q=laptop` and the brand filter, but shows every category the user could switch to. The `price_range` facet excludes nothing, so it reflects the fully filtered set — which is correct, because the user has not selected a price yet. You can exclude several tags at once with a comma-separated list, `{!ex=cat,brand}`. The cost is real: each distinct exclusion set forces Solr to build a different base DocSet, so a page with many excluded facets does more set intersection work than one with none. ## The JSON Facet API version The JSON Facet API expresses the same idea more legibly and adds nesting and metrics: ```json json.facet={ categories: { type: terms, field: category, limit: 20, domain: { excludeTags: "cat" }, facet: { avg_price: "avg(price)", in_stock: { type: query, q: "stock:[1 TO *]" } } } } ``` The `domain` block is the general mechanism for changing what a facet counts over: `excludeTags` drops tagged filters, `filter` adds extra constraints for this facet only, and `blockChildren`/`blockParent` move the domain across a nested-document join. `facet` nests sub-facets and aggregations inside each bucket, so you get an average price per category in one request — something the classic parameter API cannot express without a separate query per bucket. Aggregation functions include `sum`, `avg`, `min`, `max`, `stddev`, `percentile`, `unique` and `hll`. Two defaults differ between the APIs and catch people out. The classic `facet.mincount` defaults to 0, so zero-count buckets appear; the JSON terms facet defaults `mincount` to 1. And `facet.limit` defaults to 100 in the classic API, so a naive facet on a high-cardinality field silently truncates. ## Accuracy in a distributed collection With more than one shard, each shard returns its own top buckets and the coordinator merges them. A term that ranks just below a shard's cut-off contributes nothing from that shard, so a merged count can be too low and a bucket can be missing entirely. Two mitigations: over-request (ask each shard for more buckets than you will show, tuned with `overrequest`), and set `refine: true` on a JSON facet, which makes the coordinator go back to the shards for exact counts of the buckets it intends to return. Refinement costs a second round trip and is the honest answer when counts must be exact. ## Method and cost The classic `facet.method` chooses the algorithm: `fc` builds per-field counts from docValues or an uninverted view, `enum` intersects a filter per term through the filterCache (sensible only for very low cardinality), `uif` uninverts the field into memory, and `fcs` is the per-segment variant for fields that change frequently. The JSON API's `method` offers `dv`, `dvhash`, `enum`, `stream` and `smart`, with `smart` picking for you. In practice the biggest wins are declaring `docValues="true"` on every field you facet on, and faceting on a `string` field rather than an analyzed text field so buckets are readable values instead of tokens. ## The diagnostic habit When a facet looks wrong, ask two questions in order: what base set is this facet counting over (which filters apply, which are excluded), and is the collection sharded such that merge truncation could be hiding a bucket. Those two account for nearly every "my facet counts are wrong" report.

  • Why exclude the filter only on its own facet rather than on all of them?
    Because the other facets should reflect what the user has already chosen. If a category filter is excluded from the brand facet too, the brand counts describe a set the user is not looking at and clicking one gives a surprising result count. Exclusion is per facet precisely so each one can answer "what would happen if I changed this dimension" while holding the rest fixed.
  • In a sharded Solr collection, why can a facet count come back slightly low?
    Each shard returns only its own top buckets, so a term sitting just below one shard's cut-off contributes zero from that shard to the merged count, and a globally significant term can be missing entirely. Over-requesting more buckets per shard reduces the risk; setting `refine: true` on a JSON facet makes the coordinator re-query the shards for exact counts of the buckets it will return, at the cost of an extra round trip.
  • What does the JSON Facet API give you that the classic facet parameters do not?
    Nesting and metrics. A terms facet can contain sub-facets and aggregation functions such as `avg(price)` or `unique(user)`, computed per bucket in one request instead of one query per bucket. The `domain` block also generalises filter exclusion, per-facet extra filters, and moving the domain across nested parent and child documents, which the flat parameter API cannot express.
  • Why is docValues on the faceted field the first thing to check for facet performance?
    Faceting needs a per-document view of field values. With `docValues="true"` that column already exists on disk and is memory-mapped; without it Solr must uninvert the index into heap at query time, which is slow on the first request after a commit and a common source of heap pressure. Faceting on a string field with docValues, fed by a copyField from the analyzed text field, is the standard shape.

saying these in an interview costs you the question

  • Says the facet counts are simply a Solr bug
  • Applies the filter as a query clause instead of tagging it
  • Excludes the tag on every facet in the request
  • Assumes facet counts are exact across shards by default
  • Facets on an analyzed text field and expects readable buckets

context