skip to content

How does group_by change the result shape of a Weaviate near_text query?

level: seniorimportance: nice to knowfreq 28%

answer

  1. ten chunks, one document
  2. diversity rather than relevance
  3. groups plus a flat list, both returned
  4. cap the objects inside each bucket
  5. the candidate pool must be generous

basics

~10 s

Passing group_by=GroupBy(prop=..., number_of_groups=..., objects_per_group=...) folds the ranked hits into groups keyed by a property value. The response then exposes groups alongside the flat object list, and each object carries the group it belongs to.

solid answer

~50 s

`group_by` turns a flat ranked list into a grouped one without changing how the search itself runs. You pass `GroupBy(prop="category", number_of_groups=3, objects_per_group=2)` to `near_text`, `near_vector` or `near_object`, and Weaviate takes the ranked hits, buckets them by that property's value, keeps at most `number_of_groups` groups and at most `objects_per_group` objects inside each. The response shape changes: you get a `groups` mapping keyed by the property value, each group holding its own objects, plus a flat `objects` list where every object records which group it belongs to. The practical use is diversity. A plain top-10 over chunked documents frequently returns ten chunks from the same document, because the whole document is on-topic. Grouping by the source-document property and capping objects per group guarantees you see several distinct sources — the classic fix for a retrieval stage that keeps feeding one document's chunks into a prompt.

code

python · 19 lines
python
from weaviate.classes.query import GroupBy

chunks = client.collections.get("Chunk")

response = chunks.query.near_text(
    query="refund window",
    limit=50,
    group_by=GroupBy(
        prop="document_id",
        number_of_groups=5,
        objects_per_group=2,
    ),
)

for name, group in response.groups.items():
    print(name, len(group.objects))

for o in response.objects:
    print(o.belongs_to_group, o.properties["text"][:60])

go deeper

for a junior

Know that group_by buckets the search results by a property value and caps how many objects come back per bucket, and that the response then exposes groups as well as the flat object list.

for a middle

Explain that grouping is applied after the search rather than changing it, what number_of_groups and objects_per_group each bound, and why limit must be generous for the groups to fill.

for a senior

Show the production motivation: chunked corpora returning ten chunks of one document, and how grouping by source id fixes it more cheaply than client-side deduplication or an extra reranking call.

for a principal

Own the retrieval-contract choice between grouping at query time, one vector per document, and a reranker that enforces diversity — each moves cost to a different layer — and insist that isolation stays a filtering and tenancy concern, never a grouping one.

## The failure that motivates it Split long documents into chunks, embed each chunk, and search. A query that matches one document well matches *many of its chunks* well, because they share vocabulary and context. The top ten results are then ten chunks from the same source, and the retrieval stage has effectively returned a single document while pretending to return ten results. Any downstream consumer — a language model, a search UI, a human reviewer — sees a narrow slice of the corpus and no indication that the narrowness is an artefact of chunking rather than of the data. Deduplicating client-side is possible but wasteful: you must over-fetch by an unknown factor and hope enough distinct sources appear in the fetched window. ## What group_by does ```python from weaviate.classes.query import GroupBy response = collection.query.near_text( query="refund window", limit=50, group_by=GroupBy( prop="document_id", number_of_groups=5, objects_per_group=2, ), ) ``` The search runs exactly as it would without grouping — same query vector, same filters, same index. Grouping is applied to the ranked results: - `prop` names the property whose value forms the group key. - `number_of_groups` caps how many distinct groups come back. - `objects_per_group` caps how many hits are kept inside each group. So the example asks: give me the five best distinct documents, with up to two chunks from each. That is a fundamentally different retrieval contract from "give me the ten best chunks". ## The response shape A grouped query does not return the same object as an ungrouped one, and code that assumes otherwise breaks: - `response.groups` is a mapping from group value to a group object holding that group's own objects and a count. - `response.objects` is still a flat list, but each object records which group it belongs to. Iterate `groups` when the grouping is the point — rendering "3 results from Contract A, 2 from Contract B". Iterate `objects` when you want a flat list that happens to be diversified, which is the common case for assembling model context. ## Choosing the grouping property The property must exist on the collection and be indexed for filtering — grouping resolves values through the same inverted index that filters use. Beyond that, cardinality drives usefulness: - **Source identifier** (document id, url, product id) — the canonical use. Delivers diversity across origins. - **Category or type** — gives coverage across kinds of content: one policy page, one FAQ, one forum thread. - **Tenant or author** — mostly a modelling smell. If a query should be scoped to a tenant, filter on it rather than grouping by it; grouping is about diversity, not isolation. - **A near-unique property** — degenerate. If almost every object has its own value, every group holds one object and grouping bought nothing beyond a slower path to the same list. ## Interaction with limit `limit` still bounds the candidate pool that grouping operates on, so it must be generous relative to what you want out. Ask for `number_of_groups=5` with `limit=5` and you may find fewer than five distinct groups exist inside those five hits. A workable rule is to set `limit` to several times `number_of_groups * objects_per_group`, then verify against real queries that the groups actually fill. ## Grouping versus the alternatives **Client-side deduplication**: over-fetch, then keep the first *n* per source. Works, but the over-fetch factor is a guess, and a pathological query can return a window entirely occupied by one source. **One vector per document instead of per chunk**: removes the duplication problem at the source, but loses the passage-level precision that made chunking worth doing. **Reranking**: a reranker can enforce diversity as part of its scoring, but it costs an extra model call and still needs a diverse candidate set to work from — which grouping can supply. Grouping at query time is the cheapest of the three when the duplication is structural, because it happens inside the database with no extra round trip. ## What to watch in production The grouped result count is bounded by `number_of_groups * objects_per_group`, but it is not guaranteed to reach it — a query may simply not have hits spread across that many groups. Downstream code that assumes a fixed context size must handle a short list. And grouping does not change ranking quality. If the underlying search returns poor chunks, grouping returns poor chunks from more sources. It is a diversity tool, not a relevance tool.

  • Your grouped query asks for five groups but returns only two. What went wrong?
    Almost always the candidate pool. Grouping operates on the ranked hits that limit produced, so if those hits span only two distinct property values there are only two groups to form. Raise limit to several times number_of_groups times objects_per_group. If it persists, the corpus genuinely lacks diversity for that query.
  • Would you group by tenant to keep tenants separate in results?
    No — that conflates diversity with isolation. Grouping shapes a result list; it does not enforce a security boundary and could still surface another tenant's objects. Scope the query with a filter on the tenant property, or use the collection-level tenancy features, and reserve grouping for diversity within an already-scoped result set.
  • Does grouping improve the relevance of what comes back?
    No. The search runs identically; grouping only reorganises and trims the ranked hits. If the underlying retrieval is poor, grouping returns poor results from more distinct sources. Fix relevance at the embedding, chunking or query level, then use grouping to stop one source dominating the list.

saying these in an interview costs you the question

  • Thinking group_by re-runs or re-ranks the vector search
  • Expecting the grouped count to always reach groups times objects
  • Grouping by tenant instead of filtering for isolation
  • Grouping on a near-unique property and expecting real groups
  • Setting limit no larger than the number of groups requested

context