skip to content

In Weaviate, what does the auto_limit (autocut) parameter do to a result list?

level: middleimportance: should knowfreq 38%

answer

  1. variable result count, not fixed
  2. cut at the drop-off
  3. relative, not an absolute threshold
  4. one group, two groups, three groups
  5. always returns something, even when nothing is good

basics

~20 s

auto_limit cuts the ranked results where similarity drops off sharply. auto_limit=1 keeps only the first cluster of closely-scoring objects, auto_limit=2 keeps the first two clusters, and so on — so the number of results varies with each query instead of being fixed.

solid answer

~50 s

`auto_limit` implements autocut: instead of taking a fixed number of results, Weaviate walks the ranked list looking for significant jumps in the distance or score between consecutive hits and cuts after the *n*th jump. So `auto_limit=1` returns the leading group of results that are all similarly close, and stops at the first big drop. A query with three excellent matches returns three; a query with twelve equally good matches returns twelve. It adapts per query, which a fixed `limit` cannot do. It is most useful for retrieval feeding a language model, where a fixed top-k either starves a rich query or pads a narrow one with irrelevant filler. It composes with `limit`, which still acts as a hard ceiling, and it is a *relative* cut, so unlike a `distance` threshold it never returns zero results just because the whole corpus is a bit far from the query. That is also its weakness: if every result is bad, autocut still hands you the least-bad group.

code

python · 13 lines
python
from weaviate.classes.query import MetadataQuery

docs = client.collections.get("Document")

response = docs.query.near_text(
    query="return policy for damaged goods",
    limit=25,          # hard ceiling
    auto_limit=1,      # keep only the leading group
    return_metadata=MetadataQuery(distance=True),
)

for o in response.objects:
    print(round(o.metadata.distance, 3), o.properties["title"])

go deeper

for a junior

Recall that auto_limit makes the number of returned results vary per query by cutting where similarity drops off, rather than always returning a fixed count like limit does.

for a middle

Explain that the value counts groups of similarly-scoring results, not objects, and that autocut is a relative cut based on gaps in the ranked distances rather than an absolute similarity threshold.

for a senior

Show when to combine it with a distance cutoff and a hard limit, and diagnose the failure cases: a uniform corpus where no jump exists, and pathological first groups that blow up payload size.

for a principal

Own the retrieval contract it implies — a variable-size context set changes downstream cost, prompt budgeting and pagination design, and every distance-derived setting must be re-evaluated as part of an embedding-model migration.

## The problem it solves Every nearest-neighbour search returns whatever you ask for. `limit=5` gives five results whether the corpus contains fifty perfect matches or none at all. Both directions hurt: - **Too few**: a query with fifteen genuinely relevant documents gets truncated at five, and downstream a model answers from a third of the available evidence. - **Too many**: a narrow query with two relevant documents gets three irrelevant ones stapled on, and the model treats them as context worth using. A fixed `distance` cutoff addresses the second case but requires an absolute threshold that is only meaningful if distance distributions are stable across queries — and for heterogeneous corpora they are not. Short queries, unusual phrasings and rare terms all shift the whole distribution. ## What autocut actually does Autocut is a *relative* cut. Weaviate ranks results as usual, then examines the sequence of distances (or scores) and looks for discontinuities — places where the gap between consecutive results is large compared to the gaps before it. `auto_limit=n` means "cut after the *n*th such jump". Picture a ranked list with these cosine distances: ``` 0.11 0.12 0.13 |jump| 0.34 0.35 |jump| 0.61 0.62 0.63 ``` - `auto_limit=1` → the first three objects. - `auto_limit=2` → the first five. - `auto_limit=3` → all eight. The cut point is discovered per query. The same `auto_limit=1` returns three results for one query and eleven for another, which is the whole point. ## Using it ```python response = collection.query.near_text( query="return policy for damaged goods", limit=25, auto_limit=1, return_metadata=MetadataQuery(distance=True), ) ``` Two habits make it safe in production. **Keep a `limit` as a ceiling.** Autocut decides where to cut, but if the corpus contains hundreds of near-identical documents the first "group" can be enormous. A `limit` bounds the payload, the token count, and the latency. **Request distance metadata while tuning.** You cannot judge whether the cut landed sensibly without seeing the distances it cut between. Log them during evaluation, then drop the metadata once you trust the setting. ## Autocut versus a distance threshold They fail in opposite directions, and that is how you choose. A `distance` cutoff is absolute: it can return zero results, which is exactly what you want when the honest answer is "we have nothing about that". Its weakness is that the right absolute value depends on the embedding model, the corpus, and often on query length. Autocut is relative: it always returns *something*, because there is always a leading group. Its weakness is that a query with no good answers still yields a confident-looking cluster of bad ones. It carries no notion of "good enough", only of "noticeably better than what follows". That suggests combining them: a generous `distance` cutoff to reject the genuinely hopeless, plus `auto_limit=1` to trim what survives to the useful group, plus `limit` as a hard ceiling. Each covers the other's blind spot. ## Where it fits and where it does not Good fits: - RAG context assembly, where a variable number of chunks is more honest than a fixed k. - Search UIs that show "top results" rather than a numbered page. - Queries whose relevant-result count genuinely varies — a product catalogue where some searches have one right answer and others have forty. Poor fits: - Paginated interfaces. Pagination needs stable, offset-addressable result counts; a per-query variable cut fights that model. - Candidate generation for a reranker, where you deliberately want a fixed, generous pool and let the reranker do the quality judgment. - Uniform corpora where distances are tightly clustered by construction. If there are no meaningful jumps, autocut has nothing to detect and either returns nearly everything or cuts arbitrarily. ## Tuning Start at `auto_limit=1` — it is the setting that matches the intuition of "just the good ones". Move to 2 when evaluation shows the first group is too tight and you are dropping useful evidence at the first minor gap. Values above 3 rarely mean much, because by then you are including groups the engine already judged to be distinctly worse. As with every distance-derived setting, re-evaluate after any embedding-model change: the shape of the distance distribution is a property of the vector space, and a new model can make jumps appear or disappear.

  • Why keep a limit alongside auto_limit rather than relying on the cut alone?
    Because the leading group has no inherent size bound. A corpus with hundreds of near-duplicate documents can produce a first group of hundreds, blowing up payload size, token counts and latency. limit acts as a hard ceiling while autocut decides where within that ceiling to stop, so the two are complementary rather than redundant.
  • When would you prefer an absolute distance cutoff over autocut?
    When returning nothing is a valid and desirable answer — for example a support-article search that should say "no relevant article" rather than showing the least-bad one. Autocut is relative and always yields a leading group, so it cannot express "nothing here is good enough". Combining a loose distance cutoff with auto_limit=1 gets both behaviours.
  • Your corpus is highly uniform and autocut returns nearly everything. What is happening?
    Autocut needs discontinuities in the ranked distances to find a cut point. If documents are near-duplicates or the embedding space is tightly clustered, consecutive gaps are all similar and no jump stands out, so the first group swallows most of the list. A fixed limit, or a distance cutoff tuned against measured distances, is the better tool there.

saying these in an interview costs you the question

  • Thinking auto_limit sets a fixed number of results
  • Confusing autocut's relative cut with an absolute distance cutoff
  • Expecting autocut to return nothing when all matches are poor
  • Using autocut behind a paginated result interface
  • Dropping limit entirely once autocut is enabled

context