In Weaviate, what does a generative-* module add to a collection?
answer
- retrieval and generation in one request
- two prompt shapes, two cost profiles
- one call per object versus one per query
- no seam for reranking
- LLM latency inside a database call
basics
~20 sA generative module attaches an LLM to a collection so retrieval and generation happen in one round-trip: the search runs, the retrieved objects are fed into your prompt server-side, and the model's text comes back with the results.
solid answer
~50 sYou configure it at collection creation with `generative_config=Configure.Generative.openai()`, and then query through `collection.generate.*` instead of `collection.query.*`. The generate methods take the usual search arguments plus two prompt shapes. `single_prompt` runs the model once per returned object, with `{property}` placeholders interpolated from that object — good for per-result summaries or extractions. `grouped_task` runs the model once over the whole result set — good for the answer-from-sources shape. The response carries both: a collection-level `generated` field for the grouped task and a per-object `generated` field for the single prompt. The appeal is that one request does retrieve-and-generate, with no orchestration code and no second network hop. The limits are equally real: prompt construction is Weaviate's, so you do not control context ordering, truncation or a reranking step in between, and the LLM call now sits inside your database's request path. Teams that need control over the prompt usually retrieve with a normal query and call the model themselves.
code
python · 11 linesfrom weaviate.classes.config import Configure, Property, DataType
client.collections.create(
name="Article",
properties=[
Property(name="title", data_type=DataType.TEXT),
Property(name="body", data_type=DataType.TEXT),
],
vectorizer_config=Configure.Vectorizer.text2vec_openai(),
generative_config=Configure.Generative.openai(),
)go deeper
Know that a generative module lets a search also return LLM-generated text, configured on the collection and queried through the generate methods rather than the plain query methods.
Distinguish the two prompt shapes precisely: single_prompt once per object with property interpolation, grouped_task once over the whole result set, and know both can be returned together.
Weigh the operational cost — LLM latency and failure inside a database request, generation cost scaling with limit, and the missing seam where reranking or a relevance threshold would go.
Own the architectural call. Fusing retrieval and generation buys speed of delivery and costs you prompt control, observability and portability; decide deliberately when the product crosses the line into needing its own generation service.
## What the module is Weaviate's `generative-*` modules wire a text-generation model into the query path. Where a `text2vec-*` module turns text into vectors, a generative module turns retrieved objects into prose. It is configured per collection at creation time, alongside the vectorizer, and it does not affect how anything is stored or indexed — it only adds a second style of query. ## Configuring and calling it At creation: `generative_config=Configure.Generative.openai()`. As with vectorizer modules, the credential travels as a connection header rather than living in the schema, and the module must be enabled on the server. At query time you switch namespace. Every search you can run under `collection.query` has a counterpart under `collection.generate` — same search arguments, same filters, same limit — plus the prompt arguments. So `collection.generate.near_text(query=..., limit=3, grouped_task=...)` is "do that search, then do this with the results". ## The two prompt shapes **`single_prompt`** is evaluated once per returned object. Placeholders in braces are filled from that object's properties, so `"Summarise {title} in one sentence."` produces one generation per result. Cost scales with `limit`: ten results is ten model calls. Use it for per-item transformation — summaries, tag extraction, tone rewrites. **`grouped_task`** is evaluated once, over all returned objects together. This is the retrieval-augmented answering shape: retrieve k passages, ask one question about them, get one answer. Cost is one call regardless of `limit`, but the context window bounds how many objects can meaningfully be included. Both can be supplied in one request. The result object then carries a top-level `generated` string for the grouped task and a per-object `generated` string on each result for the single prompt, so you can render an overall answer above a list of per-item summaries from a single round-trip. ## What you gain One network round-trip instead of two, and no orchestration layer for the simple case. For a prototype, an internal search tool, or a demo, this collapses a whole service into a query argument. It also keeps retrieved text off your application server entirely — the passages go from the database to the model without a hop through your code, which occasionally matters for data-handling reasons. ## What you give up *Prompt control.* Weaviate assembles the context from the retrieved objects. You do not choose the ordering, the delimiters, how properties are laid out, or what happens when the retrieved text exceeds the model's context window. Retrieval quality work — deduplicating near-identical chunks, interleaving sources, putting the strongest evidence last — has nowhere to live. *No step in between.* The standard improvement to a naive retrieve-then-generate pipeline is to insert something between the two: a cross-encoder rerank, a relevance threshold, a fallback when nothing scores well, a query rewrite. Fusing the two calls removes the seam where that step would go. *Coupled failure and latency.* The LLM call is now inside a database request. Generation latency is measured in seconds and its failure modes — rate limits, timeouts, content filters — become database query failures. A slow model holds a database connection open. *Weaker observability.* You see the request and the generated string. Logging the exact prompt sent, replaying it, diffing it across versions, or attributing cost per prompt template is harder when the prompt was assembled elsewhere. *Portability.* Search that returns objects is portable across stores; search that returns generated prose is a Weaviate-shaped API your application now depends on. ## The pragmatic line Generative modules are excellent for getting something working and for read-only internal tools where prompt quality is not the product. Once answer quality is the product — once you are iterating on prompts, adding reranking, measuring groundedness, or switching model providers — the usual move is to retrieve with a plain query and own the generation step in your own service. Nothing about the collection has to change to make that switch: the generative config is inert for queries that do not use `generate`, so both styles can coexist while you migrate.
- How does the cost of single_prompt scale compared with grouped_task?single_prompt is one model call per returned object, so raising limit from 3 to 20 multiplies generation cost and latency by roughly seven. grouped_task is one call whatever the limit, bounded instead by the context window — past a certain number of passages you are paying for tokens the model cannot usefully attend to. Pick by shape, then tune limit with cost in mind.
- Why do teams often stop using generative modules once answer quality becomes the product?Because the prompt is assembled by Weaviate. There is no place to rerank, deduplicate, threshold on relevance, control context ordering, or handle 'nothing good was retrieved'. Owning the generation call in your own service restores that seam, and the collection needs no change — generative config is simply unused by plain query calls.
- What happens to a query if the generative model call fails?The database request fails with it, because generation happens inside the query. A provider rate limit or timeout surfaces as a failed search rather than as a degraded answer over successful retrieval. If you want retrieval to survive generation problems, the two calls have to be separate in the first place.
saying these in an interview costs you the question
- Thinking the generative module changes how objects are stored or indexed
- Believing grouped_task runs once per returned object
- Assuming you can edit the prompt template Weaviate builds around your text
- Expecting retrieval to still succeed when the LLM call fails
- Treating generate as free because it is one request